ClassEval
ClassEval
Dataset Card for FudanSELab ClassEval
Dataset Summary
We manually build ClassEval of 100 class-level Python coding tasks, consists of 100 classes and 412 methods, and average 33.1 test cases per class.
For 100 class-level tasks, diversity is maintained by encompassing these tasks over a wide spectrum of topics, including Management Systems, Data Formatting, Mathematical Operations, Game Development, File Handing, Database Operations and Natural Language Processing.
For… See the full description on the dataset page: https://huggingface.co/datasets/FudanSELab/ClassEval.modified-classeval-code-summarization
Modified ClassEval (MCE) Dataset
This dataset is a modification of the ClassEval benchmark, designed for evaluating code summarization models beyond the function level. It explores the impact of function and class contexts on summary quality. The dataset includes modifications for evaluating at both function and class levels.
Paper: Code Summarization Beyond Function Level
Dataset Structure:
The dataset contains samples with the following fields:
class_id: Identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/sm1rk/modified-classeval-code-summarization.ClassEval
ClassEval (modernised)
A lightly patched fork of FudanSELab/ClassEval,
the 100-task class-level Python code generation benchmark, fixed so that it
still runs correctly on a current Python and a current NumPy.
The benchmark itself is unchanged. Every patch either repairs a test that no
implementation could pass, or repairs the reference solution. No task was made
easier, no prompt (skeleton) was touched, and nothing about what a model is
asked to write has changed.… See the full description on the dataset page: https://huggingface.co/datasets/ilintar/ClassEval.ClassEval-LineInfilling
ClassEval-LineInfilling
ClassEval-LineInfilling is a derived fill-in-the-middle benchmark constructed from the original FudanSELab/ClassEval dataset. The original ClassEval benchmark contains human-crafted class-level Python programming tasks with class skeletons, unit tests, ground-truth implementations, and method-level metadata.
This derived dataset reformulates ClassEval into a line-level infilling task. For each non-empty line in the body of every non-constructor method, we… See the full description on the dataset page: https://huggingface.co/datasets/annachaaang/ClassEval-LineInfilling.ClassEvalQABenchmarkclasseval-structured-v1
ClassEval Structured v1
Derived from FudanSELab/ClassEval (Du et al. 2023, arXiv:2308.01861).
License: CC BY-NC 4.0 (upstream data license). Non-commercial use only.
One row per (task_id, variant) with variant in {1,2,3} (100 tasks × 3 = 300 rows; Hub split test).
solution_code, test, and methods_info_json come from upstream.
We add rendered prompts/targets and stratification fields.
Missing bodies use ....
Variants:
Signatures and docstrings kept; every method body is ....… See the full description on the dataset page: https://huggingface.co/datasets/dhruveshpatel/classeval-structured-v1.
