helloadhavan/python-docstrings
Python Docstring Diff Dataset This dataset contains training samples for models that generate Python documentation patches. Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation. The dataset is designed for training or evaluating language models that assist with: Automatic code documentation Docstring generation Code review automation Developer tooling Dataset Structure Each entry contains… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/python-docstrings.
254
1---2dataset_info:3 features:4 - name: instruction5 dtype: string6 - name: code7 dtype: string8 - name: response9 dtype: string10 - name: file11 dtype: string12 splits:13 - name: train14 num_bytes: 24783293415 num_examples: 1644016 download_size: 8643184017 dataset_size: 24783293418configs:19- config_name: default20 data_files:21 - split: train22 path: data/train-*23license: mit24task_categories:25- text-generation26language:27- en28tags:29- code30pretty_name: python docstring dataset31size_categories:32- 10K<n<100K33---34# Python Docstring Diff Dataset35 36This dataset contains training samples for models that generate Python documentation patches.37Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation.38 39The dataset is designed for training or evaluating language models that assist with:40 41* Automatic code documentation42* Docstring generation43* Code review automation44* Developer tooling45* Dataset Structure46 47Each entry contains the following fields:48 49Field Description50-------------------51instruction| Task instruction given to the model52code| Python source code with docstrings removed53response| A unified diff patch that adds the correct docstrings54file| Original file path from the source project55 56## Task Format57 58The model receives a Python file missing its documentation and must produce a unified diff that adds appropriate docstrings.59 60Example input:61 62```python63def load_json(path):64 with open(path) as f:65 return json.load(f)66```67 68Example expected output:69```diff70--- a/file.py71+++ b/file.py72@@73 def load_json(path):74+ """Load JSON data from a file path."""75 with open(path) as f:76 return json.load(f)77```78 79## Data Sources80 81The dataset was generated by scanning Python packages in github.82Docstrings were extracted from functions, classes, async functions, methods, and modules using Python's AST parser.83Low-quality documentation was filtered out using heuristics such as:84 85* Minimum docstring length86* Removal of TODO or placeholder documentation87* Deduplication of similar examples88 89## Intended Use90 91This dataset is useful for training models that perform:92 93* automatic docstring generation94* documentation patch creation95* codebase documentation improvement tools96* AI-assisted code review systems97 98## License99 100This dataset is released under the MIT License.