Team Ai
Datasetpublic

helloadhavan/python-docstrings

Python Docstring Diff Dataset This dataset contains training samples for models that generate Python documentation patches. Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation. The dataset is designed for training or evaluating language models that assist with: Automatic code documentation Docstring generation Code review automation Developer tooling Dataset Structure Each entry contains… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/python-docstrings.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
2likes54downloads
README.md100 linesDownload Raw Back to root
1---2dataset_info:3  features:4  - name: instruction5    dtype: string6  - name: code7    dtype: string8  - name: response9    dtype: string10  - name: file11    dtype: string12  splits:13  - name: train14    num_bytes: 24783293415    num_examples: 1644016  download_size: 8643184017  dataset_size: 24783293418configs:19- config_name: default20  data_files:21  - split: train22    path: data/train-*23license: mit24task_categories:25- text-generation26language:27- en28tags:29- code30pretty_name: python docstring dataset31size_categories:32- 10K<n<100K33---34# Python Docstring Diff Dataset35 36This dataset contains training samples for models that generate Python documentation patches.37Each example provides a Python source file with its docstrings removed and a corresponding unified diff patch that restores the documentation.38 39The dataset is designed for training or evaluating language models that assist with:40 41* Automatic code documentation42* Docstring generation43* Code review automation44* Developer tooling45* Dataset Structure46 47Each entry contains the following fields:48 49Field	Description50-------------------51instruction|	Task instruction given to the model52code|	Python source code with docstrings removed53response|	A unified diff patch that adds the correct docstrings54file|	Original file path from the source project55 56## Task Format57 58The model receives a Python file missing its documentation and must produce a unified diff that adds appropriate docstrings.59 60Example input:61 62```python63def load_json(path):64    with open(path) as f:65        return json.load(f)66```67 68Example expected output:69```diff70--- a/file.py71+++ b/file.py72@@73 def load_json(path):74+    """Load JSON data from a file path."""75     with open(path) as f:76         return json.load(f)77```78 79## Data Sources80 81The dataset was generated by scanning Python packages in github.82Docstrings were extracted from functions, classes, async functions, methods, and modules using Python's AST parser.83Low-quality documentation was filtered out using heuristics such as:84 85* Minimum docstring length86* Removal of TODO or placeholder documentation87* Deduplication of similar examples88 89## Intended Use90 91This dataset is useful for training models that perform:92 93* automatic docstring generation94* documentation patch creation95* codebase documentation improvement tools96* AI-assisted code review systems97 98## License99 100This dataset is released under the MIT License.