Corpus-NZ/Code-Syntax-Expanded
Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.
025
1version https://git-lfs.github.com/spec/v12oid sha256:fab5de517a71101b96adfc5ceaa9d0579f2465bdd2362e7841569b441d089b403size 10739009744 