Team Ai
Datasetpublic

Corpus-NZ/Code-Syntax-Expanded

Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows. 📊 Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes24downloads
1 commits on main
4381d3b1mo ago

Duplicate from Gugu8/Code-Syntax-Expanded

Gugu8