Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Samsoup /Code-Generation-Quality-Estimation Code Generation Quality Estimation This repository contains model-ready task context, generated code, and complete-case execution-resource targets for five public LLM code-generation cohorts. It provides deterministic 70/10/20 group-aware split versions using seeds 42, 1234, and 2026. Configurations There are 15 configurations: one for each dataset and split seed. Each configuration has train, validation, and test splits. Dataset Complete rows Groups Models… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/Code-Generation-Quality-Estimation.tabulartabular-regression100K<n<1M3 likes114 downloads3mo agoHugging Face02sujalgawas /multilang-code-quality-reviewstext10K<n<100K0 likes79 downloads2mo agoHugging Face03sujalgawas /multilang-code-quality-100ktext10K<n<100K0 likes61 downloads2mo agoHugging Face04sujalgawas /big-multilang-code-quality-reviewstext10K<n<100K0 likes32 downloads2mo agoHugging Face05CatQualia /code-quality-corpusgated CatQualia code-quality corpus — semantic smell classes with before/after fixes 39,383 rows · 25,541,792 bytes · JSON Lines, one object per line. What this is Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.tabular10K<n<100K0 likes31 downloads26d agoHugging Face06AadiBhatia /code-edit-quality Code Editing Quality — SFT-Ready (ShareGPT Format) Quality-filtered splits of a 50K code-editing SFT dataset in ShareGPT conversation format, produced by LLM-based distillation that evaluates 9 quality criteria per sample. Format Each sample has a conversations field with ShareGPT-style turns: system: Code editing system prompt human: Instruction + source code gpt: Edited code Compatible with axolotl, LLaMA-Factory, and other SFT frameworks that support ShareGPT format.… See the full description on the dataset page: https://huggingface.co/datasets/AadiBhatia/code-edit-quality.texttext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face07v-i-s-h-w-a-s /code-quality-assistant-comprehensivetext1K<n<10K0 likes6 downloads1y agoHugging Face08happylife365 /code-quality-poor Low Quality Code Dataset Overview This dataset contains 444 Python code samples with poor quality scores (62-74 out of 100). These samples can be used for: Training models to recognize bad code patterns Contrastive learning (good vs bad code) Code quality classification tasks Statistics Metric Value Total samples 444 Quality range 62-74 Average quality ~70 Source Distribution Source Count BigOBench 425… See the full description on the dataset page: https://huggingface.co/datasets/happylife365/code-quality-poor.tabularn<1K0 likes4 downloads8mo agoHugging Face09LGizkde /code_qualitytabularn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.