Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-alchemy /code-alchemy CodeAlchemy CodeAlchemy is a synthetic code dataset (~976.6B tokens, ~162M rows) designed for training and evaluating code language models. It consists of 5 training subsets covering a range of code-related tasks, and 2 evaluation subsets. All files are Parquet with zstd compression with on-disk size ~873 GB. Raw source files are not included due to ownership considerations and must be manually fetched as instructed below. Dataset Statistics Config… See the full description on the dataset page: https://huggingface.co/datasets/open-alchemy/code-alchemy.tabulartext-generation100M<n<1B15 likes2.2k downloads3mo agoHugging Face02choucsan /mimo-claude-code-traces-1k MIMO Claude Code Traces MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools. The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.tabulartext-generation1K<n<10K11 likes2k downloads2mo agoHugging Face03louisbrulenaudet /code-sante-publique Code de la santé publique, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sante-publique.tabulartext-generation1K<n<10K1 likes298 downloads1y agoHugging Face04CSJianYang /CodeArena Dataset Summary To bridge the gap between the model-generated response and human preference, we present a rigorous human-curated benchmark CodeArena to emulate the complexity and diversity of real-world coding tasks, where 397 high-quality samples spanning 40 categories and 40 languages, carefully curated from user queries. Data Example An example of 'validation' looks as follows: { "id": "60670a8d9b1e39dd845fb1639d0d8b86", "messages": "[{'role': 'user'… See the full description on the dataset page: https://huggingface.co/datasets/CSJianYang/CodeArena.tabularquestion-answeringn<1K16 likes273 downloads2y agoHugging Face05louisbrulenaudet /code-commande-publique Code de la commande publique, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-commande-publique.tabulartext-generation1K<n<10K0 likes252 downloads1y agoHugging Face06Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes187 downloads6mo agoHugging Face07Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes163 downloads5mo agoHugging Face08sup2ch /ru-en-code-curriculum RuEn Code Curriculum RuEn Code Curriculum is a curated Russian-English dataset for continued pretraining (CPT) and supervised fine-tuning (SFT) of small code-oriented language models. This public release contains only records classified as redistributable. Local-training-only web and code sources used by the internal curriculum are intentionally excluded. Dataset summary Configuration Split Records Tokens sft train 53,278 13,997,239 sft reserve 19,225… See the full description on the dataset page: https://huggingface.co/datasets/sup2ch/ru-en-code-curriculum.tabulartext-generation100K<n<1M1 likes161 downloads17d agoHugging Face09louisbrulenaudet /code-civil Code civil, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-civil.tabulartext-generation1K<n<10K3 likes144 downloads1y agoHugging Face10louisbrulenaudet /code-impots Code général des impôts, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impots.tabulartext-generation1K<n<10K5 likes139 downloads1y agoHugging Face11Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes131 downloads10mo agoHugging Face12tandevllc /offsec_redteam_codesgated OffSec RedTeam Codes Token count: ~30B tokens. OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled. ⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.tabulartext-generation1M<n<10M17 likes121 downloads11mo agoHugging Face13louisbrulenaudet /code-collectivites-territoriales Code général des collectivités territoriales, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-collectivites-territoriales.tabulartext-generation1K<n<10K0 likes120 downloads1y agoHugging Face14adityabhushannagar /code-alchemy-rust CodeAlchemy Rust Rust-only derivative of open-alchemy/code-alchemy. It preserves the five training configs, two evaluation configs, original splits, row order, columns, values, and task/evaluation fields. Rows were selected from the source-native language labels: Rust and rust in training data and dev-eval rs in trace-eval Labels remain unchanged in the output. code-trace.external_packages is normalized to list<string> because source Parquet shards physically alternate between… See the full description on the dataset page: https://huggingface.co/datasets/adityabhushannagar/code-alchemy-rust.tabulartext-generation1M<n<10M0 likes105 downloads2mo agoHugging Face15louisbrulenaudet /code-consommation Code de la consommation, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-consommation.tabulartext-generation1K<n<10K0 likes99 downloads1y agoHugging Face16louisbrulenaudet /code-construction-habitation Code de la construction et de l'habitation, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-construction-habitation.tabulartext-generation1K<n<10K1 likes98 downloads1y agoHugging Face17louisbrulenaudet /code-route Code de la route, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-route.tabulartext-generation1K<n<10K0 likes89 downloads1y agoHugging Face18louisbrulenaudet /code-pensions-civiles-militaires-retraite Code des pensions civiles et militaires de retraite, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-civiles-militaires-retraite.tabulartext-generationn<1K0 likes84 downloads2y agoHugging Face19typedef-ai /fenic-codebasetabulartext-generation10K<n<100K0 likes83 downloads2mo agoHugging Face20louisbrulenaudet /code-environnement Code de l'environnement, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-environnement.tabulartext-generation1K<n<10K0 likes80 downloads1y agoHugging Face21ChrisScr770 /tennessee-code Tennessee Code Unannotated (Titles 29, 34, 36, 37, 39, 40 and 41) Statute text from the Tennessee Code Unannotated, Titles 29, 34, 36, 37, 39, 40 and 41. The text is split into 9,050 retrieval-sized chunks. Each chunk keeps its place in the section, its subsection labels, its version qualifiers and the cross-references in its text, resolved to other chunks where possible. The chunk text is Markdown, so the cite, the heading and every subsection label are part of the text itself.… See the full description on the dataset page: https://huggingface.co/datasets/ChrisScr770/tennessee-code.tabulartext-retrieval10K<n<100K1 likes78 downloads2d agoHugging Face22louisbrulenaudet /code-propriete-intellectuelle Code de la propriété intellectuelle, non-instruct (2025-03-10) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-propriete-intellectuelle.tabulartext-generation1K<n<10K0 likes69 downloads2y agoHugging Face23LLMTeamAkiyama /cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder データ件数: 269,863 平均トークン数: 11674 最大トークン数: 31,184 合計トークン数: 3,150,447,484 ファイル形式: JSONL ファイルサイズ: 不明 加工内容 synthetic_sftを使用 トークン処理が重たいので、文字数でフィルター seed_question < 6000 generation < 80000 thinkタグ除去 が中途半端なものを除外 トークナイズ処理(速度向上アップデート 繰り返し除去 tabularquestion-answering100K<n<1M0 likes66 downloads1y agoHugging Face24louisbrulenaudet /code-impots-annexe-i Code général des impôts, annexe I, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impots-annexe-i.tabulartext-generationn<1K0 likes65 downloads1y agoHugging Face25jasonlingg /envoy-qasper-code-trajectories Envoy QASPER Code-Execution Trajectory Pilot This is a small, fully disclosed pilot of executable research-agent trajectories. Claude Sonnet 5 generated Python actions against a persistent document REPL. The Envoy pipeline executed every action and retained the real observations. An AI coding assistant then reviewed answer support, stopping behavior, and replay. This release is useful for studying trajectory validation and citation failures. It is not a production-ready SFT… See the full description on the dataset page: https://huggingface.co/datasets/jasonlingg/envoy-qasper-code-trajectories.tabularquestion-answeringn<1K0 likes65 downloads21d agoHugging Face26fai-adh /fon-code-switching-evaluation French-Fon Code-Switching Evaluation Benchmark Overview This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios. The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon. The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.tabularquestion-answeringn<1K0 likes62 downloads2d agoHugging Face27louisbrulenaudet /code-domaine-etat-collectivites-mayotte Code du domaine de l'Etat et des collectivités publiques applicable à la collectivité territoriale de Mayotte, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-etat-collectivites-mayotte.tabulartext-generationn<1K0 likes61 downloads1y agoHugging Face28louisbrulenaudet /code-impots-annexe-iv Code général des impôts, annexe IV, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impots-annexe-iv.tabulartext-generationn<1K0 likes58 downloads1y agoHugging Face29louisbrulenaudet /code-impositions-biens-services Code des impositions sur les biens et services, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-impositions-biens-services.tabulartext-generation1K<n<10K0 likes56 downloads1y agoHugging Face30louisbrulenaudet /code-justice-militaire-nouveau Code de justice militaire (nouveau), non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-justice-militaire-nouveau.tabulartext-generationn<1K0 likes55 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.