Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /code_x_glue_ct_code_to_text Dataset Card for "code_x_glue_ct_code_to_text" Dataset Summary CodeXGLUE code-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text The dataset we use comes from CodeSearchNet and we filter the dataset as the following: Remove examples that codes cannot be parsed into an abstract syntax tree. Remove examples that #tokens of documents is < 3 or >256 Remove examples that documents contain special tokens (e.g. <img ...> or… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text.texttranslation1M<n<10M79 likes4.3k downloads3y agoHugging Face02codeparrot /xlcost-text-to-code XLCoST is a machine learning benchmark dataset that contains fine-grained parallel data in 7 commonly used programming languages (C++, Java, Python, C#, Javascript, PHP, C), and natural language (English).texttext-generation100K<n<1M51 likes1.4k downloads4y agoHugging Face03google /code_x_glue_tc_text_to_code Dataset Card for "code_x_glue_tc_text_to_code" Dataset Summary CodeXGLUE text-to-code dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Code/text-to-code The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/. Supported Tasks and Leaderboards machine-translation: The dataset can be used to train a model for generating Java code from an English natural… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tc_text_to_code.texttranslation100K<n<1M30 likes784 downloads3y agoHugging Face04google /code_x_glue_cc_code_to_code_trans Dataset Card for "code_x_glue_cc_code_to_code_trans" Dataset Summary CodeXGLUE code-to-code-trans dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans The dataset is collected from several public repos, including Lucene(http://lucene.apache.org/), POI(http://poi.apache.org/), JGit(https://github.com/eclipse/jgit/) and Antlr(https://github.com/antlr/). We collect both the Java and C# versions of the codes and find the parallel… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_to_code_trans.texttranslation10K<n<100K17 likes620 downloads3y agoHugging Face05google /code_x_glue_tt_text_to_text Dataset Card for "code_x_glue_tt_text_to_text" Dataset Summary CodeXGLUE text-to-text dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Text-Text/text-to-text The dataset we use is crawled and filtered from Microsoft Documentation, whose document located at https://github.com/MicrosoftDocs/. Supported Tasks and Leaderboards machine-translation: The dataset can be used to train a model for translating Technical documentation between… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_tt_text_to_text.texttranslation100K<n<1M2 likes474 downloads3y agoHugging Face06lyleokoth /image-to-code-v2imagen<1K2 likes315 downloads2y agoHugging Face07prithivMLmods /OpenSCAD-Code-to-GLB OpenSCAD-Code-to-GLB Dataset Description OpenSCAD-Code-to-GLB is a derived dataset built on top of CodeCAD-HF-v2. It extends the original by converting OpenSCAD parametric code into rendered GLB (GL Transmission Format Binary) 3D model files. Each record pairs a natural language prompt, the corresponding OpenSCAD code, and a rendered GLB file of the resulting 3D mesh. Dataset Statistics Property Value Split Train Number of Rows 4,795… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenSCAD-Code-to-GLB.3dtext-to-3d1K<n<10K1 likes298 downloads7d agoHugging Face08ISLAM-PO /arabic-to-code-8-langs-3m Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,239,045 Dataset Server estimate: 1,995,159 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.texttext-generation1M<n<10M0 likes283 downloads13d agoHugging Face09feedback-to-code /SWE-bench__style-3__fs-oracle Dataset Card for "SWE-bench__style-3__fs-oracle" More Information needed textn<1K0 likes235 downloads3y agoHugging Face10codeparrot /github-jupyter-code-to-text Dataset description This dataset consists of sequences of Python code followed by a a docstring explaining its function. It was constructed by concatenating code and text pairs from this dataset that were originally code and markdown cells in Jupyter Notebooks. The content of each example the following: [CODE] """ Explanation: [TEXT] End of explanation """ [CODE] """ Explanation: [TEXT] End of explanation """ ... How to use it from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/github-jupyter-code-to-text.texttext-generation10K<n<100K27 likes233 downloads3y agoHugging Face11feedback-to-code /SWE-bench__style-3__fs-oracle_large_tokenlength Dataset Card for "SWE-bench__style-3__fs-oracle" More Information needed textn<1K0 likes221 downloads3y agoHugging Face12anisiraj /code-image-to-text Code Snippet Image → Text A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of transcribing an image of a code snippet back into its source text — syntax-aware OCR. Each example pairs a syntax-highlighted PNG of code with the exact code text that produced it. It spans 8 programming languages and deliberately mixes two capture types: block — a complete function / unit (6–45 lines). fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.imageimage-to-text10K<n<100K0 likes169 downloads3mo agoHugging Face13michaelnath /bad_code_to_good_code_dataset Dataset Card for "bad_code_to_good_code_dataset" More Information needed text1M<n<10M0 likes105 downloads4y agoHugging Face14sinatra-rd /math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages ​​in a similar way to the code interpreter. Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog Jsonl format: {"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.textn<1K0 likes62 downloads1y agoHugging Face15GPS-Lab /Compliance-to-Codetextn<1K1 likes52 downloads1y agoHugging Face16knguyennguyen /code_x_glue_ct_code_to_text_java_pythontext100K<n<1M0 likes47 downloads2y agoHugging Face17feedback-to-code /mozzarella Mozzarella-0.3.1 Motivation Mozzarella is a dataset matching issues (= problem statements) and corresponding pull requests (PRs = problem solutions) of a selection of well maintained Java GitHub repositories. The original purpose was to serve as training and evaluation data for ML models concerned with fault localization and automated program repair of complex code bases. However, there might be more use cases that could benefit from this data. Inspired by SWEBench… See the full description on the dataset page: https://huggingface.co/datasets/feedback-to-code/mozzarella.text1K<n<10K1 likes46 downloads2y agoHugging Face18OrangeeSofty /Metro_Code_Chapters_18_to_22_Data Dataset Card for Metro System Requirements in Design Summary This dataset provides detailed requirements for systems used in metro network design, collected from Chapter 18-22 of the Code for Design of Metro (GB 50157-2013). The dataset is annotated using a description, categories format, aimed at facilitating the training and fine-tuning of large language models (LLMs) for information extraction tasks in complex product systems, particularly within metro transit… See the full description on the dataset page: https://huggingface.co/datasets/OrangeeSofty/Metro_Code_Chapters_18_to_22_Data.texttext-classification10K<n<100K0 likes43 downloads2y agoHugging Face19metaphilabs /frontend-figma-to-code CREW: Figma to Code A benchmark for evaluating AI coding agents on Figma-to-code generation — converting real-world Figma community designs into production-ready React + Tailwind CSS applications. Each task gives an agent full access to a Figma file via MCP tools. The agent must extract the design system, generate components, build successfully, and deploy a live preview. Outputs are evaluated through human preference (ELO) and automated verifiers. Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/metaphilabs/frontend-figma-to-code.textn<1K0 likes39 downloads7mo agoHugging Face20SultanR /arxiv-to-code-agentic-tool-calling arxiv-to-code-agentic-tool-calling Multi-turn tool-calling dataset where an assistant implements ML papers in PyTorch through file-creation and command-execution tool calls. Built from lucidrains' (Phil Wang) open-source paper implementations. There are ~217 repositories on Codeberg, each implementing a different ML paper. This dataset reverse-engineers those into synthetic coding conversations. What's in it 199 conversations, each covering one repository. Every… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/arxiv-to-code-agentic-tool-calling.tabulartext-generationn<1K0 likes39 downloads8mo agoHugging Face21Reubencf /adaption-react-screenshot-to-code This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-react_screenshot_to_code This dataset contains 1,000 paired examples for training multimodal screenshot-to-code systems, specifically targeting React and TypeScript implementations. Each entry consists of a source webpage screenshot, a detailed visual description, and the corresponding generated React/TSX source code. The samples demonstrate high-fidelity UI… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-react-screenshot-to-code.textn<1K0 likes29 downloads3mo agoHugging Face22obaydata /v2c-video-to-code-demo V2C: Video to Code/Game Demo Summary Video-to-Code (V2C) demo dataset containing game video recordings paired with AI-generated game code. Each sample includes the source video, chain-of-thought reasoning, game requirements analysis, and the final generated HTML game code. Games Included Game Video Duration Code Output Description 2048 ~175s (720x1440) index.html Number puzzle game recreation Flappy Bird variant ~91s (582x1280) generated_game.html… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/v2c-video-to-code-demo.texttext-generationn<1K0 likes26 downloads6mo agoHugging Face23Adam-Ben-Khalifa /strudel-nl-to-codetextn<1K0 likes25 downloads6mo agoHugging Face24archit11 /hyperswitch-issue-to-code_v2 Rust Commit Dataset - Hyperswitch Dataset Description This dataset contains Rust commit messages paired with their corresponding code patches from the Hyperswitch repository. Dataset Summary Total Examples: 319 Language: Rust Source: Hyperswitch GitHub repository Format: Prompt-response pairs for supervised fine-tuning (SFT) Data Fields prompt: The commit message describing the change response: The git patch/diff showing the actual code changes… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-issue-to-code_v2.texttext-generationn<1K0 likes24 downloads1y agoHugging Face25omgbobbyg /Zip-Code-to-Timezone Dataset Card for Zip Code to Timezone Offset Mapping This dataset maps Zip Codes and Postal Codes for the USA and Canada to the relevant timezone offset. Dataset Details Dataset Description In addition to providing a mapping from a Zip Code or Postal Code to timezone offset, it also contains the timezone offset for DST (if observed). Curated by: Bobby Gill, BlueLabel Acknowledgements Based off the Work Here:… See the full description on the dataset page: https://huggingface.co/datasets/omgbobbyg/Zip-Code-to-Timezone.tabular10K<n<100K0 likes22 downloads3y agoHugging Face26doejn771 /code_x_glue_ct_code_to_text_java_pythontext100K<n<1M0 likes22 downloads2y agoHugging Face27dk-bot /code_x_glue_ct_code_to_texttext1K<n<10K0 likes20 downloads2y agoHugging Face28feedback-to-code /cwa-server-task-instancestextn<1K0 likes16 downloads3y agoHugging Face29TachyHealth /ADA_Dental_Code_to_SBS_V2tabularn<1K2 likes16 downloads3y agoHugging Face30Romoamigo /oop-bad-code-to-good-code-cpptext10K<n<100K0 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.