datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-to-code-8-langs-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,239,045
Dataset Server estimate: 1,995,159
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages in a similar way to the code interpreter.
Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog
Jsonl format:
{"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.Compliance-to-Codeadaption-react-screenshot-to-code
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-react_screenshot_to_code
This dataset contains 1,000 paired examples for training multimodal screenshot-to-code systems, specifically targeting React and TypeScript implementations. Each entry consists of a source webpage screenshot, a detailed visual description, and the corresponding generated React/TSX source code. The samples demonstrate high-fidelity UI… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-react-screenshot-to-code.cwa-server-task-instancesConv-to-Bench-Code
Conv-to-Bench: Evaluating LLMs via User-Assistant Dialogues
This repository contains the code-domain dataset generated by the Conv-to-Bench framework, presented at the 3rd Workshop on Navigating and Addressing Data Problems for Foundation Models (DATA-FM @ ICLR 2026). The framework automatically transforms authentic multi-turn dialogues between users and assistants into structured, verifiable requirement checklists for LLM evaluation.
Overview
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/huglabs/Conv-to-Bench-Code.Server_Text_Dataset_1disease-code-to-name-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
disease_code_to_name
This dataset consists of pairs mapping unique disease identification codes (e.g., DIS000940) to their corresponding medical condition names. The samples cover a range of chronic diseases including Alzheimer's, Parkinson's, Breast Cancer, Hypertension, and Diabetes. It serves as a lookup resource for translating standardized disease identifiers into human-readable labels.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/disease-code-to-name-v1.disease-code-to-name
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
disease_code_to_name
This dataset consists of pairs mapping unique disease identification codes (e.g., DIS000940) to their corresponding medical condition names. The samples cover a range of chronic diseases including Alzheimer's, Parkinson's, Breast Cancer, Hypertension, and Diabetes. It serves as a lookup resource for translating standardized disease identifiers into human-readable labels.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/disease-code-to-name.retrieval_bench_1python-text-to-codecode-end-to-endcommitpack_original_cm_old_code_to_diff
