Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01klieret /swe-bench-dummy-test-datasettextn<1K0 likes50k downloads1y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes46k downloads2y agoHugging Face03albertvillanova /datasets-tests-compressiontextn<1K0 likes43k downloads5y agoHugging Face04agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K208 likes39k downloads2y agoHugging Face05databricks /databricks-dolly-15k Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.textquestion-answering10K<n<100K1.3k likes36k downloads3y agoHugging Face06OraRL /OraRL-Data OraRL-Data [🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code] We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL. It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.imagevideo-text-to-text10K<n<100K1 likes24k downloads2mo agoHugging Face07open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes24k downloads6h agoHugging Face08common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes21k downloads1y agoHugging Face09efficient-deep-research /synthesized_datasettext10K<n<100K0 likes19k downloads1y agoHugging Face10MedOtter /brats2023-gli-dataset BraTS2023 GLI Dataset Dataset Description The BraTS2023 Glioma (GLI) dataset for brain tumor segmentation. This dataset contains multi-modal MRI scans with dense segmentation annotations. Multi-Modal MRI Each patient case includes 4 MRI modalities: T1n: Native T1-weighted MRI T1c: Post-contrast T1-weighted MRI T2w: T2-weighted MRI T2f: T2-FLAIR MRI All 4 modalities share the same segmentation mask. Dataset Structure Each sample in… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/brats2023-gli-dataset.textimage-segmentation1K<n<10K4 likes19k downloads1y agoHugging Face11RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M95 likes19k downloads1y agoHugging Face12Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M216 likes18k downloads1mo agoHugging Face13KAS2003 /xfield-radar-dataset-20260915 XField radar dataset — formal snapshot, 2026-09-15 Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json. This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.text10K<n<100K0 likes16k downloads25d agoHugging Face14zouhar /bio-mqm-datasetThis dataset is compiled from the official Amazon repository (all respective licensing applies). It contains system translations, multiple references, and their quality evaluation on the MQM scale. It accompanies the ACL 2024 paper Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains. Watch a brief 4 minutes-long video. Abstract: We introduce a new, extensive multidimensional quality metrics (MQM) annotated dataset covering 11 language pairs in the biomedical domain. We use this… See the full description on the dataset page: https://huggingface.co/datasets/zouhar/bio-mqm-dataset.texttranslation10K<n<100K8 likes16k downloads2y agoHugging Face15nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K111 likes13k downloads1y agoHugging Face16tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B19 likes12k downloads6mo agoHugging Face17nvidia /Nemotron-Cascade-2-SFT-Data Nemotron-Cascade-2-SFT-Data We release the SFT data used for training Nemotron-Cascade-2. Data sources Math Our non-proof math prompts are sourced from Nemotron-Cascade-1-SFT and Nemotron-Math-v2, with responses generated by DeepSeek-V3.2, DeepSeek-V3.2-Speciale, and GPT-OSS-120B. For mathematical proofs, prompts are taken from Nemotron-Math-Proofs-v1 and generated using DeepSeek-V3.2-Speciale. Science We collect science prompts from… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data.text10M<n<100M75 likes9.4k downloads7mo agoHugging Face18xxxspatialencoderwds4 /data_4 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.tabularobject-detectionn<1K1 likes8.7k downloads20d agoHugging Face19soma114 /soma-competition-datasettabularn<1K0 likes8.5k downloads11d agoHugging Face20xxxspatialencoderwds3 /data_3 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.tabularobject-detectionn<1K0 likes6.9k downloads20d agoHugging Face21RooseveltHonaker /kalshi-trades-data Kalshi Public Trade History An independent archive of public Kalshi market trade data retrieved from Kalshi market-data API endpoints. This dataset is not affiliated with or endorsed by Kalshi. Each historical_raw_N.jsonl file is an immutable numbered shard. Records contain a market ticker, market window, reported volume, and the public trades returned for that market. Trade fields can include timestamps, public trade IDs, prices, quantities, taker direction, and block-trade… See the full description on the dataset page: https://huggingface.co/datasets/RooseveltHonaker/kalshi-trades-data.text100K<n<1M1 likes6.8k downloads21m agoHugging Face22nvidia /Nemotron-VLM-Dataset-v2 Nemotron-VLM-Dataset v2 Versions Date Commit Changes 2025-11-05 head Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes. 2025-10-28 214051e Initial Release Dataset Description Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples. This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.textvisual-question-answering1M<n<10M98 likes6.8k downloads10mo agoHugging Face23LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B90 likes6.5k downloads6mo agoHugging Face24PortPy-Project /PortPy_Dataset PortPy: Planning and Optimization for Radiation Therapy Data Overview PortPy equips researchers with a robust benchmark patient dataset, sourced from the FDA-approved Eclipse commercial treatment planning system through its API. This dataset embodies all necessary elements for optimizing various machine configurations such as beam angles, aperture shapes, and leaf movements. It includes Dose Influence Matrix (AKA dose deposition matrix, dij matrix): The dose… See the full description on the dataset page: https://huggingface.co/datasets/PortPy-Project/PortPy_Dataset.textn<1K1 likes6.3k downloads5mo agoHugging Face25xxxspatialencoderwds2 /data_2 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.tabularobject-detectionn<1K5 likes6.2k downloads20d agoHugging Face26xxxspatialencoderwds1 /data_1 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds1/data_1.tabularobject-detectionn<1K0 likes6k downloads20d agoHugging Face27nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M708 likes6k downloads1y agoHugging Face28Gramscii-IT /european-open-data-catalogue Open Data catalogue This repository publishes independently versioned metadata and licensed source snapshots: A discovery catalogue with 69000 dataset entries, covering the providers listed in the discovery table below. 3 independently pinned availability indexes with 911,795 joint combinations across 35 datasets, built from complete source responses within the explicitly declared scope. Licensed Cruscotto source snapshots, stored separately from the metadata, preserve the… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.text100K<n<1M1 likes5.8k downloads16h agoHugging Face29lfsm /ja-datasettext100K<n<1M0 likes5.5k downloads3y agoHugging Face30lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes5.4k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.