Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M14 likes24k downloads3h agoHugging Face02chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes21k downloads5mo agoHugging Face03AyoubChLin /Company-document-dataset-v2 Company Documents v2 Generation complete: all 13 document types have completed export and upload checkpoints. Synthetic, born-digital business documents rendered from four open sample databases, with exact gold labels: 353,580 PDFs (404,514 pages) of 13 document types in English and French, issued by 60 synthetic companies, each with its own letterhead, numbering and wording. Successor of CompanyDocuments (2,677 PDFs, 4 types). Dataset overview property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentdocument-question-answering100K<n<1M2 likes12k downloads3d agoHugging Face04takschdube /moltbook-dataset Moltbook Dataset A longitudinal dataset of social interactions from Moltbook — an AI-agent social platform where autonomous "Molties" post, comment, and interact. Collected automatically and published as timestamped snapshots for temporal analysis. Dataset Statistics Metric Count Posts (platform total) 4,427,314 Comments (platform total) 13,538,160 Posts (collected) 422,497 Comments (collected) 3,723,066 Agents 57,217 Social graph edges 826… See the full description on the dataset page: https://huggingface.co/datasets/takschdube/moltbook-dataset.tabulartext-generation1M<n<10M4 likes6.5k downloads2h agoHugging Face05SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face06johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M4 likes5k downloads8mo agoHugging Face07GenerTeam /pretrain_data_eukaryote GENERator-v2-Eukaryote Gene-Centric Pretraining Corpus This repository provides the gene-centric pretraining corpus underlying GENERator-v2-Eukaryote, a large-scale DNA language model for eukaryotic genome understanding. The dataset is constructed by leveraging RefSeq annotations to extract biologically meaningful functional genomic regions, which serve as the foundation for large-context DNA language model pretraining. 📌 Dataset Construction Overview The core… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/pretrain_data_eukaryote.tabulartext-generationn<1K4 likes2.6k downloads5mo agoHugging Face08xupy21 /ICPC_Data ICPC World Finals — a discriminative subset, with model traces 24 ICPC World Finals problems (2021–2025), together with 1440 full contest transcripts of an LLM attempting them under simulated contest rules across three arms: with no hint, with the official editorial as a hint, and with a hint written by a second model that gets 10 rounds of measured feedback to improve it. Selection The agent Every contest run in this dataset comes from:… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.tabulartext-generationn<1K1 likes2.5k downloads9d agoHugging Face09mjaso /flashmini-data-v1 FlashMini data v4 (card) Deterministic FlashMini training corpus. Canonical documents live in Parquet+ZSTD shards under shards/; each shard carries a manifest with sha256, counts, and distributions; the frozen corpus identity is corpus_fingerprint_sha256. Sources and redistribution: each source carries one of mirror_allowed, recipe_only, gated_recipe_only, review_required, generated_owned (fail-closed; see registry/sources.yaml + source_snapshot.lock.json). Content shards are… See the full description on the dataset page: https://huggingface.co/datasets/mjaso/flashmini-data-v1.tabulartext-generation1M<n<10M0 likes2.4k downloads17d agoHugging Face10Cyrile /dataset-the-stack-v2-dedup-sub The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++) TempestTeam/dataset-the-stack-v2-dedup-sub Dataset Summary This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part of the BigCode Project. It contains only files written in the following programming languages: Python 🐍 Java ☕ JavaScript 📜 C ⚙️ C++ ⚙️ Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.tabulartext-generation10M<n<100M6 likes2.3k downloads2y agoHugging Face11data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2.2k downloads3y agoHugging Face12M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.1k downloads3y agoHugging Face13OpenSafetyLab /Salad-Data Data Description ✊ How to use from datasets import load_dataset dataset = load_dataset("OpenSafetyLab/Salad-Data", name='base_set', split='train') 📊 Statistical Overview of Base Question Type Data Source Nums Self-instructed Finetuned GPT-3.5 15,433 Open-Sourced HH-harmless 4,184 HH-red-team 659 Advbench 359 Multilingual 230 Do-Not-Answer 189 ToxicChat 129 Do Anything Now 93 GPTFuzzer 42 Total 21,318… See the full description on the dataset page: https://huggingface.co/datasets/OpenSafetyLab/Salad-Data.tabulartext-classification10K<n<100K34 likes2.1k downloads3mo agoHugging Face14bakrianoo /jabarti-llm-dataset jabarti-llm-dataset Cleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Every pretrain record is a contiguous span of 120-1500 characters with the article title and section headings removed. Provenance is in ds_source. Configs and Splits Config Split Rows Training phase Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.tabulartext-generation1M<n<10M58 likes1.4k downloads17d agoHugging Face15RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes1.3k downloads19d agoHugging Face16BytedTsinghua-SIA /Open-MOPD-Data Open-MOPD Data This repository contains the training and evaluation data released with Open-MOPD, including mixed-domain supervised fine-tuning data, the shared RL/OPD prompt mixture, and six evaluation benchmarks. Dataset contents Configuration Description Examples rl_prompt_mix Shared math, code, and instruction-following prompts for RL and OPD 86,931 sft_openr1_math_93k Math SFT data in a unified think-tag format 93,733 sft_ocr_50k Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.tabulartext-generation1M<n<10M2 likes1.3k downloads2mo agoHugging Face17inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face18BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face19ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face20DataHound26 /FineWeb-2021 FineWeb-Edu 2021 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2021. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2021 Rows 139,636,993… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2021.tabulartext-generation100M<n<1B0 likes1k downloads23d agoHugging Face21DataHound26 /FineWeb-2020 FineWeb-Edu 2020 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2020. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2020 Rows 123,382,457… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2020.tabulartext-generation100M<n<1B0 likes1k downloads23d agoHugging Face22Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes1k downloads2y agoHugging Face23d3LLM /trajectory_data_dream_32 d3LLM Trajectory Dataset Project Page | Paper | GitHub | Blog This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation". Introduction d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.tabulartext-generation100K<n<1M0 likes947 downloads5mo agoHugging Face24quantcodeeval /task_data QuantCodeEval A benchmark for evaluating LLM coding agents on quantitative-strategy code reproduction from finance research papers. Status: Anonymous artifact for the 30-task benchmark. Release mirrors The release is mirrored at two anonymous locations: Hugging Face Datasets — complete anonymous release: https://huggingface.co/datasets/quantcodeeval/task_data anonymous.4open.science — browseable mirror: https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.tabulartext-generationn<1K2 likes906 downloads2mo agoHugging Face25trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes901 downloads7mo agoHugging Face26DataHound26 /FineWeb-2022 FineWeb-Edu 2022 — Cleaned and Shuffled This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2022. This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year. Dataset summary Item Value Year 2022 Rows 106,753,442… See the full description on the dataset page: https://huggingface.co/datasets/DataHound26/FineWeb-2022.tabulartext-generation100M<n<1B0 likes892 downloads23d agoHugging Face27UCLNLP /monoweb-dataset MonoWeb Dataset MonoWeb is a multilingual pretraining corpus derived from FineWeb-Edu (English) and FineWeb2 (German, Spanish, French) by systematically removing all mixed-language documents. Released alongside the paper: The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining Dataset Structure monoweb-dataset/ ├── eng/ # Full English corpus (FineWeb-Edu) ├── deu/… See the full description on the dataset page: https://huggingface.co/datasets/UCLNLP/monoweb-dataset.tabulartext-generation100M<n<1B0 likes864 downloads6mo agoHugging Face28longevity-genie /cell2sentence4longevity-data Dataset Card: longevity-genie/cell2sentence4longevity-data Summary This repository contains preprocessed single-cell RNA-seq (scRNA‑seq) datasets prepared as “cell sentences” for training and evaluation of cells2sentence-style models. Each cell is represented as a space‑separated sequence of top expressed gene symbols, enabling language‑model style training for tasks such as biological age prediction and other downstream applications. This dataset targets fine‑tuning and… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/cell2sentence4longevity-data.tabulartext-generation10M<n<100M0 likes782 downloads11mo agoHugging Face29anonymous-md /EDGAR_FILINGS_DATASET SFD: SEC Filings Dataset (v1) SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser described in: The SEC Filings Dataset: Reconstructing U.S. Corporate and Financial… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-md/EDGAR_FILINGS_DATASET.tabulartext-generation1M<n<10M3 likes772 downloads5mo agoHugging Face30beita6969 /R2Flow-Dataset R² Flow Dataset Training and test splits of R² Flow: Recursive Self-Improvement via Recursive Skill Evolution (code: beita6969/r2flow). Six in-distribution (IID) benchmarks supply the training tasks and the IID test sets. Six out-of-distribution (OOD) benchmarks are evaluation-only; each is posed as one IID task type. Benchmark Role Training Test Test source Posed as HotpotQA IID 512 128 distractor validation – TriviaQA IID 512 128 rc.nocontext validation – AIME… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/R2Flow-Dataset.tabularquestion-answering1K<n<10K0 likes735 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.