Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads25d agoHugging Face02m-a-p /FineFineWeb-test FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.tabulartext-classification1M<n<10M5 likes501 downloads2y agoHugging Face03boxin-wbx /test Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.tabulartext-classification10K<n<100K0 likes277 downloads3y agoHugging Face04test-alexpouliquen /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.documenttext-generationn<1K0 likes186 downloads5mo agoHugging Face05summykai /chem-rlvr-TEST ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards Dataset Description ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools. Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.tabularquestion-answering1K<n<10K0 likes182 downloads1y agoHugging Face06JetBrains-Research /agent-trajectories-swe-bench-test-minus-verified Agent Trajectories: SWE-bench Test \ Verified — Mixed Teachers (gpt-5.2 / gpt-5-mini) Summary Full multi-turn agent trajectories collected from the SWE-bench Test minus Verified split (i.e., SWE-bench Test instances that are not part of SWE-bench Verified). Intended for SFT of agent models on coding tasks. Data Collection Each trajectory was produced by a GT-aware lookahead agent that, at every turn: Sampled a candidate response from both gpt-5.2 and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/agent-trajectories-swe-bench-test-minus-verified.tabulartext-generation1K<n<10K0 likes165 downloads7mo agoHugging Face07nlile /math_benchmark_test_saturation LLM Leaderboard Data for Hendrycks MATH Dataset (2022–2024) This dataset aggregates yearly performance (2022–2024) of large language models (LLMs) on the Hendrycks MATH benchmark. It is specifically compiled to explore performance evolution, benchmark saturation, parameter scaling trends, and evaluation metrics of foundation models solving complex math word problems. Original source data: Math Word Problem Solving on MATH (Papers with Code) About Hendrycks' MATH… See the full description on the dataset page: https://huggingface.co/datasets/nlile/math_benchmark_test_saturation.tabularquestion-answeringn<1K0 likes135 downloads2y agoHugging Face08PhillyMac /The_OSHA_Test_Project The OSHA Test Project This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/The_OSHA_Test_Project.tabulartext-generation1K<n<10K0 likes104 downloads18d agoHugging Face09Sudnya /test-subset-classic-eda nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 100 rounds. Every model turn is one row, including the ones that went nowhere. What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.tabulartext-generation1K<n<10K0 likes83 downloads19d agoHugging Face10Duke313 /ClawBench-test ClawBench Can AI Agents Complete Everyday Online Tasks? ClawBench evaluates AI agents on 153 everyday tasks (such as booking flights, ordering groceries, submitting job applications) across 144 live websites. We capture 5 layers of behavioral data (session replay, screenshots, HTTP traffic, agent reasoning traces, and browser actions), collect human ground-truth for every task, and score with an agentic evaluator that provides step-level traceable diagnostics. Paper… See the full description on the dataset page: https://huggingface.co/datasets/Duke313/ClawBench-test.tabulartext-generationn<1K0 likes77 downloads5mo agoHugging Face11JackHsieh /luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tabulartext-generation100K<n<1M0 likes77 downloads1mo agoHugging Face12Aerdax /kagent-humanize-test kagent-humanize-test 2026-09-26 第四版:C18(人设与规划只做输入、范例按场景挑两条、S6 不默认加 emoji)、C19(正式度、话量、口癖频率三项风格参数)、C20(一条气泡等于一个说话动作:S6 只按模型换行分条,不再按逗号切)之后重跑全量。人设仍为小雅。第一版(整改前)结果保留在指标表「kAgent 整改前」列。 200 段中文私聊对话。客户消息取自公开语料原文(与 Aerdax/reply-agent-humanize-test 同一测试集 kapibala_humanize_test),销售消息全部由 kAgent 服务(POST /v1/reply:generate,kAgent commit b026827 加 C18、C19、C20 工作区改动)顺序生成:每一轮请求的 history 是客户原文加此前 kAgent 生成的草稿气泡;kAgent 弃权(abstain)的轮次不产生销售消息,对话里客户消息会连续出现。只保留草稿,不涉及发送。 子集 子集 每行 行数… See the full description on the dataset page: https://huggingface.co/datasets/Aerdax/kagent-humanize-test.tabulartext-generation1K<n<10K0 likes74 downloads14d agoHugging Face13Aerdax /reply-agent-humanize-test reply-agent-humanize-test 200 段中文私聊对话。客户消息取自公开语料原文,销售消息全部由 claude-sonnet-5 按 prompt_zh.md 顺序生成:每一轮生成时模型看到的历史是客户原文加此前生成的回复。 子集 子集 每行 行数 字段 conversations(默认) 一段对话 200 id、messages(role 为 customer 或 sales,content;sales 每条一个气泡) meta 一段对话的元数据 200 id、source、real_customer、domain、scene、lang、n_messages、n_customer_turns、n_sales_bubbles、n_generated_turns、model、prompt_version、product_info generated_replies 一轮生成 1116… See the full description on the dataset page: https://huggingface.co/datasets/Aerdax/reply-agent-humanize-test.tabulartext-generation1K<n<10K0 likes73 downloads17d agoHugging Face14armand0e /hermes-testThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. My Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by nex-agi/nex-n2-pro:free. Sessions: 2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/hermes-test.tabulartext-generationn<1K1 likes69 downloads4mo agoHugging Face15Manju46 /kotlin-test-pairs KotlinTestPairs 9,856 Kotlin source↔unit-test pairs ("focal method" pairs) mined from permissively licensed public GitHub code. Applies the methods2test (MSR 2022) methodology to Kotlin, and is the sibling of SwiftTestPairs. ⚠️ This dataset contains NO source code Rows are references + derived metadata: repository, file paths, content MD5s, and measured properties. The code stays where it has always been, in the upstream corpus. This avoids redistributing anyone's… See the full description on the dataset page: https://huggingface.co/datasets/Manju46/kotlin-test-pairs.tabulartext-generation1K<n<10K0 likes66 downloads24d agoHugging Face16mrinaldi /TestiMolegated Dataset Card for TestiMole -- A multi-billion tokens Italian text corpus Testimole is a large linguistic resource for Italian obtained through a massive web scraping effort. As of June 2024, it is one of the largest datasets for the Italian language, if not the largest, publicly available, consisting of almost 100B tokens counted with the Tiktoken cl100k BPE tokenizer. It consists mainly of conversational data (Italian Usenet hierarchies, Italian message boards, Italian… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/TestiMole.tabulartext-classification100M<n<1B12 likes60 downloads4mo agoHugging Face17room-b007 /test-medicina Medschool-Test, or "Test di Medicina" Is your LLM able to pass a National Entrance Exam for the Italian Medical School? This the GitHub repo for our Hugging Face dataset designed for evaluating Large Language Models (LLMs) on a broad range of questions from the national entrance exams for the Italian medical school (ORIGINAL WEBSITE). The dataset includes multiple-choice questions from various subjects such as biology, chemistry, physics, mathematics, world… See the full description on the dataset page: https://huggingface.co/datasets/room-b007/test-medicina.tabulartext-generation1K<n<10K4 likes55 downloads2y agoHugging Face18JackHsieh /8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/8B-reason-only.stride-32-test.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10K<n<100K0 likes52 downloads25d agoHugging Face19rodriguescarson /adaption-floor-test-300 Legal (300-row probe) Requests (Pooled) Requests from a pooled chat corpus, selected by domain labelling. Rows 300 Domain legal (300-row probe) Format data.parquet, one row per example Licence odc-by Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded. original_completion The target response as uploaded. enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-floor-test-300.tabulartext-generationn<1K0 likes51 downloads14d agoHugging Face20Testing333555 /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.tabulartext-generationn<1K1 likes50 downloads6mo agoHugging Face21test-time-compute /game-of-24 Game of 24 Dataset Dataset Description The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once. This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies. Example Input: 4 5 6 10 Output: (5 * (10 - 4)) - 6 = 24 Step-by-step solution: 10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.tabularquestion-answering1K<n<10K3 likes49 downloads1y agoHugging Face22Minuri /sinhala-test-set-50k Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.tabulartext-generation10K<n<100K0 likes49 downloads6mo agoHugging Face23JackHsieh /4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids 4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. Each record is the model's whole assistant turn: a <think> block, then a note — a long, dense passage of plain prose reasoning about the next 8 tokens after a cut, written from the document prefix alone (the generator never sees the continuation). The chat template is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-think.stride-32-test.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10K<n<100K0 likes48 downloads10d agoHugging Face24JackHsieh /4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids 4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B with thinking turned off. Each thought is a note — a long, dense passage of plain prose reasoning about the next 8 tokens after a cut, written from the document prefix alone (the generator never sees the continuation). The chat template is domain-free (it never mentions arXiv, papers or… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-nothink.stride-32-test.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10K<n<100K0 likes48 downloads10d agoHugging Face25StrataSynth /stratasynth-agent-stress-test StrataSynth Agent Stress Test Part of the StrataSynth Synthetic Identity Engineering corpus. 2,068 turns · 100 conversations · 23 columns per turn High-stakes dialogue explicitly designed to push conversational AI systems. Relationship endings, manager-subordinate conflicts, inheritance disputes. Average relationship tension: 0.69 — the highest of the corpus. This is where Synthetic Identity Engineering is most visible: under maximum pressure, identities either hold or collapse.… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-agent-stress-test.tabulartext-generation1K<n<10K0 likes43 downloads19d agoHugging Face26JackHsieh /4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv 4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by stock Qwen/Qwen3-4B-Instruct-2507 -- no fine-tuning, prompted with the reason-only-nothink.jinja reasoning template. Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each. Why this checkpoint: Stock Qwen3-4B-Instruct-2507 prompted with reason-only-nothink.jinja, a byte-identical copy of the reason-only.jinja template used for… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-reason-only.stride-train8off4-test32.k-8.L-512.statml-arxiv.tabulartext-generation1M<n<10M0 likes42 downloads2mo agoHugging Face27Manju46 /swift-test-pairs SwiftTestPairs 20,801 Swift source↔unit-test pairs ("focal method" pairs) mined from permissively licensed public GitHub code. Applies the methods2test (MSR 2022) methodology to Swift, which had no equivalent dataset. ⚠️ This dataset contains NO source code Rows are references + derived metadata: repository, file paths, content MD5s, and measured properties. Reconstruct file contents from the two public upstream datasets with resolve.py (⚠️ streams ~4 GB from the… See the full description on the dataset page: https://huggingface.co/datasets/Manju46/swift-test-pairs.tabulartext-generation10K<n<100K0 likes42 downloads26d agoHugging Face28JackHsieh /luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags A pre-tokenized, tag-wrapped variant of JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids. Each thought_text is wrapped as <|note|>{thought_text}<|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids). Tag ids and the spliced key are inserted as ids, never re-tokenized. Delimiter ids: <|note|> = 151669, <|/note|> = 151670. These are the first… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tags.tabulartext-generation100K<n<1M0 likes36 downloads1mo agoHugging Face29Maxros /test CS4570 Java Refactoring Benchmark A benchmark of Java refactoring commits mined from GitHub repositories, split into high-resource and low-resource tiers based on repository popularity. Built for the TU Delft CS4570 ML for Software Engineering course (2026). Dataset split Tier Repositories Instances Star range High-resource 57 618 1000-7298 Low-resource ~60 198 13-90 Total 816 Use the resource_tier field to filter by tier. repo_stars and… See the full description on the dataset page: https://huggingface.co/datasets/Maxros/test.tabulartext-generationn<1K0 likes28 downloads4mo agoHugging Face30esadek /test-tracestabulartext-generationn<1K0 likes28 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.