Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes34k downloads2y agoHugging Face02agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.1k downloads2y agoHugging Face03EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.3k downloads2y agoHugging Face04ksolovev /fine-news-sample Fine-News Sample Fine-News Sample contains 1,000,000 rows sampled from the Fine-News corpus. The sample covers all 117 capture months and 388 language-and-script labels in that corpus. Each selected row preserves its article text, source metadata, and sampling weight. The full corpus derives from INFINI-NEWS, which extracts article text from Common Crawl News web archives. At a glance Measure Value Rows 1,000,000 Distinct document IDs 1,000,000 Sum… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news-sample.texttext-generation1M<n<10M0 likes2.7k downloads2d agoHugging Face05asenion-ai /sampled-local-resumes sampled-local-resumes This dataset contains synthetic resume data sampled from local folders (20% sample from each folder). License This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details. Attribution Copyright 2025 Fairly AI Inc. dba Asenion This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0. You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.text-generation1K<n<10K0 likes1.3k downloads1y agoHugging Face06bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes1.3k downloads4y agoHugging Face07TheFinAI /dolma3_300B_samplegated Dolma 3 — 300B-token sample 🌐 The Fin AI Pretraining / reference corpus released by The Fin AI. Source: Dolma 3 mix (AllenAI) — https://huggingface.co/allenai. Source A ~300B-token sample of AllenAI's Dolma 3 mix; Dolma is released under ODC-BY 1.0. Structure Rows: 187,823,645 Columns: source, date, text, token_count, category Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/dolma3_300B_sample"… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/dolma3_300B_sample.tabulartext-generation100M<n<1B0 likes1.1k downloads3d agoHugging Face08DynaMath /DynaMath_Sample Dataset Card for DynaMath [💻 Github] [🌐 Homepage][📖 Preprint Paper] Dataset Details 🔈 Notice DynaMath is a dynamic benchmark with 501 seed question generators. This dataset is only a sample of 10 variants generated by DynaMath. We encourage you to use the dataset generator on our github site to generate random datasets to test. 🌟 About DynaMath The rapid advancements in Vision-Language Models (VLMs) have shown significant potential in tackling… See the full description on the dataset page: https://huggingface.co/datasets/DynaMath/DynaMath_Sample.imagemultiple-choice1K<n<10K9 likes726 downloads2y agoHugging Face09FredyRivera-dev /LLaDA-Sample-10BT Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.text-generation1B<n<10B2 likes570 downloads4mo agoHugging Face10JamesConley /fineweb-sample-22.95B-512 FineWeb-Sample-22.95B-512 Dataset Description This dataset contains approximately 22.95 billion tokens (22,948,244,480 tokens) sampled from the FineWeb-Edu dataset. Each text sample is capped at a maximum of 512 tokens. Dataset Statistics Total Tokens: ~22.95B (22,948,244,480) Max Tokens per Sample: 512 Max Characters per Sample: 5,120 (10 chars/token estimate) Source Dataset: FineWeb-Edu 350BT Random Seed: 42 Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/JamesConley/fineweb-sample-22.95B-512.texttext-generation10M<n<100M0 likes419 downloads11mo agoHugging Face11xu-song /cc100-samplesThe cc100-samples is a subset which contains first 10,000 lines of cc100. Languages To load a language which isn't part of the config, all you need to do is specify the language code in the config. You can find the valid languages in Homepage section of Dataset Description: https://data.statmt.org/cc-100/ E.g. dataset = load_dataset("cc100-samples", lang="en") VALID_CODES = [ "am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de", "el"… See the full description on the dataset page: https://huggingface.co/datasets/xu-song/cc100-samples.texttext-generation1M<n<10M6 likes409 downloads2y agoHugging Face12FredyRivera-dev /LLaDA-Sample-ES Dataset: LLaDA-Sample-ES Base: crscardellino/spanish_billion_words Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~ 652,089 Shards: 65… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-ES.text-generation100M<n<1B1 likes406 downloads4mo agoHugging Face13BEE-spoke-data /TxT360-5M-sample-en BEE-spoke-data/TxT360-5M-sample-en english only sample from LLM360/TxT360: min length 256 GPT-4 tokens max length 24576 GPT-4 tokens GPT-4 tiktoken token count: token_count count 5.000000e+06 mean 1.003614e+03 std 1.424231e+03 min 2.570000e+02 25% 4.020000e+02 50% 6.220000e+02 75% 1.050000e+03 max 2.457400e+04 Total count: 5018.07 M tokens texttext-generation10M<n<100M3 likes389 downloads10mo agoHugging Face14voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face15Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes321 downloads1y agoHugging Face16LianeMarilin /tb3-tb4-sample-ml-checkpoint-reshard-recovery TB3/TB4 Sample Dataset Card 1. Dataset Overview This repository provides a Harbor terminal-agent benchmark task for the TB3/TB4 Sample stage: ml-checkpoint-reshard-recovery. The task requires an agent to repair an offline distributed-training checkpoint resharding and recovery tool and satisfy an independent, offline, programmatic verifier. Field Value Task ID ml-checkpoint-reshard-recovery Primary Domain ML Related tags distributed-systems, storage… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/tb3-tb4-sample-ml-checkpoint-reshard-recovery.text-generationn<1K0 likes312 downloads26d agoHugging Face17Mindgard /evaded-prompt-injection-and-jailbreak-samplesgatedThis dataset originates from our paper 'Bypassing Prompt Injection and Jailbreak Detection in LLM Guardrails'. The dataset contains a mixture of prompt injections and jailbreak samples modified via character injection and adversarial ML evasion techniques (Techniques can be found within the paper above). For each sample we provide the original unaltered prompt and a modified prompt, the attack_name outlines which attack technique was used to modify the sample. Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/Mindgard/evaded-prompt-injection-and-jailbreak-samples.texttext-classification10K<n<100K24 likes243 downloads1y agoHugging Face18LianeMarilin /enterprise-agent-aa-samples Dataset Card Dataset Description Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks. Task: enterprise tool-use and agent-trajectory evaluation Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.texttext-generationn<1K1 likes243 downloads1mo agoHugging Face19HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes233 downloads2mo agoHugging Face20nht10 /cx_sampled CulturaX sampled text pools Use UPLOAD_COMPLETE.json before consuming this release. Its absence means the upload is incomplete. Pin the completed repository revision for experiments. Original, unpacked document text from uonlp/CulturaX. This is a fresh sample, independent of nguyenhuuthuat09/CulturaX_sampled. No token sequences or training caches are distributed. Each subset has a fixed validation set and a nested training prefix suitable for smaller token budgets. Token counts… See the full description on the dataset page: https://huggingface.co/datasets/nht10/cx_sampled.texttext-generation100M<n<1B0 likes230 downloads12d agoHugging Face21ks46 /urls-sampled URLs (hash-sampled) The same 74,918,894,107 URLs as ks46/urls, partitioned by xxh3_64 range into 2,048 chunks of ≈36.6 M rows instead of by SURT key range. Each chunk is a uniform random sample of the whole corpus, and a URL's chunk depends on nothing but the URL itself. Why this exists The SURT layout groups the web by host: shard 1,000 is a contiguous slice of the key space, so it holds whole sites and nothing about any other site. That is what you want for… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-sampled.texttext-generation10B<n<100B0 likes211 downloads2mo agoHugging Face22openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes199 downloads1y agoHugging Face23ll922 /RedPajama-Data-1T-Sample-Backup RedPajama Data 1T Sample Backup This dataset is a backup mirror of togethercomputer/RedPajama-Data-1T-Sample. It is provided for easier access when the original dataset is unavailable or difficult to download. Usage Original: from datasets import load_dataset ds = load_dataset( "togethercomputer/RedPajama-Data-1T-Sample", split="train", trust_remote_code=True, ) Backup: from datasets import load_dataset ds = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ll922/RedPajama-Data-1T-Sample-Backup.texttext-generation100K<n<1M0 likes198 downloads5mo agoHugging Face24lemoncmd /lldms-associative-memory-samples LLDMs Associative Memory — Generated Samples Model-generated text for the paper: Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov, Matteo Negri Accepted to EMNLP 2026 (Main Conference). arXiv:2604.26841 · paper · code · checkpoints 29.5 million generated sequences (~3.8B tokens) sampled from the released checkpoints — one generation run per (model size, training-set fraction). These… See the full description on the dataset page: https://huggingface.co/datasets/lemoncmd/lldms-associative-memory-samples.text-generation10M<n<100M0 likes193 downloads1mo agoHugging Face25DeceptionPro /EDR_Telemetry_SampleThis dataset contains raw Endpoint Detection & Response (EDR) telemetry captured during controlled Deception.Pro malware sandbox operations on an enterprise Active Directory network. Unlike most malware sandboxes — which detonate samples for roughly 30 minutes — our operations run for hours or days per analysis, capturing the full arc of adversary behavior. The data represents a full-fidelity snapshot of system activity recorded while threat actors interacted with a live deception environment… See the full description on the dataset page: https://huggingface.co/datasets/DeceptionPro/EDR_Telemetry_Sample.question-answeringn<1K8 likes189 downloads5mo agoHugging Face26AxiomSetLabs /stem-scientific-code-sample AxiomSet Labs STEM Scientific-Code Sample A 30-task sample of STEM reasoning and scientific-code problems across five domains. Domains Biology: 6 tasks Chemistry: 6 tasks Materials Science: 6 tasks Mathematics: 6 tasks Physics: 6 tasks Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions. Files data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.texttext-generationn<1K0 likes174 downloads2mo agoHugging Face27superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face28insop /ultrafineweb-v1.0-sample-100BT ultrafineweb-v1.0-sample-100BT A random sample of Ultra-FineWeb v1.0 (English) (files data/ultrafineweb_en/*.parquet at revision 02c85641e3): 60,307,595 of its 1,159,254,991 documents, ~99.9B tokens (estimated at 2.39 UTF-8 bytes per token of an 8k BPE tokenizer), globally shuffled. split documents shards train 60,248,654 1,023 validation 58,941 1 Columns text: the document text (the source's content). score, source: copied unchanged from the… See the full description on the dataset page: https://huggingface.co/datasets/insop/ultrafineweb-v1.0-sample-100BT.texttext-generation10M<n<100M0 likes168 downloads9d agoHugging Face29Dynamicresponselabs /JASON-High-Stakes-AI-Evaluation-Samples J.A.S.O.N. Evaluation Sample Previews V01-V29 Dynamic Response Labs develops specialized data and evaluation resources for high-stakes AI. This public preview introduces the breadth of the J.A.S.O.N. Framework through 29 domain volumes spanning financial stress, operational disruption, coercion and exploitation, cyber incidents, healthcare finance, automated systems, and other consequential contexts. The collection contains 31 compact preview records. It is designed to help… See the full description on the dataset page: https://huggingface.co/datasets/Dynamicresponselabs/JASON-High-Stakes-AI-Evaluation-Samples.texttext-generationn<1K0 likes163 downloads13d agoHugging Face30insop /ultrafineweb-l1-hq-sample-100BT ultrafineweb-l1-hq-sample-100BT A random sample of Ultra-FineWeb L1 English HQ (files data/ultrafineweb_l1_en_hq/*/*.parquet at revision 02c85641e3): 45,129,641 of its 144,908,921 documents, ~99.8B tokens (estimated at 2.51 UTF-8 bytes per token of an 8k BPE tokenizer), globally shuffled. split documents shards train 45,085,116 1,023 validation 44,525 1 Columns text: the document text (the source's content). meta: copied unchanged from the… See the full description on the dataset page: https://huggingface.co/datasets/insop/ultrafineweb-l1-hq-sample-100BT.texttext-generation10M<n<100M0 likes158 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.