Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01legacy-datasets /wikipediaWikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).text-generationn<1K675 likes56k downloads3y agoHugging Face02legacy-datasets /c4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset by AllenAI.text-generation100M<n<1B242 likes17k downloads3y agoHugging Face03lesserfield /4chan-datasetsPlease see repo to turn the text file into json/csv format Deleted some boards, since they are already archived by https://archive.4plebs.org/ texttext-generation34 likes5.6k downloads3y agoHugging Face04Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes4.8k downloads3y agoHugging Face05google-research-datasets /taskmaster2Taskmaster is dataset for goal oriented conversations. The Taskmaster-2 dataset consists of 17,289 dialogs in the seven domains which include restaurants, food ordering, movies, hotels, flights, music and sports. Unlike Taskmaster-1, which includes both written "self-dialogs" and spoken two-person dialogs, Taskmaster-2 consists entirely of spoken two-person dialogs. In addition, while Taskmaster-1 is almost exclusively task-based, Taskmaster-2 contains a good number of search- and recommendation-oriented dialogs. All dialogs in this release were created using a Wizard of Oz (WOz) methodology in which crowdsourced workers played the role of a 'user' and trained call center operators played the role of the 'assistant'. In this way, users were led to believe they were interacting with an automated system that “spoke” using text-to-speech (TTS) even though it was in fact a human behind the scenes. As a result, users could express themselves however they chose in the context of an automated interface.text-generation1K<n<10K7 likes2.5k downloads3y agoHugging Face06yulan-team /YuLan-Mini-Datasets YuLan-Mini Datasets 🔥 Updated (April 11, 2025): For a clearer presentation of the information, see the table at this link: link. This datasets contains: Classified data using python-edu-scorer and fineweb-edu-classifier Synthesized data (math, code, instruction, ...) Retrieved data using math, code, and reasoninig-classifier Notice Since we have used BPE-Dropout, in order to ensure accuracy, the data we uploaded is tokenized.… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets.text-generation10M<n<100M10 likes2.5k downloads1y agoHugging Face07M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.1k downloads3y agoHugging Face08defunct-datasets /amazon_reviews_multiWe provide an Amazon product reviews dataset for multilingual text classification. The dataset contains reviews in English, Japanese, German, French, Chinese and Spanish, collected between November 1, 2015 and November 1, 2019. Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID and the coarse-grained product category (e.g. ‘books’, ‘appliances’, etc.) The corpus is balanced across stars, so each star rating constitutes 20% of the reviews in each language. For each language, there are 200,000, 5,000 and 5,000 reviews in the training, development and test sets respectively. The maximum number of reviews per reviewer is 20 and the maximum number of reviews per product is 20. All reviews are truncated after 2,000 characters, and all reviews are at least 20 characters long. Note that the language of a review does not necessarily match the language of its marketplace (e.g. reviews from amazon.de are primarily written in German, but could also be written in English, etc.). For this reason, we applied a language detection algorithm based on the work in Bojanowski et al. (2017) to determine the language of the review text and we removed reviews that were not written in the expected language.summarization100K<n<1M102 likes2k downloads3y agoHugging Face09JiaaqiLiu /SkillArena-datasets SkillArena Offline Datasets Offline evaluation data for SkillArena — a validated automatic benchmark generation framework for AI agent skills, targeting NeurIPS 2026 Datasets & Benchmarks Track. Overview This dataset provides domain-specific task input data for 289 AI agent skills across 13 domains. Each skill has 50 curated data files designed as meaningful agent task inputs — files an agent could receive and act upon (analyze, transform, validate, or generate from). The… See the full description on the dataset page: https://huggingface.co/datasets/JiaaqiLiu/SkillArena-datasets.text-generation10K<n<100K0 likes1.9k downloads7mo agoHugging Face10YanZhanPKU /dLLM-PRM-Gap-Datasets dLLM PRM Gap · Datasets 💻 Code &nbsp; • &nbsp; 🤗 Collection &nbsp; • &nbsp; 📄 Paper This dataset repository contains the selected trajectory corpus and small diagnostic artifacts for dLLM PRM Gap, a controlled study of process reward model (PRM) guidance and outcome reward model (ORM) reranking in discrete diffusion reasoning. Public release. This dataset accompanies the arXiv paper and the dLLM PRM Gap code release. News 2026-09: Accepted at NeurIPS… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/dLLM-PRM-Gap-Datasets.text-generation10K<n<100K1 likes1.7k downloads11d agoHugging Face11THUIAR /MMLA-Datasets Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark 1. Introduction MMLA is the first comprehensive multimodal language analysis benchmark for evaluating foundation models. It has the following features: Large Scale: 61K+ multimodal samples. Various Sources: 9 datasets. Three Modalities: text, video, and audio Both Acting and Real-world Scenarios: films, TV series, YouTube, Vimeo, Bilibili, TED, improvised scripts, etc. Six Core… See the full description on the dataset page: https://huggingface.co/datasets/THUIAR/MMLA-Datasets.textzero-shot-classification10K<n<100K4 likes1.2k downloads1y agoHugging Face12inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face13google-research-datasets /schema_guided_dstc8The Schema-Guided Dialogue dataset (SGD) was developed for the Dialogue State Tracking task of the Eights Dialogue Systems Technology Challenge (dstc8). The SGD dataset consists of over 18k annotated multi-domain, task-oriented conversations between a human and a virtual assistant. These conversations involve interactions with services and APIs spanning 17 domains, ranging from banks and events to media, calendar, travel, and weather. For most of these domains, the SGD dataset contains multiple different APIs, many of which have overlapping functionalities but different interfaces, which reflects common real-world scenarios.text-generation10K<n<100K15 likes935 downloads3y agoHugging Face14legacy-datasets /mc4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI.text-generationn<1K155 likes828 downloads3y agoHugging Face15defunct-datasets /amazon_us_reviewsAmazon Customer Reviews (a.k.a. Product Reviews) is one of Amazons iconic products. In a period of over two decades since the first review in 1995, millions of Amazon customers have contributed over a hundred million reviews to express opinions and describe their experiences regarding products on the Amazon.com website. This makes Amazon Customer Reviews a rich source of information for academic researchers in the fields of Natural Language Processing (NLP), Information Retrieval (IR), and Machine Learning (ML), amongst others. Accordingly, we are releasing this data to further research in multiple disciplines related to understanding customer product experiences. Specifically, this dataset was constructed to represent a sample of customer evaluations and opinions, variation in the perception of a product across geographical regions, and promotional intent or bias in reviews. Over 130+ million customer reviews are available to researchers as part of this release. The data is available in TSV files in the amazon-reviews-pds S3 bucket in AWS US East Region. Each line in the data files corresponds to an individual review (tab delimited, with no quote and escape characters). Each Dataset contains the following columns: - marketplace: 2 letter country code of the marketplace where the review was written. - customer_id: Random identifier that can be used to aggregate reviews written by a single author. - review_id: The unique ID of the review. - product_id: The unique Product ID the review pertains to. In the multilingual dataset the reviews for the same product in different countries can be grouped by the same product_id. - product_parent: Random identifier that can be used to aggregate reviews for the same product. - product_title: Title of the product. - product_category: Broad product category that can be used to group reviews (also used to group the dataset into coherent parts). - star_rating: The 1-5 star rating of the review. - helpful_votes: Number of helpful votes. - total_votes: Number of total votes the review received. - vine: Review was written as part of the Vine program. - verified_purchase: The review is on a verified purchase. - review_headline: The title of the review. - review_body: The review text. - review_date: The date the review was written.summarization100M<n<1B75 likes660 downloads3y agoHugging Face16community-datasets /wiki_snippets Dataset Card for "wiki_snippets" Dataset Summary Wikipedia version split into plain text snippets for dense semantic indexing. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in English: wiki40b_en_100_0: Wiki-40B wikipedia_en_100_0: Wikipedia Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.tabulartext-generation10M<n<100M6 likes577 downloads2y agoHugging Face17zhangdw /to-tool-call-datasets 🛠️ To-Tool-Call Datasets A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training &nbsp;&nbsp;&nbsp;&nbsp; To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention. Quick Start · At a Glance · Format · Sources · Training Notes [!IMPORTANT] This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.texttext-generation1K<n<10K4 likes569 downloads4mo agoHugging Face18yulan-team /YuLan-Mini-Datasets-Phasae-27The tokenized datasets for YuLan-Mini phase 27, where each line has been packed to 28K tokens. Usage dataset = [] dataset_path = "/path/to/YuLan-Mini-Datasets-Phasae-27" seed = 42 for data_name in sorted(os.listdir(dataset_path)): d = load_dataset( os.path.join(dataset_path, data_name), split="train", num_proc=8, ) dataset.append(d) print(f"Num subsets: {len(dataset)}") dataset = concatenate_datasets(dataset).shuffle(seed=seed)… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets-Phasae-27.question-answering10B<n<100B0 likes526 downloads2y agoHugging Face19yulan-team /YuLan-Mini-Text-Datasets News [2025.04.11] Add dataset mixture: link. [2025.03.30] Text datasets upload finished. This is text dataset. 这是文本格式的数据集。 Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here. 由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。 For more information, please refer to our datasets details and preprocess details. Contributing We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.tabulartext-generation100M<n<1B12 likes521 downloads1y agoHugging Face20proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes435 downloads6mo agoHugging Face21yulan-team /YuLan-Mini-Datasets-Phasae-26The tokenized datasets for YuLan-Mini phase 26, where each line has been packed to 28K tokens. Usage dataset = [] dataset_path = "/path/to/YuLan-Mini-Datasets-Phasae-26" seed = 42 for data_name in sorted(os.listdir(dataset_path)): d = load_dataset( os.path.join(dataset_path, data_name), split="train", num_proc=8, ) dataset.append(d) print(f"Num subsets: {len(dataset)}") dataset = concatenate_datasets(dataset).shuffle(seed=seed)… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets-Phasae-26.question-answering1M<n<10M0 likes434 downloads2y agoHugging Face22amd /AIG-Datasets AMD AIG GPU Kernel Datasets AMD AIG-Datasets is a collection of GPU-kernel generation, translation, optimization, and ROCm-library supervision data. It contains PyTorch/CUDA-to-HIP, HIP-to-HIP, PyTorch-to-Triton, and production-grounded rocBLAS/rocSOLVER entries, together with metadata, samples, conversion utilities, and reproducible evaluation tools. The repository is organized into versioned releases. New training and evaluation workflows should use the unified-schema datasets… See the full description on the dataset page: https://huggingface.co/datasets/amd/AIG-Datasets.text-generation4 likes429 downloads2mo agoHugging Face23defunct-datasets /the_pile_books3This dataset is Shawn Presser's work and is part of EleutherAi/The Pile dataset. This dataset contains all of bibliotik in plain .txt form, aka 197,000 books processed in exactly the same way as did for bookcorpusopen (a.k.a. books1). seems to be similar to OpenAI's mysterious "books2" dataset referenced in their papers. Unfortunately OpenAI will not give details, so we know very little about any differences. People suspect it's "all of libgen", but it's purely conjecture.text-generation100K<n<1M153 likes426 downloads3y agoHugging Face24thefinalboss /fractus-datasets Fractus Datasets — the neuroscience-grounded training corpus A proprietary, neuroscience-derived training corpus for the Fractus Continuous Thought Engine — ~3–4B tokens mapping real brain mechanisms to software/AI architecture, plus cognitive skills, code, esoteric tradition, and lexical knowledge. Curator: Philippe-Antoine Robert · rpa.tu@proton.me · 2026 What this dataset collection IS Fractus is a non-transformer Continuous Cognitive Agent whose architecture… See the full description on the dataset page: https://huggingface.co/datasets/thefinalboss/fractus-datasets.texttext-generationn<1K0 likes407 downloads2mo agoHugging Face25YanZhanPKU /SLCA-GRPO-Datasets SLCA-GRPO · Datasets 📄 Paper (arXiv:2609.29050) &nbsp; • &nbsp; 💻 Code &nbsp; • &nbsp; 🤗 Collection This repository contains every processed data file read by SLCA-GRPO, the segment-locked credit assignment estimator for tool-calling RL introduced in "SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL". A tool-calling rollout opens with structured tool-call tokens and closes with free-form summary text; SLCA-GRPO normalises the two segment rewards… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/SLCA-GRPO-Datasets.texttext-generation100K<n<1M2 likes385 downloads15d agoHugging Face26defunct-datasets /bookcorpusopenBooks are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This version of bookcorpus has 17868 dataset items (books). Each item contains two fields: title and text. The title is the name of the book (just the file name) while text contains unprocessed book text. The bookcorpus has been prepared by Shawn Presser and is generously hosted by The-Eye. The-Eye is a non-profit, community driven platform dedicated to the archiving and long-term preservation of any and all data including but by no means limited to... websites, books, games, software, video, audio, other digital-obscura and ideas.text-generation10K<n<100K39 likes384 downloads3y agoHugging Face27p11-p11 /chess_datasets Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.texttext-generation1M<n<10M0 likes370 downloads2y agoHugging Face28reasonwang /ToolGen-Datasets How to use? Before making use of this dataset, you may need to add the tokens to the vocabulary. For HuggingFace transformers tokenizer, the following is an example code snippet to add tokens. from unidecode import unidecode import transformers with open('virtual_tokens.txt', 'r') as f: virtual_tokens = f.readlines() virtual_tokens = [unidecode(vt.strip()) for vt in virtual_tokens] model_name_or_path = "meta-llama/Meta-Llama-3-8B" # Load tokenizer and add tokens into… See the full description on the dataset page: https://huggingface.co/datasets/reasonwang/ToolGen-Datasets.texttext-generation100K<n<1M8 likes369 downloads2y agoHugging Face29Similoluwa /capstone-datasets AIMS AI Research Foundations – Capstone Datasets Ready-to-use train / validation / test datasets for the AIRF Capstone Project Recipes, a set of short, end-to-end recipes for implementing and evaluating AI capstone projects, for university lecturers and learners across Africa, built around small open-weight models (Gemma 1B / 4B). Every dataset is a configuration of this repository. Every configuration has train, validation and test splits with labels or reference answers.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/capstone-datasets.tabulartext-classification10K<n<100K0 likes360 downloads8d agoHugging Face30Dionysianspirit /SafeCRS-datasets SafeCRS Datasets Datasets for training safety-aware conversational recommendation models. The default viewer shows the SafeRec SFT test split — the safety-annotated evaluation set with one explicit user sensitivity trait per sample. SafeRec SFT Dataset (test split preview) Each sample contains a user conversation requesting movie recommendations, an assigned sensitivity trait, constraint-filtered ground truth, and the constraint text injected into the prompt.… See the full description on the dataset page: https://huggingface.co/datasets/Dionysianspirit/SafeCRS-datasets.texttext-generationn<1K0 likes349 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.