Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gtfintechlab /ipo-text SEC IPO Filings Dataset A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants. Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.tabulartext-classification100K<n<1M6 likes9.6k downloads8mo agoHugging Face02Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.6k downloads2mo agoHugging Face03PleIAs /WTO-Text Dataset Card for WTO Documents Dataset Dataset Overview Title: WTO Documents Dataset Source: World Trade Organization Documents Online Description: The WTO Documents Dataset is a comprehensive collection of official documentation from the World Trade Organization (WTO). This dataset is sourced from the WTO's official Documents Online platform, which provides access to documents in the three official languages (English, French, and Spanish) from 1995 onwards. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/WTO-Text.tabular100K<n<1M9 likes3.6k downloads2y agoHugging Face04DorayakiLin /TextOnly_FromRLBench_CloseBox_24K_unfixedtabular1M<n<10M0 likes2.4k downloads11mo agoHugging Face05textql /Argo-Bench Argo-Bench One simulated year of a New York City food-delivery company, exported to an Oracle E-Business Suite 12.2 warehouse of 235 tables and 7.54 billion rows. Leaderboard & Demo · How it works · Paper · Code An enterprise ERP you can download The data that enterprise analytics runs on is rarely public. Large companies keep their orders, payouts, ledgers and customer records in ERP systems such as Oracle E-Business Suite and SAP, extended with custom tables of… See the full description on the dataset page: https://huggingface.co/datasets/textql/Argo-Bench.tabular1B<n<10B5 likes2k downloads9d agoHugging Face06textql /Decision-Bench Decision-Bench One simulated year of a New York City food-delivery company, exported to an Oracle E-Business Suite 12.2 warehouse of 235 tables and 7.54 billion rows. Leaderboard & Demo · How it works · Paper · Code An enterprise ERP you can download The data that enterprise analytics runs on is rarely public. Large companies keep their orders, payouts, ledgers and customer records in ERP systems such as Oracle E-Business Suite and SAP, extended with custom… See the full description on the dataset page: https://huggingface.co/datasets/textql/Decision-Bench.tabular1B<n<10B1 likes1.3k downloads9d agoHugging Face07crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes1.3k downloads5mo agoHugging Face08Rapidata /text-2-video-human-preferences Rapidata Video Generation Preference Dataset This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set: Sora Hunyouan Pika 2.0 Runway ML Alpha Luma Ray 2 Explore our latest model rankings on our website. If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.imagetext-to-video1K<n<10K21 likes1.3k downloads2y agoHugging Face09Rapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes1.2k downloads1y agoHugging Face10trl-lab /SQaLe-text-to-SQL-dataset 🧮 SQALE: A Large-Scale Semi-Synthetic Dataset SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas. It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions. The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.tabulartext-generation100K<n<1M21 likes901 downloads7mo agoHugging Face11matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes857 downloads3y agoHugging Face12Rapidata /text-2-video-human-preferences-veo3 Rapidata Video Generation Veo 3 Human Preference In this dataset, ~46k human responses from ~20k human annotators were collected to evaluate Veo3 video generation model on our benchmark. This dataset was collected in roughly 35 minutes using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.imagevideo-classification1K<n<10K20 likes837 downloads1y agoHugging Face13lightblue /text_ratingsTodo - Write dataset card tabular1M<n<10M4 likes776 downloads2y agoHugging Face14ourafla /Mental-Health_Text-Classification_Dataset Mental Health Text Classification Dataset (4-Class) Dataset Description This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models. The repository includes: An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.texttext-classification10K<n<100K13 likes774 downloads10mo agoHugging Face15semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes607 downloads3y agoHugging Face16nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes550 downloads2y agoHugging Face17Quoron /EEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics. The raw EEG data and the datasheet are available at https://osf.io/xh3g5/. See code repository for benchmark results. EEG data acquisition: Explanations of the variables: event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.tabulartext-classification10K<n<100K8 likes540 downloads1y agoHugging Face18yulan-team /YuLan-Mini-Text-Datasets News [2025.04.11] Add dataset mixture: link. [2025.03.30] Text datasets upload finished. This is text dataset. 这是文本格式的数据集。 Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here. 由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。 For more information, please refer to our datasets details and preprocess details. Contributing We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.tabulartext-generation100M<n<1B12 likes521 downloads1y agoHugging Face19while-ai /text-to-sql-shop Text-to-SQL on a seeded store schema, with checkpoints Recipe: recipes/04-train/text-to-sql · Collection: Analyst A question about an online store's database in, one PostgreSQL query out, graded by a program: run the query, compare the result set to the gold query's result. The schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and the benchmark runner are the recipes/04-train/text-to-sql recipe in the open-source whileai SDK. Splits |… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.tabular10K<n<100K0 likes496 downloads18d agoHugging Face20rihim /icl-selfplay-vs-text-results Self-play vs. text pretraining: ICL results These are the full evaluation results from comparing the in-context learning (ICL) of: the self-play learners from Self-Play Pretraining with Zero Data, and same-size models trained on ordinary web text for the same number of tokens. Code and write-up: github.com/mihir-s-05/icl-selfplay-vs-text. Text-model checkpoints: rihim/icl-selfplay-vs-text-checkpoints. What was scored Model families (the arm column): arm… See the full description on the dataset page: https://huggingface.co/datasets/rihim/icl-selfplay-vs-text-results.tabular10K<n<100K0 likes478 downloads12d agoHugging Face21Rapidata /text-2-video-human-preferences-moonvalley-marey Rapidata Video Generation Marey Pro Human Preference In this dataset, ~75k human responses from ~15k human annotators were collected to evaluate Marey video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-moonvalley-marey.imagevideo-classification1K<n<10K7 likes462 downloads1y agoHugging Face22Rapidata /text-2-video-human-preferences-veo3.1 Rapidata Video Generation Veo 3.1 Human Preference In this dataset, ~74k human responses from ~23k human annotators were collected to evaluate the Veo 3.1 video generation model on our benchmark. This dataset was collected using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please consider liking it ❤️… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-veo3.1.imagevideo-classification1K<n<10K9 likes453 downloads11mo agoHugging Face23matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes417 downloads3y agoHugging Face24Glide-py /spider-text-to-sql Spider Text-to-SQL with LLM-Judge Labels This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4. Files File Description spider_dataset.parquet Full dataset with predictions and labels scripts/ Reproduction scripts (see below) Dataset statistics Source: Spider 1.0 training split (train_spider.json) Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.tabulartext-generation1K<n<10K0 likes412 downloads4mo agoHugging Face25enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes408 downloads2y agoHugging Face26matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes403 downloads3y agoHugging Face27nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes402 downloads3y agoHugging Face28cloud0day3 /alania-domain-text-tr Alania Turkish Domain Text English · Türkçe 1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud. We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr.tabulartext-to-speech1M<n<10M19 likes380 downloads7d agoHugging Face29datapointai /text-to-speech-human-preferences-315kgated Text-to-speech human preferences: 315K votes across 15 models This gated dataset contains the evaluation record behind Datapoint Audio Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech models in a complete round-robin over 300 English prompts. The prompt set covers eight practical voice-agent categories, and every generated sample is included as a typed audio record. The source evaluation collected 357,651 completed responses. The published benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.audiotext-to-speech100K<n<1M41 likes374 downloads1mo agoHugging Face30ksopyla /cogito-text-world Cogito Text World Part of Cogito Capability Checks — does your model actually use what it read? Small synthetic tests with exact answers, guessing floors and a length ladder (1k → 128k tokens). Project page: ai.ksopyla.com/projects/cogito-capability-datasets TL;DR. Train a small language model from scratch on simple English stories with facts woven in, then ask it about those facts in documents of 1k to 128k tokens (trained at 4k). Seven question types, from copying a sign to… See the full description on the dataset page: https://huggingface.co/datasets/ksopyla/cogito-text-world.tabulartext-generation100K<n<1M0 likes360 downloads6h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.