Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amongglue /muse_textbookstext1M<n<10M3 likes11k downloads3y agoHugging Face02gtfintechlab /ipo-text SEC IPO Filings Dataset A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants. Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.tabulartext-classification100K<n<1M6 likes9.6k downloads8mo agoHugging Face03XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes5.1k downloads7d agoHugging Face04agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes4k downloads11mo agoHugging Face05opencompass /TextEdit TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models Danni Yang, Sitao Chen, Changyao Tian If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details. 🎉 News [2026/03/06] TextEdit benchmark released. [2026/03/06] Evaluation code and initial baselines released. [2026/03/06] Leaderboard updated with latest models. 📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.imageimage-to-image1K<n<10K9 likes3.9k downloads7mo agoHugging Face06MedRAG /textbooks The Textbooks Corpus in MedRAG This HF dataset contains the chunked snippets from the Textbooks corpus used in MedRAG. It can be used for medical Retrieval-Augmented Generation (RAG). Dataset Details Dataset Descriptions Textbooks is a collection of 18 widely used medical textbooks, which are important references for students taking the United States Medical Licensing Examination (USLME). In MedRAG, the textbooks are processed as chunks with no more than 1000… See the full description on the dataset page: https://huggingface.co/datasets/MedRAG/textbooks.textquestion-answering100K<n<1M64 likes2.7k downloads3y agoHugging Face07myduy /vnexpress_plain_texttext100K<n<1M0 likes1.5k downloads1y agoHugging Face08sailor2 /sea-pdf-texttext10M<n<100M1 likes1.3k downloads2y agoHugging Face09myduy /vtv_plain_texttext100K<n<1M0 likes1.1k downloads1y agoHugging Face10ndurkee /muse_textbookstext100K<n<1M0 likes1k downloads3y agoHugging Face11keisuke-miyako /text-commands-2026-0426 Commands (reverse description) Clean summary of 4D language reference. This dataset was generated with Mistral Large 3. Example In the context of 4D version 21, when developing compiled database applications, a critical challenge arises during the execution of long-running or tightly bound loops that do not yield control back to the system. Such loops can monopolize processor resources, leading to unresponsive behavior and preventing the execution of other essential… See the full description on the dataset page: https://huggingface.co/datasets/keisuke-miyako/text-commands-2026-0426.text1K<n<10K0 likes1k downloads6mo agoHugging Face12keisuke-miyako /text-commands-2026-0405textn<1K0 likes751 downloads6mo agoHugging Face13Nbardy /science-theory-textbookstext10K<n<100K9 likes672 downloads3y agoHugging Face14semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes607 downloads3y agoHugging Face15keisuke-miyako /text-commands-2026-0425 Commands Clean summary of 4D language reference. Abstract LLMs are generally incapable of understanding 4D code. LoRA by exposure to raw source code would actually increase the rate of hallucination as the model gets confused between 4D code and C#, Visual Basic, or JavaScript. CPT, or continued pre-training, based on grammar and vocabulary should moderate the model's attention before extensive fine-tuning using raw source code. This dataset was generated with Grok 4.20… See the full description on the dataset page: https://huggingface.co/datasets/keisuke-miyako/text-commands-2026-0425.text1K<n<10K0 likes540 downloads6mo agoHugging Face16ssz1111 /SpokenWOZ-Train-Text What is SpokenWOZ? SpokenWOZ is a large-scale multi-domain speech-text dataset for spoken task-oriented dialogue modeling, which consists of 203k turns, 5.7k dialogues and 249 hours audios from realistic human-to-human spoken conversations. Why SpokenWOZ? The majority of existing TOD datasets are constructed via writing or paraphrasing from annotators rather than being collected from realistic spoken conversations. The written TDO datasets may not be representative of the… See the full description on the dataset page: https://huggingface.co/datasets/ssz1111/SpokenWOZ-Train-Text.text1K<n<10K1 likes528 downloads9mo agoHugging Face17yachay /text_coordinates_regions Dataset Card for Multilingual Geo-Tagged Social Media Posts (by 123 world regions) Dataset Summary The "Regions" dataset is a multilingual corpus that encompasses textual data from the 123 most populated regions worldwide, with each region's data organized into separate .json files. This dataset consists of approximately 500,000 text samples, each paired with its geographic coordinates. Key Features: Textual Data: The dataset contains 500,000 text samples.… See the full description on the dataset page: https://huggingface.co/datasets/yachay/text_coordinates_regions.textfeature-extraction100K<n<1M10 likes515 downloads3y agoHugging Face18SkySyrup /muse_textbookstext100K<n<1M1 likes513 downloads3y agoHugging Face19while-ai /text-to-sql-shop Text-to-SQL on a seeded store schema, with checkpoints Recipe: recipes/04-train/text-to-sql · Collection: Analyst A question about an online store's database in, one PostgreSQL query out, graded by a program: run the query, compare the result set to the gold query's result. The schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and the benchmark runner are the recipes/04-train/text-to-sql recipe in the open-source whileai SDK. Splits |… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.tabular10K<n<100K0 likes496 downloads18d agoHugging Face20keisuke-miyako /text-commands-2026-0412text1K<n<10K0 likes489 downloads6mo agoHugging Face21malaysia-ai /mosaic-dedup-text-dataset-filtered Mosaic format for filtered dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-filtered-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset-filtered.textn<1K0 likes482 downloads3y agoHugging Face22evalitahf /textual_entailmentThe Textual Entailment dataset contains 800 pairs of Italian sentences, extracted from Wikipedia, and annotated for the presence of textual entailment. A pair of texts consists of T (for text) and H (hypothesis). Textual entailment is defined as a directional relationship between such pairs. The hypothesis must be fully entailed by the text. The dataset has been created and used for the Textual Entailment Task (http://www.evalita.it/2009/tasks/te), organised as part of the EVALITA 2009… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/textual_entailment.texttext-classificationn<1K0 likes462 downloads2y agoHugging Face23keisuke-miyako /text-commands-2026-0421text1K<n<10K1 likes455 downloads6mo agoHugging Face24keisuke-miyako /text-commands-2026-0422 Commands Clean summary of 4D language reference. Abstract LLMs are generally incapable of understanding 4D code. LoRA by exposure to raw source code would actually increase the rate of hallucination as the model gets confused between 4D code and C#, Visual Basic, or JavaScript. CPT, or continued pre-training, based on grammar and vocabulary should moderate the model's attention before extensive fine-tuning using raw source code. This dataset was generated with Mistral… See the full description on the dataset page: https://huggingface.co/datasets/keisuke-miyako/text-commands-2026-0422.text1K<n<10K0 likes409 downloads6mo agoHugging Face25malaysia-ai /mosaic-dedup-text-dataset Mosaic format for dedup text dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-dedup-text-dataset-4096.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-dedup-text-dataset.textn<1K0 likes395 downloads3y agoHugging Face26keisuke-miyako /text-commands-2026-0408text1K<n<10K0 likes382 downloads6mo agoHugging Face27codaco /text CoDaCo - texts dataset This dataset was created using codaco.app. Description All data contributed to this campaign goes to the global CoDaCo datasets. Labels This dataset includes the following labels: Summaries Entities Tags Emotions AI generated Quality rating License This dataset is licensed under CC BY 4.0. You are free to share and adapt it for any purpose, including commercially, as long as you give appropriate credit. See… See the full description on the dataset page: https://huggingface.co/datasets/codaco/text.textn<1K0 likes365 downloads22h agoHugging Face28Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes355 downloads3y agoHugging Face29keisuke-miyako /text-commands-2026-0429text1K<n<10K0 likes354 downloads6mo agoHugging Face30Nbardy /wild-science-theory-textbookstext10K<n<100K3 likes342 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.