Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes475 downloads5mo agoHugging Face02Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes375 downloads1y agoHugging Face03Unified-Language-Model-Alignment /Anthropic_HH_Golden Dataset Card for Anthropic_HH_Golden This dataset is constructed to test the ULMA technique as mentioned in the paper Unified Language Model Alignment with Demonstration and Point-wise Human Preference (under review, and an arxiv link will be provided soon). They show that replacing the positive samples in a preference dataset by high-quality demonstration data (golden data) greatly improves the performance of various alignment methods (RLHF, DPO, ULMA). In particular, the ULMA… See the full description on the dataset page: https://huggingface.co/datasets/Unified-Language-Model-Alignment/Anthropic_HH_Golden.text10K<n<100K40 likes221 downloads3y agoHugging Face04Plim /language_model_frtext1M<n<10M0 likes192 downloads4y agoHugging Face05xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K10 likes170 downloads3y agoHugging Face06beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging Face07astha /languagemodelsforRNNdecompositionThis repository is for the paper "Decomposing a Recurrent Neural Network into Modules for Enabling Reusability and Replacement". To use the data, there are two directories: language datasets: Contains the necessary Tatoeba files used for the experiments. We have experimented with 4 languages(English, French, Italian and German). language_models: Contains all trained language models and scripts to train them. It's organized in this way: language_models/{X}: contains language models for X… See the full description on the dataset page: https://huggingface.co/datasets/astha/languagemodelsforRNNdecomposition.text100K<n<1M1 likes115 downloads4y agoHugging Face08Arpuuu /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Arpuuu/nepalitext-language-model-dataset.texttext-generation10M<n<100M0 likes106 downloads6mo agoHugging Face09robotamski /language_models_lab_2text1M<n<10M0 likes93 downloads22d agoHugging Face10abidlabs /repro-how-much-can-language-models-memorize-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes90 downloads3mo agoHugging Face11linneripe /language_modelstext1M<n<10M1 likes81 downloads24d agoHugging Face12rosimeirecosta /c_corpus_br_finetuning_language_model_bert Dataset Card for "c_corpus_br_finetuning_language_model_bert" More Information needed text100K<n<1M3 likes68 downloads4y agoHugging Face13mznaser /Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset Data for the paper Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173 It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result. If you use the data, please cite the paper (BibTeX under Citation). The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.tabulartext-classification10K<n<100K0 likes52 downloads19d agoHugging Face14pierreguillou /lener_br_finetuning_language_model Dataset Card for "LeNER-Br language modeling" Dataset Summary The LeNER-Br language modeling dataset is a collection of legal texts in Portuguese from the LeNER-Br dataset (official site). The legal texts were downloaded from this link (93.6MB) and processed to create a DatasetDict with train and validation dataset (20%). The LeNER-Br language modeling dataset allows the finetuning of language models as BERTimbau base and large. Language Portuguese from… See the full description on the dataset page: https://huggingface.co/datasets/pierreguillou/lener_br_finetuning_language_model.text10K<n<100K7 likes51 downloads4y agoHugging Face15rosimeirecosta /c_corpus_br_finetuning_language_model_deberta Dataset Card for "c_corpus_br_finetuning_language_model_deberta" More Information needed text100K<n<1M3 likes51 downloads4y agoHugging Face16luciolrv /lener_br_finetuning_language_model Dataset Card for "lener_br_finetuning_language_model" More Information needed text1K<n<10K0 likes49 downloads3y agoHugging Face17langtech-languagemodeling /piqa_es Dataset Card for PIQA (Spanish Version) Dataset summary This dataset provides the Spanish translation and adaptation of the validation set of PIQA (Physical Interaction: Question Answering). The original dataset was designed to evaluate physical commonsense reasoning in language models through questions about everyday situations. Each example presents a physical goal and two possible solutions, only one of which is correct. This Spanish adaptation enables… See the full description on the dataset page: https://huggingface.co/datasets/langtech-languagemodeling/piqa_es.textquestion-answering1K<n<10K0 likes47 downloads19d agoHugging Face18fair-forward /evals-for-every-language-modelstabularn<1K0 likes43 downloads4mo agoHugging Face19interlinguistic-language-modeling /ilm_detext1M<n<10M0 likes43 downloads6mo agoHugging Face20interlinguistic-language-modeling /ilm_estext1M<n<10M0 likes40 downloads6mo agoHugging Face21langtech-languagemodeling /ALIABOOST-C2textn<1K0 likes39 downloads7mo agoHugging Face22ravikumar1478 /masked_language_modeling_for_Telugu_languagetext10K<n<100K1 likes38 downloads3y agoHugging Face23DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V2 白泽中医药大模型语料库 版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.text10K<n<100K3 likes38 downloads1y agoHugging Face24kgourgou /hugging-face-language-models Data from the configs of the 184 most popular language models on Hugging Face tabularn<1K2 likes35 downloads2y agoHugging Face25heitorefer /repro-how-good-is-post-hoc-watermarking-with-language-model-rephrasing-traces Agent traces Agent sessions published from a Trackio Logbook. text1K<n<10K0 likes35 downloads2mo agoHugging Face26wilmamuller /languagemodelstext1M<n<10M0 likes35 downloads24d agoHugging Face27fineset-io /protein-language-models-papers Protein Language Models Papers — FineSet A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.tabulartext-classificationn<1K0 likes34 downloads4mo agoHugging Face28patrikgerard /uk_language_modeling_v2text10M<n<100M1 likes33 downloads2y agoHugging Face29dairafm05 /2-language-modelstext1M<n<10M0 likes33 downloads20d agoHugging Face30langtech-languagemodeling /aliaboost_when2call_estextn<1K0 likes30 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.