Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01prompt-agnostic-language-models /pal-results0 likes579 downloads3mo agoHugging Face02CCB /cis5300-language-models CIS 5300 Language Models Dataset Dataset for Homework 3 of CIS 5300 (Natural Language Processing) at Penn. Cities config Country-of-origin classification over short city-name strings, drawn from nine countries (Afghanistan, China, Germany, Finland, France, India, Iran, Pakistan, South Africa). from datasets import load_dataset cities = load_dataset("CCB/cis5300-language-models", "cities") Split Rows Has labels? train 12,392 yes validation 1,548 yes test 1… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-language-models.text10K<n<100K0 likes475 downloads5mo agoHugging Face03beatsprom /multimodal-vision-language-video-models-2026 👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition) A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators. Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.tabularfeature-extractionn<1K8 likes169 downloads2mo agoHugging Face04Setloop /What-Gradients-Add-to-Text-Leakage-in-Split-Language-Models-Counted-per-Token-and-per-Documentgated Paper A: reproduction and peer-review release This dataset contains the final paper, the evidence used for its three experiments, and the historical material explicitly discussed in the paper. Final paper The current branch includes the 2 October 2026 punctuation revision of the 26-page manuscript. Its scientific content and experimental evidence are unchanged. Tag v1.0 retains the frozen release and its original manuscript; the updated PDF and source package are… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/What-Gradients-Add-to-Text-Leakage-in-Split-Language-Models-Counted-per-Token-and-per-Document.3 likes167 downloads8d agoHugging Face05hassan-wajid /Spatial-Blind-Spots-in-Vision-Language-Modelslicense: mit model_evaluated: name: Qwen3-VL-2B-Instruct url: https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct evaluation_notebook: https://www.kaggle.com/code/wajidhassanmoosa/blind-spot-qwen3-2b evaluation_setup: | The model evaluated in this study is Qwen3-VL-2B-Instruct. Evaluation was conducted using the Hugging Face Transformers library with automatic device mapping (device_map="auto") and "bfloat16" dtype selection. For each example: The image was provided as part of a… See the full description on the dataset page: https://huggingface.co/datasets/hassan-wajid/Spatial-Blind-Spots-in-Vision-Language-Models.imagen<1K7 likes148 downloads7mo agoHugging Face06astha /languagemodelsforRNNdecompositionThis repository is for the paper "Decomposing a Recurrent Neural Network into Modules for Enabling Reusability and Replacement". To use the data, there are two directories: language datasets: Contains the necessary Tatoeba files used for the experiments. We have experimented with 4 languages(English, French, Italian and German). language_models: Contains all trained language models and scripts to train them. It's organized in this way: language_models/{X}: contains language models for X… See the full description on the dataset page: https://huggingface.co/datasets/astha/languagemodelsforRNNdecomposition.text100K<n<1M1 likes115 downloads4y agoHugging Face07robotamski /language_models_lab_2text1M<n<10M0 likes93 downloads22d agoHugging Face08abidlabs /repro-how-much-can-language-models-memorize-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes90 downloads3mo agoHugging Face09DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V3 白泽中医药大模型语料库 版本:3.0语料数量:157,438 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 157,438 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V3.2 likes81 downloads1y agoHugging Face10linneripe /language_modelstext1M<n<10M1 likes81 downloads24d agoHugging Face11mznaser /Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset Data for the paper Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173 It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result. If you use the data, please cite the paper (BibTeX under Citation). The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.tabulartext-classification10K<n<100K0 likes52 downloads19d agoHugging Face12fair-forward /evals-for-every-language-modelstabularn<1K0 likes43 downloads4mo agoHugging Face13DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V2 白泽中医药大模型语料库 版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.text10K<n<100K3 likes38 downloads1y agoHugging Face14kgourgou /hugging-face-language-models Data from the configs of the 184 most popular language models on Hugging Face tabularn<1K2 likes35 downloads2y agoHugging Face15wilmamuller /languagemodelstext1M<n<10M0 likes35 downloads24d agoHugging Face16fineset-io /protein-language-models-papers Protein Language Models Papers — FineSet A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.tabulartext-classificationn<1K0 likes34 downloads4mo agoHugging Face17dairafm05 /2-language-modelstext1M<n<10M0 likes33 downloads20d agoHugging Face18mtjikuzu /small-language-models-youtube-lesson mtjikuzu/small-language-models-youtube-lesson Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset("mtjikuzu/small-language-models-youtube-lesson") 0 likes31 downloads5mo agoHugging Face19Filippashuggingface /language_modelstext1M<n<10M0 likes30 downloads8d agoHugging Face20lovisadi /Lab-2-Language-Modelstext1M<n<10M0 likes28 downloads16d agoHugging Face21sati-01 /language_modelstext1M<n<10M0 likes28 downloads16d agoHugging Face22SeanWu25 /NEJM-AI_Benchmarking_Medical_Language_Models A Comparative Study of Open-Source Large Language Models Dataset Overview Welcome to the dataset repository for our paper, "A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology." The preprint of the paper can be accessed here. Files This repository contains two key files: NEJM_All_Questions_And_Answers.csv: This file includes all the questions and corresponding answers used in the study.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWu25/NEJM-AI_Benchmarking_Medical_Language_Models.n<1K10 likes27 downloads3y agoHugging Face23eval-aware /Large-Language-Models-Often-Know-When-They-Are-Being-Evaluatedgated Dataset Card for Evaluation Awareness Benchmark Dataset Summary This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. The dataset contains 976 conversational transcripts with rich metadata, including: True evaluation transcripts from prompt-injection tests, red-teaming tasks, and coding challenges Organic/real transcripts from actual user queries, scraped chats, and… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated.tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face24el2303vi-s /language-detection-models0 likes24 downloads10d agoHugging Face25harpreetsahota /Instruction-Following-Evaluation-for-Large-Language-Models Instruction-Following Evaluation Dataset 📜 Overview This dataset, specifically designed for the evaluation of large language models in instruction-following tasks, is directly inspired by the methodologies and experiments described in the paper titled "Instruction-Following Evaluation for Large Language Models". The dataset's creation and availability on HuggingFace are aimed at enhancing research and application in the field of natural language understanding… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/Instruction-Following-Evaluation-for-Large-Language-Models.textn<1K7 likes23 downloads3y agoHugging Face26zxliu /ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models" text100K<n<1M3 likes22 downloads2y agoHugging Face27DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V1 白泽中医药大模型语料库 版本:1.0语料数量:4,735 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 4,735 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V1.text1K<n<10K1 likes17 downloads1y agoHugging Face28TahaGorji /Data-for-training-language-models.-Structuredtext1K<n<10K0 likes12 downloads1y agoHugging Face29Raiadarsh12345 /Scaling_Language_Models0 likes4 downloads8mo agoHugging Face30french-datasets /astha-languagemodelsforRNNdecompositionCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/astha/languagemodelsforRNNdecomposition. 0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.