Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mohameddalii /coda-llm-data Coda LLM Project & Dataset Repository This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM (Granite-4.2-8B Najdi Sales Agent). Model Repository: mohameddalii/coda-llm Dataset / Code Repository: mohameddalii/coda-llm-data 📁 Repository Structure coda-llm-data/ ├── data/ │ ├── raw/ # Raw generated multi-turn dialogues across domains │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/mohameddalii/coda-llm-data.texttext-generation1K<n<10K0 likes12k downloads34m agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K23 likes9.3k downloads2y agoHugging Face03tokyotech-llm /swallow-math-v2 SwallowMath-v2 Resources 📑 arXiv: Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath-v2 is a large-scale mathematical dataset containing 32 billion tokens, developed as the successor to SwallowMath-v1. Building on the success of v1, this release aims to construct a larger-scale and more permissively licensed corpus to support open and… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math-v2.texttext-generation10M<n<100M35 likes6.8k downloads11mo agoHugging Face04sarahcen /llm-election-data-2024 Data Release for Large-Scale, Longitudinal Survey of Large Language Models (LLMs) During the 2024 US Elections Overview This repository contains the questions asked of and responses given by LLMs during the 2024 US elections, collected for a longitudinal survey conducted from July 23, 2024 to November 12, 2024. The study is described in detail in the paper "Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season" by Sarah H. Cen, Andrew… See the full description on the dataset page: https://huggingface.co/datasets/sarahcen/llm-election-data-2024.text-generation4 likes6.6k downloads10mo agoHugging Face05Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes4.8k downloads3y agoHugging Face06tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B55 likes4k downloads11mo agoHugging Face07LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K37 likes3.4k downloads7mo agoHugging Face08addisonwu05 /llm-polysemy-outputs Polysemy Outputs Raw model generations for the paper "Where did the ambiguity go? Examining how multimodal models interpret polysemous words." Each polysemous word (e.g. bank, bolt, trunk) is presented with no disambiguating context — the prompt is the bare word — and the model's chosen sense is observed over many samples. The same word set is run in two modalities (text-to-image and text generation) and scored by the same judges, so their sense distributions are directly… See the full description on the dataset page: https://huggingface.co/datasets/addisonwu05/llm-polysemy-outputs.text-to-image0 likes2.9k downloads2mo agoHugging Face09epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes2.6k downloads3y agoHugging Face10Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K2 likes2.6k downloads3mo agoHugging Face11LLM-OS-Models /Qwen-Terminal-ToolBench-Processed-Tokenized Qwen Terminal ToolBench Processed Datasets Qwen-family processed/template-applied and selected tokenized terminal datasets. Contents qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.text-generation0 likes2.2k downloads4mo agoHugging Face12elmoghany /Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text Dataset Overview A collection of 27 domains (“topics”) and 3100 question-answer pair. Each topic comes with average 117 QA pairs.Every QA entry comes with: references: one or more source files the answer is extracted from time with each reference comes the starting and ending time the answer is extracted from the reference video_files: the video files where the answer can be found (future) video title & description from metadata.csv File structure You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.question-answering1K<n<10K3 likes2.1k downloads1y agoHugging Face13fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes1.6k downloads3mo agoHugging Face14Open-Style /Open-LLM-Benchmark Open-LLM-Benchmark Dataset Description The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator. License: CC-BY 4.0 Dataset Structure An example of model response files looks as follows: { "question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.texttext-generation100K<n<1M1 likes1.6k downloads2y agoHugging Face15Podtech /llm-jp-corpus-v4-ja_wiki llm-jp-corpus-v4 — ja_wiki Mirror of the ja/ja_wiki sub-corpus of LLM-jp Corpus v4, built by the LLM-jp Corpus Building WG (NII). Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4 Sub-corpus: ja_wiki Files: 6 × jsonl.gz (1.9 GB compressed) Format: one JSON object per line, with a text key and a meta key (document id, URL, and other provenance fields). Directory layout mirrors the upstream repository. License CC BY-SA 3.0 — inherited from the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_wiki.texttext-generation1M<n<10M0 likes1.6k downloads2mo agoHugging Face16Vintage-LLM /EEBO EEBO-TCP (Markdown) Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II. TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are… See the full description on the dataset page: https://huggingface.co/datasets/Vintage-LLM/EEBO.tabulartext-generation10K<n<100K0 likes1.5k downloads12d agoHugging Face17bakrianoo /jabarti-llm-dataset jabarti-llm-dataset Cleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Every pretrain record is a contiguous span of 120-1500 characters with the article title and section headings removed. Provenance is in ds_source. Configs and Splits Config Split Rows Training phase Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.tabulartext-generation1M<n<10M59 likes1.4k downloads17d agoHugging Face18mii-llm /gazzetta-ufficiale Gazzetta Ufficiale 👩🏻‍⚖️⚖️🏛️📜🇮🇹 La Gazzetta Ufficiale della Repubblica Italiana, quale fonte ufficiale di conoscenza delle norme in vigore in Italia e strumento di diffusione, informazione e ufficializzazione di testi legislativi, atti pubblici e privati, è edita dall’Istituto Poligrafico e Zecca dello Stato e pubblicata in collaborazione con il Ministero della Giustizia, il quale provvede alla direzione e redazione della stessa. L'Istituto Poligrafico e Zecca dello Stato… See the full description on the dataset page: https://huggingface.co/datasets/mii-llm/gazzetta-ufficiale.texttext-generation1M<n<10M41 likes1.1k downloads3y agoHugging Face19Exgentic /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.tabulartext-generation1K<n<10K23 likes1.1k downloads4mo agoHugging Face20tokyotech-llm /swallow-math SwallowMath October 21, 2025: Newer versions are available: SwallowCode-v2 and SwallowMath-v2 have been released with improved rewriting pipelines. Resources 🐙 GitHub: Explore the project repository, including pipeline code and prompts at rioyokotalab/swallow-code-math. 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowCode, our companion dataset for code generation. What is it?… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-math.texttext-generation1M<n<10M49 likes1.1k downloads7mo agoHugging Face21LLM-OS-Models /KoHRM-Text-1.4B-sft-lora-data KoHRM-Text-1.4B SFT and LoRA Prepared Data This dataset repo stores curated KoHRM SFT/LoRA subsets in the same tokenized HRM-Text V1Dataset format used by training. It is intended for quick behavior alignment experiments after KoHRM pretraining. Model repo: https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B Code repo: https://github.com/LLM-OS-Models/KoHRM-text Format Each folder is a prepared V1Dataset: <dataset-name>/ metadata.json tokenizer_info.json… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-sft-lora-data.text-generation0 likes1.1k downloads4mo agoHugging Face22secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 889 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 446 targets OSS-Fuzz: 442 targets PoC task support: 889 targets Exploit task support: 447 targets Patch task support: 447 targets Compressed bundle size: 42.41 GiB Vulnerability classes: {'logic_bug': 450, 'memory_vulnerability': 439} Primary languages: {'C': 118… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generationn<1K0 likes1.1k downloads2d agoHugging Face23Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 64 runs, 2,782,350 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes977 downloads14d agoHugging Face24zkolter /llm_speedrun LLM Speedrun token streams Pre-tokenized training artifacts for the LLM speedrun exercises. File Description Tokens tokenizer_50M.bpe JSON-serialized BPE tokenizer — fineweb-edu-10BT.shuffle.bin Shuffled FineWeb-Edu sample/10BT token stream 9,440,023,113 smoltalk.shuffle.bin Shuffled SmolTalk data/all token stream 875,269,408 The .bin files are headerless, little-endian unsigned 16-bit token IDs and can be memory-mapped with NumPy: from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/zkolter/llm_speedrun.text-generation1 likes975 downloads19d agoHugging Face25bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes925 downloads2y agoHugging Face26tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B78 likes865 downloads7mo agoHugging Face27zenml /llmops-database The ZenML LLMOps Database To learn more about ZenML and our open-source MLOps framework, visit zenml.io. Dataset Summary The LLMOps Database is a comprehensive collection of over 500 real-world generative AI implementations that showcases how organizations are successfully deploying Large Language Models (LLMs) in production. The case studies have been carefully curated to focus on technical depth and practical problem-solving, with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.textfeature-extraction1K<n<10K23 likes827 downloads21h agoHugging Face28SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes754 downloads4mo agoHugging Face29llm-jp /scaling-data-constrained-llms Scaling Data-Constrained Language Models with Synthetic Data This repository provides the pre-training corpora used in Scaling Data-Constrained Language Models with Synthetic Data (Findings of EACL 2026). Overview This repository contains multiple corpora designed to study data augmentation strategies for pre-training Japanese LLMs under a data-constrained data setting. Starting from a limited Japanese Web corpus and a larger English Web corpus, we construct three… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/scaling-data-constrained-llms.texttext-generation100M<n<1B5 likes737 downloads7mo agoHugging Face30lenamerkli /LLMinstruct Dataset Card for lenamerkli/LLMinstruct This dataset consists of instruct finetuning data from all of my projects. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/LLMinstruct Uses This dataset is useful for instruct-tuning or fine-tuning large language models. Use Recommendations I recommend to use only the following data for training: all data marked as containing no mistakes the drawback… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/LLMinstruct.texttext-generation10K<n<100K1 likes717 downloads25d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.