Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wayslab /llm-network-study-data LLM-Network-Study-Data Per-request network captures (.pcapng) collected by the LLM-Network-Study benchmark harness (benchmark.py and the per-workload test scripts). Each directory holds one capture file per request, named request_<id>_run<n>_<timestamp>.pcapng. A directory name encodes four dimensions: <capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results Dimension legend Dimension Values Meaning Capture env ethernet Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.tabularn<1K0 likes19k downloads1mo agoHugging Face02open-llm-leaderboard /contentstabular1K<n<10K25 likes14k downloads2y agoHugging Face03OpenCoder-LLM /opc-fineweb-code-corpus OpenCoder Dataset The OpenCoder dataset is composed of the following datasets: opc-sft-stage1: the sft data used for opencoder sft-stage1 opc-sft-stage2: the sft data used for opencoder sft-stage2 opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.tabular100M<n<1B66 likes7k downloads2y agoHugging Face04llm-blender /Unified-FeedbackCollections of pairwise feedback datasets. openai/summarize_from_feedback openai/webgpt_comparisons Dahoas/instruct-synthetic-prompt-responses Anthropic/hh-rlhf lmsys/chatbot_arena_conversations openbmb/UltraFeedback argilla/ultrafeedback-binarized-preferences-cleaned berkeley-nest/Nectar Codes to reproduce the dataset: jdf-prog/UnifiedFeedback Dataset formats { "id": "...", "conv_A": [ { "role": "user", "content": "...", }, { "role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.tabular1M<n<10M18 likes4.1k downloads3y agoHugging Face05tokyotech-llm /swallow-code-v2 SwallowCode-v2 Resources 📑 arXiv: Read our paper for detailed methodology and results at arXiv:2505.02881. 🤗 Sister Dataset: Discover SwallowMath-v2, our companion dataset for mathematical reasoning. 💻 What is it? SwallowCode-v1 was a high-quality Python code dataset generated through an LLM-based rewriting pipeline. However, it had two significant limitations: (1) it was distributed under the Llama 3.3 Community License, and (2) its size was limited to… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code-v2.tabulartext-generation100M<n<1B55 likes4k downloads11mo agoHugging Face06NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Package version recorded in run metadata: 0.2.0 (some runs lack version metadata). Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes3.2k downloads12d agoHugging Face07Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K2 likes2.6k downloads3mo agoHugging Face08LLMDH /otherdocument100K<n<1M0 likes2.1k downloads1y agoHugging Face09LLMDH /OpenScience Open Science Dataset Overview Open Science is a large-scale, permissively licensed text dataset derived from OpenAlex, containing over 100B (105,390,332,599) words. OpenAlex is an open database of scholarly publications, authors, institutions, and research outputs that serves as a comprehensive source for academic literature. Key Features Truly Open: Contains only permissively licensed data suitable for both commercial and non-commercial use Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/LLMDH/OpenScience.tabular1M<n<10M1 likes2k downloads2y agoHugging Face10fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes1.6k downloads3mo agoHugging Face11Vintage-LLM /EEBO EEBO-TCP (Markdown) Early English Books Online, Text Creation Partnership: hand-keyed transcriptions of books printed in England, and English books printed abroad, 1473-1700. Sermons, pamphlets, laws, almanacs, ballads, science, literature. Converted from TCP's XML to Markdown. 60,329 texts, about 1.5 billion words, 25,369 from TCP Phase I and 34,960 from Phase II. TCP keyed each text twice and proofed it to a 99.995% accuracy target. Characters the keyers could not read are… See the full description on the dataset page: https://huggingface.co/datasets/Vintage-LLM/EEBO.tabulartext-generation10K<n<100K0 likes1.5k downloads12d agoHugging Face12llm-jp /leaderboard-contents-v2tabularn<1K1 likes1.5k downloads20d agoHugging Face13bakrianoo /jabarti-llm-dataset jabarti-llm-dataset Cleaned, section-chunked training corpus for a small bilingual LLM (Arabic + English), combining a curated Egyptian-history collection with general Wikipedia coverage from CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Every pretrain record is a contiguous span of 120-1500 characters with the article title and section headings removed. Provenance is in ds_source. Configs and Splits Config Split Rows Training phase Purpose… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset.tabulartext-generation1M<n<10M57 likes1.4k downloads17d agoHugging Face14nishan-chatterjee /llm-bias-detection LLM Bias Detection Evaluation Traces Evaluation data accompanying Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models (arXiv:2609.08637). Licence scope: CC BY 4.0 covers the authors' original documentation, templates, selection/arrangement and author-generated tables. It does not relicense source text or annotations. IBM retains CC BY-SA 3.0; hate-corpus components retain CC BY 4.0, CC0 or MIT as documented in… See the full description on the dataset page: https://huggingface.co/datasets/nishan-chatterjee/llm-bias-detection.tabulartext-classification10M<n<100M0 likes1.4k downloads22d agoHugging Face15johnathansun /llm-cognitive-choice Do LLMs Choose Like Humans? Data for Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making, by Johnathan Sun, Andrei Shleifer, and Yonatan Belinkov. The dataset contains the product choice trials, model responses, stimuli, and ratings used in the paper. The files follow the layout expected by the analysis code. The download is about 3.42 GB and includes 2,240,800 recorded model responses across 51 files. Browse the files in Data Studio. Use the… See the full description on the dataset page: https://huggingface.co/datasets/johnathansun/llm-cognitive-choice.tabular1M<n<10M0 likes1.4k downloads16d agoHugging Face16OpenCoder-LLM /RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode! Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available. RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.tabular100M<n<1B32 likes1.3k downloads2y agoHugging Face17JQL-AI /JQL-LLM-Edu-Annotations 📚 JQL Educational Quality Annotations from LLMs This dataset provides 17,186,606 documents with high-quality LLM annotations for evaluating the educational value of web documents, and serves as a benchmark for training and evaluating multilingual LLM annotators as described in the JQL paper. 📝 Dataset Summary Multilingual document-level quality annotations scored on a 0–5 educational value scale by three state-of-the-art LLMs: Gemma-3-27B-it, Mistral-3.1-24B-it… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/JQL-LLM-Edu-Annotations.tabular10M<n<100M2 likes1.3k downloads1y agoHugging Face18hebrew-llm-leaderboard /chat-resultstabularn<1K0 likes1.2k downloads1mo agoHugging Face19Exgentic /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.tabulartext-generation1K<n<10K23 likes1.1k downloads4mo agoHugging Face20secmlr /llm-fv-security-targets LLM-FV Security Targets This dataset contains 889 independently validated, containerized security-agent targets produced by the ucsb-mlsec/llm-fv pipelines. Contents GitHub Global Security Advisories: 446 targets OSS-Fuzz: 442 targets PoC task support: 889 targets Exploit task support: 447 targets Patch task support: 447 targets Compressed bundle size: 42.41 GiB Vulnerability classes: {'logic_bug': 450, 'memory_vulnerability': 439} Primary languages: {'C': 118… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/llm-fv-security-targets.tabulartext-generationn<1K0 likes1.1k downloads2d agoHugging Face21InterstellarCG /frontier-synth-dialogue-llmtabular10K<n<100K0 likes988 downloads5mo agoHugging Face22Stereotypes-in-LLMs /hiring-bias-mitigation-responses Hiring-bias mitigation — model responses Every response produced in the mitigation study of LLM hiring decisions: 64 runs, 2,782,350 responses, from 5 open-weight models in English and Ukrainian, at baseline and under each mitigation family (baseline, embedding, prompt, scrub, sft). Each run is one subset. All released artifacts: the Hiring Bias Mitigation collection. Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data. Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.tabulartext-generation1M<n<10M0 likes977 downloads13d agoHugging Face23Necent /llm-jailbreak-prompt-injection-datasetgated LLM Jailbreak & Prompt-Injection Dataset A unified safety dataset combining 30+ public sources for training LLM guardrails, content moderation classifiers, and response-safety filters. Schema (orthogonal multi-label, WildGuard-style) Instead of a single binary is_dangerous, every example carries four orthogonal labels matching the structure used by AI2 WildGuard, IBM Granite Guardian, and Azure Prompt Shields: Column Type Description prompt str The user/attack… See the full description on the dataset page: https://huggingface.co/datasets/Necent/llm-jailbreak-prompt-injection-dataset.tabulartext-classification1M<n<10M56 likes933 downloads6mo agoHugging Face24zhenghaozhu /llm-hkmmlu-leaderboard-requeststabularn<1K0 likes874 downloads1y agoHugging Face25fromthesky /pldr-llm-training-dynamics-data PLDR-LLM Training Dynamics Data Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden. Monograph: Hugging Face Paper Page. Scientific code and readers: GitHub repository. Numerical evidence: Hugging Face dataset. Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden. Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.tabularothern<1K0 likes868 downloads4d agoHugging Face26tokyotech-llm /swallow-code SwallowCode Notice May 21, 2025: We have deleted ablation/exp1-the-stack-v2-train-smol-ids-python because it was flagged as potentially containing unsafe data collected from the Python subset of https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids. However, since this dataset can be reconstructed from the-stack-v2-train-smol-ids, there is no issue in terms of reproducibility. May 21, 2025: ClamAV has flagged “Win.Trojan.MSShellcode-88” in… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/swallow-code.tabulartext-generation100M<n<1B78 likes865 downloads7mo agoHugging Face27FT-LLM-2026-RAMEN /droid_1.0.1tabular10M<n<100M0 likes800 downloads9mo agoHugging Face28FilipT /llm-multitudes LLMs Contain Multitudes Pairwise preference and utility judgements from 5 LLMs under 5 deployment contexts. This dataset is the full set of model judgements behind the anonymous NeurIPS 2026 submission "LLMs Contain Multitudes: How Deployment Context Reshapes Model-level Preferences and Values". It captures how the same five LLMs choose between the same options when only the surrounding deployment context (writing a Reddit post, a news article, a school essay, a vlog script, or… See the full description on the dataset page: https://huggingface.co/datasets/FilipT/llm-multitudes.tabulartext-classification1M<n<10M0 likes768 downloads4mo agoHugging Face29Kevynf /llm-graph-poisoning-data Generation-Time Poisoning of LLM-Generated Social Networks This dataset contains synthetic personas, LLM-generated social graphs, cached text embeddings, and evaluation metrics for clean generation and three generation-time attack families. All names and profiles are synthetic and do not represent real people. Dataset variants Variant Nodes Generator Graph seeds per condition Attack rates p50 50 Qwen3-Max 10 10%, 20%, 30%, 40%, 50% p200 200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.tabulargraph-ml100K<n<1M0 likes761 downloads2mo agoHugging Face30SPAISS6F1 /spai-ss6-llm-1b-thai-corpus Thai Medical And Health Corpus Thai public medical and health web corpus collected for research and LLM dataset experimentation, with optional imported Thai medical/health datasets from Hugging Face stored as separate configs. Public Web Corpus Config: default Split: train Records: 3660 deduplicated articles Columns: 16 Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T17:41:38.787978+00:00 Source And Method The… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-llm-1b-thai-corpus.tabulartext-generation10M<n<100M0 likes754 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.