Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /dclm-baseline-1.0 DCLM-baseline DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks. Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime. Model Params Tokens Open dataset? CORE MMLU EXTENDED Open weights, closed datasets Llama2 7B 2T ✗ 49.2 45.8 34.1 DeepSeek 7B 2T ✗ 50.7 48.5 35.3 Mistral-0.3 7B ? ✗ 57.0 62.7 45.1 QWEN-2 7B ? ✗ 57.5 71.9 50.5 Llama3 8B 15T ✗… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0.tabular1B<n<10B320 likes482k downloads2y agoHugging Face02survivi /baseline_dapo_final21K<n<10K0 likes4.7k downloads1y agoHugging Face03survivi /baseline_dapo_positive_only1K<n<10K0 likes4.2k downloads1y agoHugging Face04IvanHU /baseline_dapo_positive_only10K<n<100K0 likes3.3k downloads1y agoHugging Face05IvanHU /baseline_dapo_final210K<n<100K0 likes2.8k downloads1y agoHugging Face06Sakib323 /GUI_BASED_PLATFORMtext100K<n<1M0 likes1k downloads7mo agoHugging Face07nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes919 downloads2y agoHugging Face08TREC-AToMiC /AToMiC-Baselines AToMiC Prebuilt Indexes Example Usage: Reproduction Toolkits: https://github.com/TREC-AToMiC/AToMiC/tree/main/examples/dense_retriever_baselines # Skip the encode and index steps, search with the prebuilt indexes and topics directly python search.py \ --topics topics/openai.clip-vit-base-patch32.text.validation \ --index indexes/openai.clip-vit-base-patch32.image.faiss.flat \ --hits 1000 \ --output… See the full description on the dataset page: https://huggingface.co/datasets/TREC-AToMiC/AToMiC-Baselines.textn<1K1 likes915 downloads3y agoHugging Face09yifanzhang114 /MME-RealWorld-Base64 MME-RealWorld Dataset This dataset contains multiple JSON files split into chunks. It includes information such as questions, images encoded in base64, and other related metadata. Usage You can load the dataset using the datasets library: from datasets import load_dataset dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld') dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld-CN') ## the image can be… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/MME-RealWorld-Base64.text10K<n<100K1 likes873 downloads2y agoHugging Face10birdsql /livesqlbench-base-lite-sqlite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 LiveSQLBench Website • 🌐 BIRD-INTERACT Project Page • 📄 Paper • 💻 LiveSQLBench GitHub • 💻 BIRD-INTERACT GitHub Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on complex, real-world… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite-sqlite.texttable-question-answeringn<1K4 likes665 downloads5mo agoHugging Face11duyle2408 /visdrone2019-mmdet-baselines-1536-seedstabularn<1K0 likes649 downloads13d agoHugging Face12birdsql /livesqlbench-base-full-v1 🚀 LiveSQLBench-Base-Full-v1 A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 Website/Leaderboard • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Lite • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-full-v1.textn<1K3 likes619 downloads4mo agoHugging Face13LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes584 downloads2y agoHugging Face14KingNish /reasoning-base-20k Dataset Card for Reasoning Base 20k Dataset Details Dataset Description This dataset is designed to train a reasoning model. That can think through complex problems before providing a response, similar to how a human would. The dataset includes a wide range of problems from various domains (science, coding, math, etc.), each with a detailed chain of thought (COT) and the correct answer. The goal is to enable the model to learn and refine its reasoning process… See the full description on the dataset page: https://huggingface.co/datasets/KingNish/reasoning-base-20k.texttext-generation10K<n<100K232 likes448 downloads1y agoHugging Face15birdsql /livesqlbench-base-lite 🚀 LiveSQLBench-Base-Lite A dynamic, contamination‑free benchmark for evaluating LLMs on complex, real‑world text‑to‑SQL tasks. 🌐 Website • 📄 Paper (coming soon) • 💻 GitHub • 🗄️ LiveSQLBench-Base-Full-v1 • 🗄️ LiveSQLBench-Large-v1 • 🗄️ Bird-Interact (ICLR 2026 Oral) Maintained by the 🦜 BIRD Team @ HKU & ☁️ Google Cloud 📊 LiveSQLBench Overview LiveSQLBench (BIRD-SQL Pro v0.5) is a contamination-free, continuously evolving benchmark designed to evaluate LLMs on… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/livesqlbench-base-lite.textn<1K5 likes416 downloads7mo agoHugging Face16baseten-admin /gpt-oss120b-generated-perfectblendtext100K<n<1M1 likes412 downloads1y agoHugging Face17baseten-admin /magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedtext1M<n<10M0 likes405 downloads1y agoHugging Face18friedrichor /Unite-Base-Retrieval-Train Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Statistics Accessing Images and Videos 2025-06-19: We've updated the compressed archives for all image and video files to enable faster extraction.If you've already downloaded the previous files, there's no need to redownload them — the content remains exactly the same. The only difference lies in the compression method, which now allows for quicker… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/Unite-Base-Retrieval-Train.textfeature-extraction1M<n<10M0 likes403 downloads1y agoHugging Face19RX5950XT /silicon-based-girlfriend-v2-dataset 矽基女友 v2 · 繁中角色扮演合成語料 繁體中文(臺灣)角色扮演的合成對話語料,2,109 筆、38,460 輪、角色輪合計 1,350 萬字元。 訓練出來的模型見 RX5950XT/silicon-based-girlfriend-v2-GGUF。 ⚠️ 全部是模型合成的資料,不是真人對話。 內容包含成人向角色扮演,不適合未成年人。 僅供研究用途。所有角色皆為虛構成年人。 內容 檔案 內容 sharegpt_dataset.json 2,109 筆多輪對話,ShareGPT 格式(id / system / conversations) grpo_prompts.json 648 題 GRPO 用的提示,與 SFT 語料零重疊 holdout_ids.json 100 筆 holdout ID 清單,這些已從 SFT 訓練集排除,供驗收用 general_probes.json 64 題通用能力探針(5 類),用來檢測微調後有無退化… See the full description on the dataset page: https://huggingface.co/datasets/RX5950XT/silicon-based-girlfriend-v2-dataset.texttext-generation1K<n<10K1 likes346 downloads23d agoHugging Face20BananaMind /BananaMind-Base-Bench-1.1gated BananaMind Base Bench 1.1 BananaMind Base Bench 1.1 is an English text-completion benchmark for base causal language models. It contains 350 individually authored examples across seven categories and reports one fixed-scale Overall Elo score. This is not an instruction-following benchmark. Models receive plain text followed by four possible continuations. The official runner selects the continuation with the highest mean conditional token log-probability. It does not use a chat… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1.tabulartext-generationn<1K11 likes325 downloads2mo agoHugging Face21casinca /PUBMED_title_abstracts_2019_baseline PUBMED Title and Abstracts 2019 Baseline This dataset contains the titles and abstracts from biomedical publications on PubMed, extracted from the 2019 baseline.It has been uploaded to Hugging Face as it is no longer hosted on the Eye and may help students w.r.t HF NLP Course CH5-4 Big data Context More infos from the HF course here. "The Pile" (825 GB) is an English text corpus created by EleutherAI for training large-scale language models. It includes a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/casinca/PUBMED_title_abstracts_2019_baseline.text1M<n<10M10 likes243 downloads2y agoHugging Face22BaseIntelligence /deepagent DeepAgent Hard, Docker-verifiable software-engineering benchmarks from real merged PRs DeepAgent ships real_pr Harbor hardness packs: live-mined multi-file pull requests, clone@SHA agent images, held-out verifier tests, and Docker dual-truth (solution reward = 1, null reward = 0). Primary product work runs through the deepagent CLI in the GitHub monorepo. Surface Ref Role HF stable pin this dataset revision main Current product on Hub (N=9) HF automation… See the full description on the dataset page: https://huggingface.co/datasets/BaseIntelligence/deepagent.tabulartext-generationn<1K0 likes243 downloads3mo agoHugging Face23W-61 /hh-harmless-base-qwen3-8b-margin-dpo-margin-logstabular1K<n<10K0 likes226 downloads7mo agoHugging Face24baseten-admin /perfectblend-Qwen3-235B-A22B-Instruct-2507-FP8-generatedtext100K<n<1M1 likes221 downloads1y agoHugging Face25raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes207 downloads3y agoHugging Face26nyu-dice-lab /lm-eval-results-bobofrut-ladybird-base-7B-v8-private Dataset Card for Evaluation run of bobofrut/ladybird-base-7B-v8 Dataset automatically created during the evaluation run of model bobofrut/ladybird-base-7B-v8 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bobofrut-ladybird-base-7B-v8-private.tabular100K<n<1M0 likes179 downloads2y agoHugging Face27baseten-admin /gpt-oss120b-generated-magpie-1m-v0.1text100K<n<1M2 likes179 downloads1y agoHugging Face28basedlsg /codex-assistant-rollouts basedlsg/codex-assistant-rollouts Real-world agentic interaction logs from Codex rollouts, documenting debugging and coding trajectories. text10K<n<100K0 likes178 downloads7mo agoHugging Face29AIML-TUDA /QA-base QA Base Data Normalized and paraphrased splits of 21 standard NLP benchmarks in English, German, French, Spanish, and Italian, intended for base model pretraining. Generation English: paraphrased with Qwen3.5-27B-FP8 (April 2026) German: translated and refined with Qwen3.5-27B-FP8 (April 2026) French: translated and refined with Qwen3.5-27B-FP8 (May 2026) Spanish: translated and refined with Qwen3.5-27B-FP8 (May 2026) Italian: translated and refined with… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/QA-base.textquestion-answering1M<n<10M0 likes176 downloads3mo agoHugging Face30xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K10 likes170 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.