Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OptimalScale /ClimbLabClimbLab is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbLab is a filtered 1.2-trillion-token corpus with 20 clusters. Based on Nemotron-CC and SmolLM-Corpus, we employed our proposed CLIMB-clustering to semantically reorganize and filter this combined dataset into 20 distinct clusters, leading to a 1.2-trillion-token high-quality corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we applied two… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbLab.texttext-generation1B<n<10B16 likes5.3k downloads1y agoHugging Face02KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes2.8k downloads16d agoHugging Face03stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes2.2k downloads3mo agoHugging Face04AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K76 likes1.5k downloads3y agoHugging Face05OptimalScale /ClimbMixClimbMix is a high-quality pre-training corpus released by NVIDIA. Here is the description: ClimbMix is a compact yet powerful 400-billion-token dataset designed for efficient pre-training that delivers superior performance under an equal token budget. It was introduced in this paper. We proposed a new algorithm to filter and mix the dataset. First, we grouped the data into 1,000 groups based on topic information. Then we applied two classifiers: one to detect advertisements and another to… See the full description on the dataset page: https://huggingface.co/datasets/OptimalScale/ClimbMix.tabulartext-generation100M<n<1B37 likes1.5k downloads1y agoHugging Face06starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M118 likes1.2k downloads2y agoHugging Face07Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes861 downloads2mo agoHugging Face08LocalDoc /climbmix-40b-az ClimbMix 40B — Azerbaijani A large-scale Azerbaijani text dataset created by translating the English karpathy/climbmix-400b-shuffle dataset into Azerbaijani using Google Translate. Dataset Summary This dataset contains approximately 40 billion tokens of Azerbaijani text, making it one of the largest publicly available Azerbaijani language corpora. It is intended for pretraining and fine-tuning large language models (LLMs) for the Azerbaijani language. Property Value… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/climbmix-40b-az.texttext-generation10M<n<100M2 likes536 downloads7mo agoHugging Face09Yujivus /nanochat-climbmix-arithmetic-base7 nanochat ClimbMix + Arithmetic: base-7 numeral world This is a deterministic base-7 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 7. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base7.texttext-generation10M<n<100M0 likes462 downloads2mo agoHugging Face10Yujivus /nanochat-climbmix-arithmetic-base6 nanochat ClimbMix + Arithmetic: base-6 numeral world This is a deterministic base-6 rendering of Yujivus/nanochat-climbmix-arithmetic-base10. It preserves the exact shard names, row order, document order, arithmetic-document placement, and non-numeric text of the source dataset. Transformation rule Every maximal ASCII digit run matching [0-9]+ is interpreted as a base-10 integer and rendered in base 6. Leading zeros are preserved as a prefix; signs, punctuation… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base6.texttext-generation10M<n<100M0 likes420 downloads2mo agoHugging Face11Yujivus /nanochat-climbmix-170 nanochat ClimbMix: first 170 train shards Convenience mirror of the exact initial ClimbMix slice downloaded by python -m nanochat.dataset -n 170. Contents Training: shard_00000.parquet through shard_00169.parquet Validation: shard_06542.parquet manifest.json: pinned source revision, file list, and byte sizes The Parquet shards are copied without modifying their rows or text. Attribution and provenance nanochat:… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-170.texttext-generation10M<n<100M0 likes373 downloads2mo agoHugging Face12Clinton /Text-to-sql-v1texttext-generation100K<n<1M73 likes355 downloads3y agoHugging Face13Giordanopsouza /clinicalbr ClinicalBr ClinicalBr is the first bilingual (Portuguese–English) clinical-decision benchmark built from 2,892 real Brazilian case reports drawn from 36 open-access medical journals spanning 18 specialties. Every case is provided as a parallel PT/EN pair and supports four evaluation tasks. Please refer to the paper for full details on the tasks, methodology, and limitations. Tasks & metrics Task Config n / lang Metric Diagnosis retrieval diagnosis 2,135… See the full description on the dataset page: https://huggingface.co/datasets/Giordanopsouza/clinicalbr.textquestion-answering10K<n<100K0 likes325 downloads12d agoHugging Face14hugo /protocolos-clinicos-br Protocolos Clínicos BR Paper | Code | Blog post Brazilian Ministry of Health official clinical guidelines (PCDTs and related) plus the synthetic training corpus derived from them, used to adapt LLMs to Brazilian clinical knowledge. This dataset was introduced in the paper "Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines". Configurations default — Original guidelines (raw text) The 178 official Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/hugo/protocolos-clinicos-br.texttext-generation10K<n<100K0 likes323 downloads3mo agoHugging Face15C-lister /ChainSWE ChainSWE ChainSWE is a benchmark of sequential, dependent bug fixes for evaluating coding agents on continuous software maintenance. It contains 100 time-ordered chains (304 bug-fix tasks) mined from six SWE-bench-family datasets across 54 Python repositories; each row is one chain over a single repository sharing one base commit and pre-built Docker image, and its bug_fixes field lists the tasks in chronological order, where each task is a self-contained SWE-bench-style… See the full description on the dataset page: https://huggingface.co/datasets/C-lister/ChainSWE.texttext-generationn<1K1 likes289 downloads3mo agoHugging Face16carosh /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI corpus.… See the full description on the dataset page: https://huggingface.co/datasets/carosh/cli-1m.texttext-generation1M<n<10M1 likes245 downloads5mo agoHugging Face17hiklikai /CLI-Bench CLI-Bench: Benchmarking AI Agents on Command-Line Tool Orchestration Abstract CLI-Bench is an evaluation benchmark for measuring AI agents' ability to learn and use command-line interface (CLI) tools to complete real-world tasks. Unlike existing benchmarks that test general coding ability or narrow tool-use scenarios, CLI-Bench evaluates tool-agnostic CLI orchestration -- the capacity to read tool documentation, plan multi-step workflows, execute commands… See the full description on the dataset page: https://huggingface.co/datasets/hiklikai/CLI-Bench.documenttext-generationn<1K0 likes172 downloads18d agoHugging Face18rayanhk19 /atencion-cliente-moda-es 🛍️ Dataset de Atención al Cliente para E-commerce de Moda — 5.000 Conversaciones Dataset de 5.000 conversaciones sintéticas en español diseñado específicamente para desarrollar, probar y evaluar agentes de IA, chatbots y sistemas de atención al cliente para e-commerce de moda. 📊 Información del dataset 5.000 conversaciones 🇪🇸 Español 👕 Sector moda y e-commerce 🤖 Diseñado para IA y chatbots 📄 Formato JSONL 🧪 Útil para entrenamiento, testing y evaluación… See the full description on the dataset page: https://huggingface.co/datasets/rayanhk19/atencion-cliente-moda-es.texttext-generationn<1K1 likes167 downloads2d agoHugging Face19mkurman /clinical-case-icd10-diagnosis Clinical History -> ICD-10 (Acute / Chronic) — CC BY-enriched 1798 de-identified clinical histories drawn from open-access case reports in PubMed Central, each paired with a single principal-diagnosis label: an ICD-10-CM code, its official descriptor, and an ACUTE/CHRONIC acuity status. input: a de-identified clinical history (presentation only; the diagnosis is removed and no PHI is present). output: {icd10_code, name, status} where status is ACUTE or CHRONIC (how the… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/clinical-case-icd10-diagnosis.texttext-classification1K<n<10K0 likes144 downloads3mo agoHugging Face20Akahsizrr /devin-cli-reasoning-distillation Devin CLI Reasoning Distillation Dataset A distillation dataset built from Devin CLI session traces, containing the model's internal reasoning traces (chain-of-thought / thinking), user prompts, assistant answers, and tool calls. The dataset is formatted to be directly compatible with SFT training pipelines that expect OpenAI-style message lists with a reasoning_content field. Dataset Summary Total rows 2,632 (2,507 train / 125 validation) Rows with… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/devin-cli-reasoning-distillation.tabulartext-generation1K<n<10K1 likes126 downloads1mo agoHugging Face21LocoreMind /qwen3.5-27b-cli-reasoning-3632x Qwen3.5-27B CLI Reasoning 3632x A synthetic reasoning dataset for CLI/terminal command assistance, distilled from Qwen3.5-27B with thinking mode enabled. Each sample contains a realistic user scenario describing a terminal task, paired with the model's reasoning chain (<think>) and a structured JSON answer (command + description). Dataset Summary Source model Qwen3.5-27B (DashScope API) Samples 3,632 Thinking mode Enabled (budget: 4096 tokens)… See the full description on the dataset page: https://huggingface.co/datasets/LocoreMind/qwen3.5-27b-cli-reasoning-3632x.texttext-generation1K<n<10K61 likes108 downloads7mo agoHugging Face22SINAI /ALIA-es-clinical-psychology-dialogues [!WARNING] DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation. Dataset Introduction The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.texttext-generationn<1K0 likes96 downloads3mo agoHugging Face23gavi56 /cli-1m CLI-1M: Industry-Diverse NL→Shell Training Corpus 975,933 natural-language → shell-command pairs · 18 industries · 6 shells · 13 languages · Apache-2.0 from datasets import load_dataset ds = load_dataset("carosh/cli-1m", revision="v1.0", split="train") # 843,461 rows — SFT-ready, license-filtered, quality-gated The most industry-diverse public dataset for NL→shell-command generation. 108× larger than NL2Bash (the previous public benchmark), and the first multilingual CLI… See the full description on the dataset page: https://huggingface.co/datasets/gavi56/cli-1m.texttext-generation1M<n<10M0 likes92 downloads1mo agoHugging Face24aisc-team-a1 /augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.texttext-generation10K<n<100K2 likes88 downloads3y agoHugging Face25Lots-of-LoRAs /task685_mmmlu_answer_generation_clinical_knowledge Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.texttext-generationn<1K0 likes84 downloads2y agoHugging Face26MaximoLopezChenlo /OncoAgent-Clinical-266K 🧬 OncoAgent Clinical Dataset — 266K Curated Multi-Source Oncology Training Dataset AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0 Dataset Description This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks. Composition Source Samples Description PMC-Patients ~100,000 Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/MaximoLopezChenlo/OncoAgent-Clinical-266K.texttext-generation100K<n<1M1 likes82 downloads5mo agoHugging Face27paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K1 likes81 downloads9mo agoHugging Face28skylenage /ClinConsensus ClinConsensus public dataset This public release contains the complete 900-case low-difficulty tier of ClinConsensus. Each de-identified Chinese clinical prompt has metadata and 30 case-specific binary rubric criteria, for 27,000 rubric criteria in total. The public file intentionally excludes reference answers, raw workbooks, source-row mappings, physician identifiers, QC notes, model responses, and judge transcripts. The complete 2,500-case benchmark is not part of this… See the full description on the dataset page: https://huggingface.co/datasets/skylenage/ClinConsensus.texttext-generationn<1K0 likes76 downloads1mo agoHugging Face29TonicAI /synthetic_clinical_conversations Synthetic Clinical Conversations Fully synthetic English clinical conversations (care-coordination calls, telehealth visits, post-discharge check-ins) paired with structured encounter records, built for training and evaluating transcript→JSON extraction models. Generated structure-first with Tonic Fabricate: the structured facts are authored as relational data with controlled vocabularies, the conversation is rendered from those facts, and the extraction target is a… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/synthetic_clinical_conversations.texttext-generation1K<n<10K0 likes73 downloads2mo agoHugging Face30b-mc2 /cli-commands-explained Overview This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.tabulartext-generation10K<n<100K5 likes72 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.