Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02BAAI /Infinity-Instructgated Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.tabulartext-generation10M<n<100M765 likes1.9k downloads10mo agoHugging Face03bcb-instruct /bcb_datatabular10M<n<100M0 likes1.6k downloads1y agoHugging Face04ENSEONG /full-math-private-n256-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes1.5k downloads7mo agoHugging Face05BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face06ENSEONG /full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes1k downloads5mo agoHugging Face07OALL /details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.tabular100K<n<1M0 likes1k downloads2y agoHugging Face08RUC-DataLab /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.tabular10K<n<100K76 likes965 downloads1y agoHugging Face09ENSEONG /preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes870 downloads5mo agoHugging Face10matlok /python-text-copilot-training-instruct-ai-research-2024-02-03 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Agora Open Source AI Research Lab: Agora GitHub Organization Agora Hugging Face This dataset is the 2024-02-03 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-03.tabulartext-generation1K<n<10K1 likes857 downloads3y agoHugging Face11sujet-ai /Sujet-Finance-Instruct-177k Sujet Finance Dataset Overview The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.tabulartext-generation100K<n<1M86 likes847 downloads3y agoHugging Face12Vanessasml /cybersecurity_32k_instruction_input_output Dataset Card The dataset Q&As are focused on identification of cyber threats, and text classification under the NIST taxonomy and ITC EBA IT risk classes Dataset Details Dataset Description This dataset includes a mix of public reports and news and aims to be used for cyber security risk model training. It includes 32k examples with instruction, input and output. The latter is the output from GPT. Curated by: [Vanessa Lopes] Language [EN] Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vanessasml/cybersecurity_32k_instruction_input_output.tabular10K<n<100K20 likes838 downloads2y agoHugging Face13bcb-instruct /datatabular1M<n<10M0 likes733 downloads1y agoHugging Face14ENSEONG /full-math-private-Qwen3-4B-Instruct-2507-bontabular100K<n<1M0 likes679 downloads7mo agoHugging Face15mlfoundations-dev /Llama-3.1-8B-Instruct_eval_5554 mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 4.7 15.8 43.2 44.7 14.1 25.8 13.1 2.1 6.7 17.0 0.3 0.3 8.9 AIME24 Average Accuracy: 4.67% ± 0.84% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554.tabular10K<n<100K0 likes554 downloads1y agoHugging Face16nvidia /Nemotron-RL-Instruction-Following-MultiTurnChat-v1 Dataset Description: The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tabular1K<n<10K6 likes512 downloads12d agoHugging Face17ENSEONG /full-math-private-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes493 downloads7mo agoHugging Face18OALL /details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.tabular100K<n<1M0 likes486 downloads2y agoHugging Face19JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes461 downloads25d agoHugging Face20ENSEONG /full-math-private-n256-flip719-Qwen3-4B-Instruct-2507-bontabular10K<n<100K0 likes449 downloads15h agoHugging Face21proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes435 downloads6mo agoHugging Face22orionweller /dolma_20bn_no_instructtabular10M<n<100M0 likes431 downloads2y agoHugging Face23matlok /python-text-copilot-training-instruct Python Copilot Instructions on How to Code using Alpaca and Yaml This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details Each row contains python code, either a class method or a global function, imported modules, base classes (if any), exceptions (ordered based off the code), returns (ordered based off the code), arguments (ordered based off the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct.tabulartext-generation100K<n<1M0 likes417 downloads3y agoHugging Face24allenai /Dolci-Instruct-RL Dolci-Instruct-RL Dataset Summary Dolci-Instruct-RL is the reinforcement learning dataset used to train the Olmo-3-7B-Instruct model.It contains 169,964 prompts spanning: Math Code Precise Instruction Following General Chat The dataset aggregates multiple curated sources, applies extensive filtering, and produces a unified RL-ready prompt set. Dataset Composition Total Samples: 169,964 Original Dataset Contribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-RL.tabular100K<n<1M15 likes417 downloads9mo agoHugging Face25daruokta /t5gemma2-indonesia-instruct-v1 T5Gemma-2 Indonesian Instruct — Mono-Repo Satu repositori dataset HF untuk seluruh data pelatihan T5-Gemma-2 bahasa Indonesia. Diorganisasi per fungsi (fondasi → spesifik → preferensi) dengan folder/subfolder, setiap config = folder dan berisi split train + validation (80:20) di level percakapan. Struktur (by fungsi) t5gemma2-indonesia-instruct-v1/ ├── README.md ├── manifest.json ├── chat_idx_map.json ├── foundation/ ← FASE 1 · fondasi Bahasa… See the full description on the dataset page: https://huggingface.co/datasets/daruokta/t5gemma2-indonesia-instruct-v1.imagetext-generation100K<n<1M0 likes411 downloads17d agoHugging Face26matlok /python-text-copilot-training-instruct-ai-research-2024-02-11 Python Copilot Instructions on How to Code using Alpaca and Yaml Training and test datasets for building coding multimodal models that understand how to use the open source GitHub projects for the Autogen and multimodal Qwen AI project: Qwen Qwen Agent Qwen VL Chat Qwen Audio This dataset is the 2024-02-11 update for the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset. Details… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-text-copilot-training-instruct-ai-research-2024-02-11.tabulartext-generationn<1K0 likes403 downloads3y agoHugging Face27ENSEONG /full-aime_2024-n256-ctx8k-Qwen3-4B-Instruct-2507-bontabular1K<n<10K0 likes378 downloads9d agoHugging Face28ENSEONG /full-aime_2023-n256-ctx8k-Qwen3-4B-Instruct-2507-bontabular1K<n<10K0 likes364 downloads9d agoHugging Face29GG-samrt /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.tabular10K<n<100K0 likes363 downloads8mo agoHugging Face30juiceb0xc0de /smollm2-135m-instruct-SAE Layer EV Mean L0 Recon Loss Dead % 0 0.9480 48.74 0.2074 0.0 1 0.9599 43.65 0.3298 0.0 2 0.9631 46.81 0.5021 0.0 3 0.9508 46.56 0.7462 0.0 4 0.9463 46.23 0.8936 0.0 5 0.9350 47.57 1.1605 0.0 6 0.9306 48.44 1.3838 0.0 7 0.9318 49.51 1.5446 0.0 8 0.9432 46.52 1.6598 0.0 9 0.9373 47.15 2.0706 0.0 10 0.9348 45.53 2.2983 0.0 11 0.9905 48.58 5.8113 0.0 12 0.9901 48.42 6.1039 0.0 13 0.9891 46.15 6.9692 0.0 14 0.9884 44.76 7.1844 0.0 15 0.9863 47.63 8.6521 0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.tabularfeature-extractionn<1K0 likes352 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.