Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dementor-research /dementor-sft-datatext100K<n<1M0 likes231 downloads2mo agoHugging Face02gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes163 downloads4mo agoHugging Face03RabotniKuma /Fast-Math-R1-SFTThis repository contains the First stage SFT dataset as presented in the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning. This dataset is used for the intensive Supervised Fine-Tuning (SFT) phase, crucial for pushing the model's mathematical accuracy. Project GitHub Repository: https://github.com/RabotniKuma/Kaggle-AIMO-Progress-Prize-2-9th-Place-Solution Dataset Construction This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/RabotniKuma/Fast-Math-R1-SFT.texttext-generation1K<n<10K4 likes113 downloads1y agoHugging Face04jamesdborin /Nemotron-SFT-Competitive-Programming-v2-prompt-only Nemotron-SFT-Competitive-Programming-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Competitive-Programming-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Competitive-Programming-v2-prompt-only.tabular100K<n<1M0 likes91 downloads3mo agoHugging Face05rjcg /efemerides-cumana-sft-qa-small Efemérides de Cumaná/Venezuela – Dataset de Preguntas y Respuestas para SFT Descripción general Este dataset contiene un conjunto curado de pares pregunta–respuesta en español, diseñado específicamente para ajuste fino supervisado (Supervised Fine-Tuning, SFT) de modelos de lenguaje de gran tamaño (LLMs). El contenido se centra en la historia de Cumaná, ciudad de Venezuela durante el período de la Guerra de Independencia, con énfasis en la interpretación histórica, el… See the full description on the dataset page: https://huggingface.co/datasets/rjcg/efemerides-cumana-sft-qa-small.textquestion-answeringn<1K0 likes69 downloads9mo agoHugging Face06tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes47 downloads8mo agoHugging Face07jamesdborin /Nemotron-SFT-CUDA-v1-prompt-only Nemotron-SFT-CUDA-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-CUDA-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-CUDA-v1-prompt-only.tabular1K<n<10K0 likes43 downloads3mo agoHugging Face08ChaoticEconomist /Classical-Mechanics-Equations-Dataset_SFT-or-LoRA Classical Mechanics Equations Dataset (SFT / LoRA Ready) A structured dataset of 64 classical mechanics equations from Newtonian, Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning rows across three task types: equation explanation, Q&A, and derivation. Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and equation understanding tasks. Overview Property Value Domain Classical Mechanics (Physics) Total rows 448 Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.texttext-generationn<1K0 likes40 downloads5mo agoHugging Face09HypoAgent /HypoAgent-SFT HypoAgent-SFT: study description → prespecified statistical analysis plan 54,163 supervised fine-tuning pairs that teach a model to read a study registration and write the analysis plan that belongs to it. The input is a study description — objective, design, population, exposure, comparator, outcome, timing, whatever the registry record actually contains. The target is a prespecified hypothesis-testing and statistical-analysis plan built for that design. This is the curated… See the full description on the dataset page: https://huggingface.co/datasets/HypoAgent/HypoAgent-SFT.texttext-generation10K<n<100K0 likes39 downloads2mo agoHugging Face10AlekseyCalvin /LYRICAL_sft_v7 SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations. This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset. Featuring a present (c. late 2025) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/LYRICAL_sft_v7.texttranslation1K<n<10K1 likes38 downloads8mo agoHugging Face11jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes37 downloads3mo agoHugging Face12AlekseyCalvin /LYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations. This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset. Featuring a present (c. late 2025) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/LYRICAL_Mix_SilverAgePoets_Songs_RuVerses_SFT.texttranslation1K<n<10K0 likes36 downloads8mo agoHugging Face13jamesdborin /Nemotron-SFT-ARC-AGI-v1-prompt-only Nemotron-SFT-ARC-AGI-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-ARC-AGI-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-ARC-AGI-v1-prompt-only.tabular100K<n<1M0 likes32 downloads3mo agoHugging Face14keversiom /VStarBench_sft_w_toolcalltextn<1K0 likes30 downloads1y agoHugging Face15saraprice /alpaca_hhh_sft_headlines_2020_2022 Alpaca-HHH-SFT-headlines-2020-2022 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca_hhh_sft_headlines_2020_2022.tabular1K<n<10K0 likes29 downloads2y agoHugging Face16AlekseyCalvin /Lyrical_Ru2En_Poems_Songs_MeterMatched_csv_SFT LYRICAL Russian2English SFT Version: Meaning+Meter-Matched Russian & Soviet Poems + Songs Manually Adapted by a Poet-Translator 1776 rows/items and 2 columns CSV version Manually translated to English by Aleksey Calvin, with a painstaking effort to cross-linguistically reproduce source texts' phrasal/phonetic, rhythmic, metric, syllabic, melodic, and other lyrical and literary features, whilst retaining adequate semantic/significational fidelity. Moreover… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_Ru2En_Poems_Songs_MeterMatched_csv_SFT.texttranslation1K<n<10K0 likes25 downloads10mo agoHugging Face17zelk12 /text_in_number_tulu-3-sft-personas-instruction-following RU Набор данных содержит в себе текст и его представление в виде 610-ти значного числа. Число полоучено при помощи модели.Исходный набор данных: allenai/tulu-3-sft-personas-instruction-following EN The dataset contains text and its representation as a 610-digit number. The number is hollowed out using model.Initial dataset: allenai/tulu-3-sft-personas-instruction-following tabulartext-generation1K<n<10K0 likes24 downloads2y agoHugging Face18Trelis /stanford-NIL-disclosure-sft NIL Policy Data is taken from the Stanford website. The maximum number of tokens (prompt + completion) in a row of data/train.csv is 100 The maximum number of tokens (prompt + completion) in a row of data/test.csv is 89 For educational and non-commercial use only. texttext-generationn<1K0 likes22 downloads3y agoHugging Face19jamesdborin /Nemotron-SFT-Safety-v1-prompt-only Nemotron-SFT-Safety-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v1-prompt-only.tabular10K<n<100K0 likes21 downloads3mo agoHugging Face20jamesdborin /Nemotron-SFT-Multilingual-v1-prompt-only Nemotron-SFT-Multilingual-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v1-prompt-only.tabular1M<n<10M0 likes21 downloads3mo agoHugging Face21jamesdborin /Nemotron-SFT-Multilingual-v2-prompt-only Nemotron-SFT-Multilingual-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Multilingual-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Multilingual-v2-prompt-only.tabular100K<n<1M0 likes21 downloads3mo agoHugging Face22saraprice /alpaca-hhh-sft-headlines-2017-2019 Alpaca-HHH-SFT-headlines-2017-2019 This is an adapted version of a filtered subset of a cleaned version of the Alpaca Dataset released by Stanford. It only contains instances that don't need input and are single-turn. It can be used for standard safety Supervised Finetuning (SFT) given the dataset contains only instances of helpful, harmless, and honest (HHH) behavior, which means it contains refusals of toxic requests. This dataset should in particular be used for SFT safety… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/alpaca-hhh-sft-headlines-2017-2019.tabular1K<n<10K0 likes20 downloads2y agoHugging Face23sbhambr1 /babiqa_for_sft_reasoning_factstext1K<n<10K0 likes20 downloads1y agoHugging Face24jamesdborin /Nemotron-SFT-Agentic-v2-prompt-only Nemotron-SFT-Agentic-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.tabular100K<n<1M0 likes20 downloads3mo agoHugging Face25breadlicker45 /bread-chat-sfttext10K<n<100K0 likes19 downloads1y agoHugging Face26Mahmoud22 /medical-o1-reasoning-SFT-Arabictext10K<n<100K0 likes18 downloads1y agoHugging Face27jamesdborin /Nemotron-SFT-Safety-v2-prompt-only Nemotron-SFT-Safety-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Safety-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Safety-v2-prompt-only.tabular100K<n<1M0 likes18 downloads3mo agoHugging Face28jamesdborin /Nemotron-SFT-SWE-v3-prompt-only Nemotron-SFT-SWE-v3-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-SWE-v3. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-SWE-v3-prompt-only.tabular100K<n<1M0 likes18 downloads3mo agoHugging Face29BRlkl /orchestrator-sft-datatabularn<1K0 likes17 downloads9mo agoHugging Face30niranjanh123 /chatgpt_filtered_sft_traces_context_awaretabularreinforcement-learningn<1K0 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.