Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes142 downloads4mo agoHugging Face02build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes132 downloads4mo agoHugging Face03build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes121 downloads4mo agoHugging Face04build-small-hackathon /hackathon-advisor-codex-traces Hackathon Advisor Codex Session Traces Real Codex session logs for the Hackathon Advisor project, selected from local Codex rollout JSONL files and redacted before publication. The event stream preserves user requests, assistant messages, tool calls, tool outputs, browser/search events, and minimal session provenance needed to audit how the project was built. Privacy filtering The publisher applied openai/privacy-filter at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.tabulartext-generationn<1K0 likes77 downloads4mo agoHugging Face05build-small-hackathon /lolaby-traces Lolaby — generation traces Pipeline traces from Lolaby, an AI-powered lullaby generator built for the Build Small Hackathon 2026 (Backyard AI track). Each trace is a complete witness of one end-to-end generation: every input the user gave, every model that ran, every prompt and raw output, every timing measurement, and the final audio. Published under CC0 so anyone can study, replay, or remix the pipeline. What's in a trace Each subfolder is one generation. Files:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lolaby-traces.audiotext-to-audion<1K0 likes73 downloads4mo agoHugging Face06build-small-hackathon /AI-Puppet-Theater-Actor-SFT AI Puppet Theater Actor SFT Synthetic supervised fine-tuning data for the Actor agent in AI Puppet Theater. The dataset teaches a small language model to respond to a single puppet-theater beat with one compact JSON object. It is intended for hackathon prototyping, schema following, and local adapter experiments, not as a general storytelling or chat dataset. Schema Each row is chat-style JSONL: { "id": "actor-sft-v0-000001", "source_mix": ["synthetic_v0"… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/AI-Puppet-Theater-Actor-SFT.texttext-generation1K<n<10K1 likes71 downloads4mo agoHugging Face07build-small-hackathon /dota2tuned-data DOTA2Tuned Data This dataset supports the DOTA2Tuned Hugging Face Build Small Hackathon app. It contains compact derived artifacts for Dota 2 draft recommendations, hero meta lookup, build timing summaries, match prediction, retrieval, and supervised fine-tuning examples. Contents sft_examples.jsonl: instruction examples generated from normalized Dota 2 recommendations, patch/stat cards, and app behaviors. Compact Parquet artifacts used by the Space: dim_hero… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dota2tuned-data.tabulartext-generation1M<n<10M0 likes69 downloads4mo agoHugging Face08build-small-hackathon /agenda-parser-models-example-agent-traces Agenda Parser — fine-tuned agent models Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step the model emits a single JSON action {"thought","tool","args"} over two toolkits — meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA, the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the dataset itself (bottom) is a gallery of example traces from the three models. tier base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.texttext-generationn<1K0 likes66 downloads4mo agoHugging Face09somosnlp-hackathon-2023 /Habilidades_Agente_v1 Description Español: Presentamos un conjunto de datos que presenta tres partes principales: 1. Dataset sobre habilidades blandas. 2. Dataset de conversaciones empresariales entre agentes y clientes. 3. Dataset curado de Alpaca en español: Este dataset toma como base el dataset https://huggingface.co/datasets/somosnlp/somos-alpaca-es, y fue curado con la herramienta Argilla, alcanzando 9400 registros curados. Los datos están estructurados en torno a un método que se describe… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/Habilidades_Agente_v1.texttext-generation10K<n<100K22 likes47 downloads3y agoHugging Face10build-small-hackathon /lost-frequency-radio-transmissions Lost Frequency Radio · Transmissions Roughly 786 short, surreal radio transmissions in chat format (system / user / assistant), in Spanish and English, for fine-tuning small models as scriptwriters for parallel-universe radio stations. Built to train the model behind Lost Frequency Radio (Hugging Face Build Small Hackathon 2026). Agent build trace (how it was made, scrubbed and shared): https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-agent-trace… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-transmissions.texttext-generationn<1K1 likes44 downloads4mo agoHugging Face11build-small-hackathon /fabella-traces Fabella Anonymized Agent Traces A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon. The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.tabulartext-generationn<1K0 likes36 downloads4mo agoHugging Face12build-small-hackathon /compliment-forest-sft Compliment Forest SFT Compliment Forest SFT teaches a small language model to turn a (name, situation) pair into a strict JSON forest of grounded encouragement. Each forest contains five distinct creature-strength clearings, a situation-specific line, an agency-oriented reflection, a short first-person spell, and a creature-only image prompt. Dataset Size Train: 1,350 records Validation: 150 records Seed: 42 Language: English Every row contains: name situation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-sft.texttext-generation1K<n<10K0 likes35 downloads4mo agoHugging Face13build-small-hackathon /genregoblin-traces GenreGoblin Agent Trace Examples This dataset contains synthetic, privacy-safe examples of GenreGoblin's visible rewrite pipeline. It is published for the Build Small Hackathon's Sharing is Caring and Best Agent quests. Each JSONL row includes: A plain input message Selected genre, intensity, and use-case Six structured trace stages A synthetic: true marker The trace is intentionally honest. It describes a structured single-agent workflow and does not claim hidden multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/genregoblin-traces.texttext-generationn<1K0 likes33 downloads4mo agoHugging Face14build-small-hackathon /job-search-distill Job Search Distillation Corpus A reasoning-trace SFT corpus for resume-aware job search. Teacher labels (search queries and fit evaluations, with full <think> reasoning preserved) generated by DeepSeek V4 Pro. Four relational configs cover the full pipeline: resumes → search queries → scraped jobs → fit evaluations. Dataset structure Config Contents resume_corpus resume_id, category, resume query_gen_pairings resume_id, teacher reasoning, list of… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/job-search-distill.tabulartext-generation10K<n<100K0 likes31 downloads4mo agoHugging Face15somosnlp-hackathon-2026 /somosnlp-2026-aerospace Dataset Card: Conjunto de Datos Aeroespacial y Cultural Completo Resumen del Dataset Este conjunto de datos ha sido diseñado específicamente para la evaluación cultural, lingüística y de alineación de Modelos de Lenguaje (LLMs) en el ámbito iberoamericano, con un foco especial en la historia aeroespacial, técnica, científica e histórica. Contiene 1.716 interacciones de tipo conversacional (multi-turn) distribuidas en múltiples países de habla hispana y portuguesa.… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/somosnlp-2026-aerospace.texttext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face16build-small-hackathon /slipstream-evm-sft Slipstream: EVM code-action forecasting traces (SFT) Supervised fine-tuning traces for distilling a code-action forecasting agent into small reasoning models. Each example is a full multi-turn trajectory in which a strong teacher forecasts a project's final cost (Estimate at Completion, EAC) and finish period from a mid-flight Earned Value Management (EVM) snapshot, by writing and running Python against a fixed toolset and then calling submit(finish, eac). This is the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/slipstream-evm-sft.texttext-generationn<1K0 likes29 downloads4mo agoHugging Face17build-small-hackathon /nightwave-traces NIGHTWAVE — Open Broadcast Trace A content-only trace of NIGHTWAVE, a 1970s all-night radio station run by a single ~1-billion-parameter model. Each record pairs the exact system prompt the app assembled with the real model output produced by MiniCPM5-1B on a Modal T4 — captured live through the Space's /api/* proxy. Built for the Build Small Hackathon (Thousand Token Wood). 🎙️ Space: https://huggingface.co/spaces/build-small-hackathon/nightwave · ▶ Demo:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/nightwave-traces.texttext-generationn<1K0 likes28 downloads4mo agoHugging Face18build-small-hackathon /tianwen-distill Tianwen Distillation Set A small, quality-filtered instruction dataset that teaches a model to read Chinese BaZi (八字) and I-Ching (六爻) charts in a plain, warm, second-person, anti-doom voice — reframing ominous symbols as growth language and ending with one concrete action. Used to fine-tune tianwen-minicpm5-1b. Size: 58 examples (cleaned from 64) Format: ShareGPT — {"messages": [{"role": "system|user|assistant", "content": ...}]} Teacher model: MiniMax-M2.7-highspeed… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/tianwen-distill.texttext-generationn<1K0 likes28 downloads4mo agoHugging Face19somosnlp-hackathon-2025 /ibero-characters-es Conjunto de datos de personajes de mitos y leyendas iberoamericanos. ⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y mejorar la cobertura de imágenes. 📚 Descripción Dataset de personajes míticos y legendarios de Iberoamérica, diseñado para preservar y promover el patrimonio cultural a través de la inteligencia artificial. 🌟 Motivación e Impacto 📱 Preservación Digital: Conservación del… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-characters-es.imagevideo-text-to-textn<1K0 likes25 downloads1y agoHugging Face20build-small-hackathon /proofkit-distill-qwen0.5b ProofKit distillation dataset ~7,000 chat examples for sequence-level (data) distillation. ProofKit's fine-tuned gpt-oss-20b teacher (visproj/proofkit-gpt-oss-20b-lora) regenerates the assistant turn over the exact prompts from visproj/proofkit-sft; the system + user turns are kept verbatim, so the set stays license-safe (no scraping, no PII). A Qwen 0.5B student is then SFT'd on this to produce visproj/proofkit-distilled-qwen0.5b (and its GGUF), which the ProofKit Space serves… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/proofkit-distill-qwen0.5b.texttext-generation1K<n<10K0 likes25 downloads4mo agoHugging Face21build-small-hackathon /compliment-forest-traces Compliment Forest Linked-Model Traces Sanitized, deterministic traces showing the complete Compliment Forest pipeline: input guard, MiniCPM author draft, MiniCPM critic decision, adaptive clearing selection, FLUX prompt handoff, and progressive completion. The three scenarios are fictional and included directly in scenario records. Runtime identity and situation fields are redacted by the trace recorder. Images are represented by prompt, seed, success status, and model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-traces.texttext-generationn<1K0 likes24 downloads4mo agoHugging Face22build-small-hackathon /PaperProf-traces PaperProf Agent Trace Step-by-step trace of PaperProf, an AI study buddy that turns course PDFs into interactive quiz sessions. What's in this dataset Each row in paperprof_trace.jsonl is one LLM call. Fields: Field Description session_id Groups steps from the same session step Step index within the session (1–4) type question_generation / answer_evaluation / mcq_generation topic Domain of the source chunk input Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.tabularquestion-answeringn<1K0 likes23 downloads4mo agoHugging Face23build-small-hackathon /elysium-training-dataset 🌿 Elysium — Agentic JSON Training Dataset The supervised fine-tuning (SFT) dataset used to train Elysium, a QLoRA fine-tune of openbmb/MiniCPM-V-4.6 that always emits a single valid ElysiumResponse JSON object (schema v1.0.0). Submission to the Build Small Hackathon. Companion model (trained on this dataset): 👉 build-small-hackathon/elysium-MiniCPM-V-4.6-F16-GGUF 📦 Dataset summary Property Value Examples 1,023 File size 6.15 MB Format JSONL (one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/elysium-training-dataset.texttext-generation1K<n<10K0 likes23 downloads4mo agoHugging Face24build-small-hackathon /professor-pip-traces Professor Pip — Open Course-Run Traces Synthetic runtime traces from Professor Pip, a kids (5–10) 3D talking-avatar teacher built for the Build Small Hackathon (Backyard AI). Each trace is one call to Pip's brain — a fine-tuned MiniCPM5-1B teacher LoRA, served as GGUF via llama.cpp on Modal — answering a child's spontaneous "raise-hand" question during a lesson, or gently redirecting an off-topic / not-for-kids prompt. Shared so others can see how a tiny, fine-tuned model holds… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/professor-pip-traces.texttext-generationn<1K0 likes22 downloads4mo agoHugging Face25somosnlp-hackathon-2025 /es-refranes-datasettexttext-generationn<1K0 likes21 downloads1y agoHugging Face26build-small-hackathon /hackathon-advisor-quest-dataset Hackathon Advisor — Quest Classification SFT Dataset Supervised fine-tuning data that teaches MiniCPM5-1B to classify a Build Small Hackathon project against 13 judging dimensions from a two-segment README + app-file prompt, emitting strict JSON with short, source-attributed evidence. Trains the LoRA at build-small-hackathon/hackathon-advisor-quest-minicpm5-lora. Files quest_sft.jsonl — the dataset (one lora_sft_example per line; the viewer split).… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-quest-dataset.tabulartext-generationn<1K0 likes17 downloads4mo agoHugging Face27build-small-hackathon /Kintsugi-Garden-traces Kintsugi Garden Evaluation Traces Paired evaluation traces from Kintsugi Garden — a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model and a four-layer voice/safety architecture; this dataset is what those layers look like under instrumentation. What's in here 114 deterministic runs over the same 19 prompts × 3 trials, evenly split between: baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.tabulartext-generationn<1K0 likes16 downloads4mo agoHugging Face28poolside-laguna-hackathon /umpalumpas Compaction Dataset Based on SWE-Smith This is a sample of a larger dataset. Work in progress. Parquet Layout This repository contains a single Parquet file: File Rows What it contains data/oracle_trajectories.parquet 10 One row per oracle trajectory, preprocessed with linked offline checkpoints, oracle memories, and oracle continuations. The compaction data is represented by linked offline H2 fixed-milestone checkpoints: offline_checkpoint_count… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/umpalumpas.tabulartext-generationn<1K0 likes14 downloads4mo agoHugging Face29build-small-hackathon /velvet-rope-playtest-transcripts Velvet Rope Playtest Transcripts Cleaned playtest transcripts for Velvet Rope, a Build Small Hackathon Gradio game where players talk past whimsical AI gatekeepers by reading moods and discovering each character's soft spot. This dataset is published for the hackathon's sharing-is-caring badge. It contains 341 turn-level rows from 96 local playtest session files. Files data/playtest_transcripts.csv - table-friendly version. data/playtest_transcripts.jsonl - one… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/velvet-rope-playtest-transcripts.tabulartext-generationn<1K0 likes13 downloads4mo agoHugging Face30build-small-hackathon /proofkit-sft ProofKit SFT dataset The supervised fine-tuning set for ProofKit's small models (~7,000 chat examples). Fully synthetic and license-safe — examples are generated deterministically from ProofKit's own templates, demo profiles, and role-knowledge records (data/finetune/build_dataset.py). No scraped prose, no private user data, no model-generated targets. Tasks section_draft, coauthor_draft (draft from rough user answers), section_revision, and strict-JSON… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/proofkit-sft.texttext-generation1K<n<10K0 likes11 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.