datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
synthetickirana-detective-build-traces
Kirana Detective — Claude Code Build Sessions
Raw Claude Code (claude-sonnet-4-6) session traces recorded while building
Kirana Detective AI for the HuggingFace Build Small Hackathon 2026.
Each .jsonl file is one coding session. Together they cover the entire
build — from first commit to final submission.
What's Inside
Sessions
Agent
Coverage
11 JSONL files
Claude Code (Sonnet 4.6)
Full project build
Sessions include
Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.readability-es-hackathon-pln-public
Dataset Card for [readability-es-sentences]
Dataset Description
Compilation of short Spanish articles for readability assessment.
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.pit-wall-chaos-tracesCodex agent traces for Pit Wall Chaos, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/pit-wall-chaos
clue-vibes-tracesCodex agent traces for Clue Vibes, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/clue-vibes
MatchWise-agent-tracepakistan-notice-helper-traces
NoticeCheck Privacy-Safe Traces
Purpose
This dataset contains compact, deterministic metadata about NoticeCheck
message-review requests. It does not contain hidden model reasoning or
autonomous-agent trajectories.
The hosted application uses MiniCPM5-1B through Transformers on Hugging Face
ZeroGPU, with NVIDIA Nemotron-Parse v1.2 for supported screenshots. The same
pipeline can run locally on an NVIDIA GPU with Docker Compose. Creating a trace
never makes an… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/pakistan-notice-helper-traces.blood-test-explainer-traces
Blood Test Explainer - agent traces
Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns.
Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.AI-Puppet-Theater-Actor-SFT
AI Puppet Theater Actor SFT
Synthetic supervised fine-tuning data for the Actor agent in AI Puppet Theater.
The dataset teaches a small language model to respond to a single puppet-theater beat with one compact JSON object. It is intended for hackathon prototyping, schema following, and local adapter experiments, not as a general storytelling or chat dataset.
Schema
Each row is chat-style JSONL:
{
"id": "actor-sft-v0-000001",
"source_mix": ["synthetic_v0"… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/AI-Puppet-Theater-Actor-SFT.TinyNarrator-agent-tracesSpaces link: https://huggingface.co/spaces/build-small-hackathon/TinyNarrator
nli-esannotations_creators:
crowdsourced
other
language_creators:
other
crowdsourced
languages:
es
licenses:
cc-by-sa-4.0
multilinguality:
monolingual
pretty_name: ESnli
size_categories:
unknown
source_datasets:
extended|snli
extended|xnli
extended|multi_nli
task_categories:
text-classification
task_ids:
natural-language-inference
Dataset Card for nli-es
Dataset Summary
A Spanish Natural Language Inference dataset put together from the sources:
the Spanish slice of the XNLI… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/nli-es.kicky-ai-codex-trace
Kicky AI - Codex agent trace (sanitized)
A redacted OpenAI Codex CLI session trace from building
Kicky AI for the Build Small
Hackathon - shared for the Sharing is Caring badge so others can see how the build went.
Format: Codex CLI JSONL session log (each record = {payload, timestamp, type}).
All secrets removed (HF / Modal / Roboflow tokens, shared secrets, emails) - verified 0 leaks.
Blog write-up: https://dcrey7.substack.com/p/world-fut-coach
dod-agent-traces
DOD Agent Traces
This dataset stores lightweight agent traces generated by DOD - Deploy or Draw, a multiplayer UNO-style game built for the Build Small Hackathon.
The traces are meant to show how the AI parts of the game participate in live gameplay:
Nemotron bot turns: the AI opponent receives the current board state and chooses whether to play a card or draw.
IT Director reactions: the LLM generates a short contextual reaction to the card that was just played and the active… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/dod-agent-traces.agenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.job-search-assistant-agent-tracereadability-es-caes
Dataset Card for [readability-es-caes]
Dataset Description
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
CAES corpus (Martínez et al., 2019): the "Corpus de Aprendices del Español" is a collection of texts produced by Spanish L2 learners from Spanish learning centers and universities. These text are produced by students… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-caes.codeflow-agent-traces
CodeFlow — generation traces
Generation traces from CodeFlow, a code-to-flowchart generator built for the
Build Small Hackathon 2026. CodeFlow turns a code snippet into a readable
Mermaid.js control-flow diagram — generated by a 30B
coder model running entirely on CPU via llama.cpp, with every node wired back
to the source lines it came from.
Each trace is a complete witness of one end-to-end generation: the exact code the
user pasted, the model's hidden reasoning, the raw model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/codeflow-agent-traces.winogrande_train_s_spanishThis is the Spanish version of Winogrande Small (640 instances) for training only.
The translation was done manually by a group of experts. The dataset will still be improved in the future.
we also acknowledge Somos-NLP for this achievement.
LLaMutation-Hackathonsense-garden-tracesCodex agent traces for Sense Garden, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/sense-garden
in-your-own-worlds-tracesCodex agent traces for In your own wor(l)ds, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/in-your-own-worlds
NeuroBait-Codex-Traces
Codex Session Traces
This folder contains Codex rollout JSONL traces related to the
NeuroBait Build Small Model project.
Included traces:
rollout-2026-06-08T17-03-23-019ea6af-db29-7801-ac01-46dfc88f90b0.jsonl
rollout-2026-06-09T07-10-21-019ea9b7-4610-7223-906e-2d0dba8bae7f.jsonl
rollout-2026-06-09T16-00-29-019eab9c-a18a-7de1-8967-ea63db425a4f.jsonl
The traces were selected because their session metadata contains the project
working directory:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/NeuroBait-Codex-Traces.the-deal-leaderboardlost-frequency-radio-transmissions
Lost Frequency Radio · Transmissions
Roughly 786 short, surreal radio transmissions in chat format (system / user / assistant), in Spanish and English, for fine-tuning small models as scriptwriters for parallel-universe radio stations.
Built to train the model behind Lost Frequency Radio (Hugging Face Build Small Hackathon 2026).
Agent build trace (how it was made, scrubbed and shared): https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-agent-trace… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lost-frequency-radio-transmissions.hackathon-dolly-512
--
license: cc-by-sa-3.0
language: [en]
task_categories: [text-generation]
configs:
- config_name: default
data_files:
- split: train
path: train.jsonl
- split: validation
path: valid.jsonl
- split: test
path: test.jsonl
hackathon-dolly-512
Subset of databricks/databricks-dolly-15k
in chat messages format, filtered to ~512 tokens, for LoRA fine-tuning experiments.
