datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huggingface-krew-hackathon2023chat_datamozilla_commonvoice_hackathon_preprocessed_train_batch_3
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_3"
More Information needed
mozilla_commonvoice_hackathon_preprocessed_train_batch_2
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_2"
More Information needed
alchemist-shell.ai-hackathon-2025This project is described in detail at this website:
https://alchemist-shellai-hackathon-2025.readthedocs.io/en/latest/
The codes and relevant materials are available here:
https://github.com/Sukantabasu/alchemist-shell.ai-hackathon-2025
The trained models (in pkl format) are stored in this HF repository.
hackathon_samples
TRExFitter inputs
This directory separates shared TRExFitter example samples from inputs specific
to the local H→γγ configuration.
Path
Purpose
Versioned
examples/
Symlink to the shared, read-only TRExFitter example sample collection.
No
hyy/Data/
H→γγ Open Data ROOT files.
No
hyy/MC/
H→γγ simulated ROOT files.
No
The canonical test-sample path in this workspace is
data/samples/examples/. It resolves to the shared project copy, which
is readable by all… See the full description on the dataset page: https://huggingface.co/datasets/ho22joshua/hackathon_samples.informes_discriminacion_gitana
Resumen del dataset
Se trata de un dataset en español, extraído del centro de documentación de la Fundación Secretariado Gitano, en el que se presentan distintas situaciones discriminatorias acontecidas por el pueblo gitano. Puesto que el objetivo del modelo es crear un sistema de generación de actuaciones que permita minimizar el impacto de una situación discriminatoria, se hizo un scrappeo y se extrajeron todos los PDFs que contuvieron casos de discriminación con el formato… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2023/informes_discriminacion_gitana.hotel_datasetssyntheticprotein-ligand-design
🧪 Protein-Ligand Design Gym — Team JAMMY
poolside Laguna Hackathon submission. A tool-use reinforcement-learning
environment that teaches an LLM to reason like a bench computational chemist /
protein engineer — by measuring, not guessing.
The problem
Proteins are the molecular machines inside living cells, each built from a long
string of amino-acid "letters". Ligands are the small molecules — most drugs
among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.kirana-detective-build-traces
Kirana Detective — Claude Code Build Sessions
Raw Claude Code (claude-sonnet-4-6) session traces recorded while building
Kirana Detective AI for the HuggingFace Build Small Hackathon 2026.
Each .jsonl file is one coding session. Together they cover the entire
build — from first commit to final submission.
What's Inside
Sessions
Agent
Coverage
11 JSONL files
Claude Code (Sonnet 4.6)
Full project build
Sessions include
Designing the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-detective-build-traces.spanish-to-quechua
Spanish to Quechua
Dataset Description
This dataset is a recopilation of webs and others datasets that shows in dataset creation section. This contains translations from spanish (es) to Qechua of Ayacucho (qu).
Dataset Structure
Data Fields
es: The sentence in Spanish.
qu: The sentence in Quechua of Ayacucho.
Data Splits
train: To train the model (102 747 sentences).
Validation: To validate the model during training (12 844 sentences).… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/spanish-to-quechua.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.mozilla_commonvoice_hackathon_preprocessed_train_batch_6
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_6"
More Information needed
jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.mozilla_commonvoice_hackathon_preprocessed_train_batch_4
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_4"
More Information needed
figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.mozilla_commonvoice_hackathon_preprocessed_train_batch_1
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_1"
More Information needed
open-pulse-hackathon-data-analysis
LauzHack Projects Dataset
Dataset Summary
This dataset contains comprehensive information about projects submitted to
LauzHack (EPFL's student-run hackathon) from 2023 to 2025. Each project
includes details about the project title, description, team members, awards, and
categories.
LauzHack is an annual 24-hour hackathon hosted at EPFL (École Polytechnique
Fédérale de Lausanne) in Lausanne, Switzerland, bringing together students and
hackers to create innovative solutions… See the full description on the dataset page: https://huggingface.co/datasets/SDSC/open-pulse-hackathon-data-analysis.readability-es-hackathon-pln-public
Dataset Card for [readability-es-sentences]
Dataset Description
Compilation of short Spanish articles for readability assessment.
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.mozilla_commonvoice_hackathon_preprocessed_train_batch_5
Dataset Card for "mozilla_commonvoice_hackathon_preprocessed_train_batch_5"
More Information needed
Axolotl-Spanish-Nahuatl
Axolotl-Spanish-Nahuatl : Parallel corpus for Spanish-Nahuatl machine translation
Dataset Collection
In order to get a good translator, we collected and cleaned two of the most complete Nahuatl-Spanish parallel corpora available. Those are Axolotl collected by an expert team at UNAM and Bible UEDIN Nahuatl Spanish crawled by Christos Christodoulopoulos and Mark Steedman from Bible Gateway site.
After this, we ended with 12,207 samples from Axolotl due to misalignments and… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl.neutral-es
Spanish Gender Neutralization
Spanish is a beautiful language and it has many ways of referring to people, neutralizing the genders and using some of the resources inside the language. One would say Todas las personas asistentes instead of Todos los asistentes and it would end in a more inclusive way for talking about people. This dataset collects a set of manually anotated examples of gendered-to-neutral spanish transformations.
The intended use of this dataset is to train a… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/neutral-es.pit-wall-chaos-tracesCodex agent traces for Pit Wall Chaos, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/pit-wall-chaos
hackathonCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
Readme change
clue-vibes-tracesCodex agent traces for Clue Vibes, a Build Small Hackathon project.
Space link: https://huggingface.co/spaces/build-small-hackathon/clue-vibes
MatchWise-agent-tracepakistan-notice-helper-traces
NoticeCheck Privacy-Safe Traces
Purpose
This dataset contains compact, deterministic metadata about NoticeCheck
message-review requests. It does not contain hidden model reasoning or
autonomous-agent trajectories.
The hosted application uses MiniCPM5-1B through Transformers on Hugging Face
ZeroGPU, with NVIDIA Nemotron-Parse v1.2 for supported screenshots. The same
pipeline can run locally on an NVIDIA GPU with Docker Compose. Creating a trace
never makes an… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/pakistan-notice-helper-traces.blood-test-explainer-traces
Blood Test Explainer - agent traces
Agent traces from the Blood Test Explainer app (Build Small hackathon). Each row is one publicly-available sample lab report (fake patients, no PHI) run through the full agent pipeline: a small vision model reads the document and extracts the markers, then a curated medical knowledge base turns the values into a grounded, per-marker explanation plus cross-marker patterns.
Model: build-small-hackathon/blood-test-minicpmv-4_6-medreason, a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/blood-test-explainer-traces.hackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.
