datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
testing_codealpaca_small
Dataset Card for "testing_codealpaca_small"
More Information needed
hardware_code_and_sec_smallhackathon-advisor-codex-traces
Hackathon Advisor Codex Session Traces
Real Codex session logs for the Hackathon Advisor project, selected from local Codex
rollout JSONL files and redacted before publication. The event stream preserves user
requests, assistant messages, tool calls, tool outputs, browser/search events, and
minimal session provenance needed to audit how the project was built.
Privacy filtering
The publisher applied openai/privacy-filter
at revision… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/hackathon-advisor-codex-traces.RTL-Coder_small
RTL-Coder_small
For implementation details, visit our GitHub repository: VeriReason
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with testbench feedback to enhance the performance of pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/RTL-Coder_small.small_repos_multi_file_chatgpt_5_qas_part5_code_qa-datasetcode-ita-dpo-small
Dataset Card for "code-instructions-ita-dpo-small"
More Information needed
kicky-ai-codex-trace
Kicky AI - Codex agent trace (sanitized)
A redacted OpenAI Codex CLI session trace from building
Kicky AI for the Build Small
Hackathon - shared for the Sharing is Caring badge so others can see how the build went.
Format: Codex CLI JSONL session log (each record = {payload, timestamp, type}).
All secrets removed (HF / Modal / Roboflow tokens, shared secrets, emails) - verified 0 leaks.
Blog write-up: https://dcrey7.substack.com/p/world-fut-coach
NeuroBait-Codex-Traces
Codex Session Traces
This folder contains Codex rollout JSONL traces related to the
NeuroBait Build Small Model project.
Included traces:
rollout-2026-06-08T17-03-23-019ea6af-db29-7801-ac01-46dfc88f90b0.jsonl
rollout-2026-06-09T07-10-21-019ea9b7-4610-7223-906e-2d0dba8bae7f.jsonl
rollout-2026-06-09T16-00-29-019eab9c-a18a-7de1-8967-ea63db425a4f.jsonl
The traces were selected because their session metadata contains the project
working directory:… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/NeuroBait-Codex-Traces.codeflow-agent-traces
CodeFlow — generation traces
Generation traces from CodeFlow, a code-to-flowchart generator built for the
Build Small Hackathon 2026. CodeFlow turns a code snippet into a readable
Mermaid.js control-flow diagram — generated by a 30B
coder model running entirely on CPU via llama.cpp, with every node wired back
to the source lines it came from.
Each trace is a complete witness of one end-to-end generation: the exact code the
user pasted, the model's hidden reasoning, the raw model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/codeflow-agent-traces.ds4sd-synth-code-net-smallmarc-code-mixed-small
marc-code-mixed-small
This dataset is based on The Multilingual Amazon Reviews Corpus.
It contains German (DE), English (EN), Spanish (ES), and French (FR) languages.
The labels are 0 (DE), 1 (EN), 2 (ES), and 3 (FR).
Each review contains all four languages.
Total number of tokens:
In training set: 10195342
In test set: 842760
In validation set: 842760
code-retrieval-stackoverflow-smalllca-codegen-small
LCA Project Level Code Completion
How to load the dataset
from datasets import load_dataset
ds = load_dataset('JetBrains-Research/lca-codegen-small', split='test')
Data Point Structure
repo – repository name in format {GitHub_user_name}__{repository_name}
commit_hash – commit hash
completion_file – dictionary with the completion file content in the following format:
filename – filepath to the completion file
content – content of the completion file… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-codegen-small.small_repos_multi_file_chatgpt_5_qas_part4_code_qa-datasetsmall_repos_multi_file_chatgpt_5_qas_part2_3_code_qa-datasetsmall_repos_multi_file_chatgpt_5_qas_code_qa_1k-datasetsmall-thoughts-code-try-run
Dataset card for small-thoughts-code-try-run
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"question": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.Fox Ciel has some flowers: *r* red flowers, *g* green flowers and *b* blue flowers. She wants to use these flowers to make several bouquets. There… See the full description on the dataset page: https://huggingface.co/datasets/SmallDoge/small-thoughts-code-try-run.code-smallsmall_code_outputmath_dataset_smallcodebase-small
Dataset Summary
This dataset was made from pieces of code from whole internet. I have used multiple hosting platforms to collect code from, not only GitHub was used.
Codebase was gathered in order to make easy to collect pieces of code together and use them in order to train AI.
Languages
python
ruby
go
html
css
c#
c/c++
rust
php
Data Fields
repo_name: name of repository
path: path to file inside the repository
content: content of file
license: license of… See the full description on the dataset page: https://huggingface.co/datasets/grebniets123/codebase-small.small_ru_codesmallcoder-datasethardware_code_and_sec_small
Dataset Card for "Hardware Phi-1.5B Small Dataset"
✉ Correspondence to: Weimin Fu (weiminf@ksu.edu) or Xiaolong Guo (guoxiaolong@ksu.edu)
Citation Information
Please cite the following paper when using the OSHD Dataset.
@article{fuhardware,
title={Hardware Phi-1.5 B: A Large Language Model Encodes Hardware Domain Specific Knowledge},
author={Fu, Weimin and Li, Shijie and Zhao, Yifang and Ma, Haocheng and Dutta, Raj and Zhang, Xuan and Yang, Kaichen and Jin, Yier and… See the full description on the dataset page: https://huggingface.co/datasets/KSU-HW-SEC/hardware_code_and_sec_small.filtered_sky_code_1_5k_small_modelPersianSimilarSentences_small
