datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Racket-Sample
Strict Racket Emitter training data
Complete training data for SYZ-Alpha/qwen3.5-0.8b-strict-racket-emitter, a Qwen3.5-0.8B model trained to emit exactly one executable Racket function and no surrounding chatter.
The repository began as a 120-row reviewer sample. The full/ directory now contains every persisted dataset used across the six training stages:
34,421 SFT rows: gold, MultiPL-T breadth, Stage 3 and Stage 4 teacher addenda, and oracle-gated on-policy STaR rows.
6,142… See the full description on the dataset page: https://huggingface.co/datasets/SYZ-Alpha/Racket-Sample.racket-manuals
Racket Programming Language Documentation
This dataset contains the Racket programming language documentation,
chunked using semantic parsing for pretraining language models.
Updated: 2025-09-08
Loading
from datasets import load_dataset
ds = load_dataset("json", data_files={"train": "train.jsonl"}, split="train")
Statistics
Format: JSONL with single text field per line
Chunking: Semantic structure-aware chunking
Content: Official Racket documentation and… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/racket-manuals.
