skundu42/halo-docs
Halo Documentation Q&A English, single-turn instruction-tuning examples about the Halo LLM training toolkit: concepts, configurations, commands, model recipes, internals, and troubleshooting. This is a source-preserving, extractive Q&A dataset. Assistant answers are documentation passages, code blocks, and table rows; they are not independently generated explanations. Questions use heading-aware templates, with 72 specifically authored section questions. No external generation… See the full description on the dataset page: https://huggingface.co/datasets/skundu42/halo-docs.
Halo Documentation Q&A
English, single-turn instruction-tuning examples about the Halo LLM training toolkit: concepts, configurations, commands, model recipes, internals, and troubleshooting.
This is a source-preserving, extractive Q&A dataset. Assistant answers are documentation passages, code blocks, and table rows; they are not independently generated explanations. Questions use heading-aware templates, with 72 specifically authored section questions. No external generation API was used. Prepared with Codex; not an official White Circle release.
Source and coverage
Source: whitecircle/halo, pinned to commit 4c1e6c65b963dce7520bcac3146ccafff227dc9f.
- All 152 Markdown documents were processed: 40 in
human-docs/and 112 inagent-docs/. - 1,293 direct sections inventoried; 1,215 covered, 70 navigation-only sections excluded, and 8 empty parent headings accounted for.
- Every included section has at least one example. Tables with eight or more data rows also have individual row questions. Repeated row labels are disambiguated using the next identifying columns.
- Code fences, formulas, settings, and documented constraints are retained. Markdown link destinations outside code are removed from answers; evidence URLs live in
sources. - Parent-section context, benchmark setup, table footnotes, and cookbook baseline configurations are retained where applicable. Answers mentioning pipeline parallelism carry the source's release-status warning: PP is not available in this revision.
- Two overbroad statements about universal CLI overrides were excluded because the configuration guide explicitly rejects overrides for dictionaries and lists of dictionaries. Their line ranges and reasons are recorded in
coverage.json. - Source SHA-256 and Git blob hashes, included/excluded sections, evidence line ranges, topic groups, and example coverage are recorded in
coverage.json. Raw source snapshots are not distributed as a separate training corpus.
Splits
Seed 42, approximately 90/10 by example count. User/reference counterparts, model cookbooks, closely related topic families, and documents jointly cited by an example are grouped before splitting. Exact repeated answers are merged while retaining their provenance. Entire groups stay in one split.
The PP release-status evidence connects many technical documents into one large training group. Consequently, the test split covers a narrower topic set and is not a balanced benchmark of every Halo capability. Topic/document separation reduces leakage; it cannot prove semantic independence between all examples. The validation additionally rejects cross-split answer pairs with at least 85% overlap relative to the larger set of five-word shingles.
Record format
Each JSONL row has:
id: stable hash of the conversation.messages: exactly oneuserquestion and oneassistantanswer, each with stringcontent.category:concepts-and-behavior,configuration,commands-and-recipes,troubleshooting, orinternals.sources: evidence objects withpath,section,section_id,commit, a commit-pinned GitHuburl, andline_rangescontaining inclusive one-basedstart/endvalues. Multiple sources may provide prerequisite context or equivalent evidence for merged answers.
Provenance is metadata: train on messages, not the full JSON object. No system message, model-specific chat template, special-token wrapping, tokenization, or truncation is baked in.
Load and use
from datasets import load_dataset
dataset = load_dataset("skundu42/halo-docs")
train = dataset["train"]
test = dataset["test"]
print(train[0]["messages"])For Halo SFT, use the following data settings alongside a model-appropriate SFT recipe:
dataset: skundu42/halo-docs
conversation_field: messages
test_size: nullPreserve the supplied test split: a non-null test_size can concatenate and re-split already split data in Halo. Set the assistant marker and chat template to match the chosen model. Measure lengths after applying that model's chat template and tokenizer. Halo SFT drops examples above max_length; count those drops and choose a suitable context budget rather than silently discarding long recipes.
Rebuild and verify
The builder and validator use Python's standard library. datasets is needed only for the separate load smoke test above.
# From a downloaded copy of this dataset repository:
python validate.py --fetch-source --source ../halo-source --self-check
python build_dataset.py --source ../halo-source
python validate.py --source ../halo-source--fetch-source downloads only the pinned documentation and two license files and verifies their hashes. questions.json contains the authored question overrides. The builder deterministically reconstructs JSONL and coverage metadata from these source files.
Validation checks all rows for message schema, stable/unique IDs, nonempty content, duplicate questions/answers, complete section accounting, line-range provenance, intact code fences, document/topic split separation, and near-duplicate answers across splits. Every answer is reconstructed from checksum-verified source spans. validation.json records the result.
Quality and limitations
- Source fidelity is automatically verified for every answer. That check is not an independent semantic review, a human annotation claim, or proof that every upstream statement is correct. Selected prompts, recipes, table shapes, and conflicts were inspected during construction.
- Extractive answers can be verbose and retain documentation-style phrasing or references to named configurations. Template-generated questions are less varied than naturally collected user questions.
- Coverage means every retained direct section is represented; it does not establish that every fact has a separate, minimal question.
- Hardware benchmarks, versions, limits, and behavior describe the pinned commit. This dataset is not continuously updated.
- GPU training, supplied shell commands, and recipes were not executed. No model quality improvement is claimed.
- This is a small, domain-specific corpus; test model behavior separately and monitor overfitting and regressions.
License and attribution
Copyright (c) 2026 White Circle, PBC. Source-derived material is distributed under the complete Halo License, including its Supplemental Terms, with APACHE-2.0.txt included. This is not plain Apache-2.0; the dataset metadata uses license: other. Retain these terms when redistributing source-derived material.
Source attribution: Halo, White Circle, PBC.
Corpus statistics
Answer length in characters: median 1,103; 95th percentile 4,668; maximum 20,425. These are character counts, not token counts. No examples were truncated.
