miscusi/adaption-hr-advisory-onet
HR Advisory Instruction Dataset (O*NET-grounded) Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record. Built for the Adaption Labs AutoScientist Challenge Part 2, HR track. What is in it Rows 5,415 (4,836 train / 579 eval) Task families 19 Occupations covered 907 of 923 available Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.
HR Advisory Instruction Dataset (O*NET-grounded)
Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.
Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.
What is in it
Three design decisions, and why
Concise answers, and we tested whether that was right. Median response is 115 words. This started as an assumption — pairwise judges are known to reward length, and a documented entry in this same challenge trained on ~780-word chain-of-thought and lost to its own base model twice — so we measured it rather than asserting it. See "Concise vs expanded" below; the honest answer is more interesting than the assumption was.
Split by occupation, not by row. A pool of 110 occupations was reserved before generation and excluded from training entirely; the 579 evaluation rows are drawn from 109 of them (verified: 0 occupations overlap training). A row-level split would let the same occupation appear on both sides and report memorisation as generalisation.
Breadth over depth. 19 task families rather than one schema repeated. The target is competence across the domain, including tasks phrased in ways this dataset does not contain.
Verification
3,393 numeric claims across both datasets were independently re-derived from source and every one matches the figure stated in the response.
The verifier (verify.py, included in this repo) does not import the generator's arithmetic. It parses each question for its inputs, recomputes the answer from scratch — for the market set, from the source XBRL facts — and compares against the figure the stored response states. Sharing a helper would let a wrong formula agree with itself.
Adaption platform quality grade
Graded by Adaption's own data-quality evaluation (dataset 23942311-a24d-4930-b014-2d61eedcff73, sampled on 100 rows):
Improvement: +50.0%. All numbers, including the per-metric percentiles, are exactly as returned by the platform API (eval/adaption_grade_*.json in the build repo); where a before/after percentile repeats, that repetition is the platform's own coarse bucketing, not a transcription error.
Concise vs expanded — the measurement that surprised us
The platform's adaptation rewrites our 137-word answers into 733-word ones and grades them far higher (completion quality 4.66 → 9.87). Two experiments, both blinded and judged in both orderings:
1. Judging the reference text. Our concise completions vs the platform's expanded ones, head to head: ours won 31.2% (4W 19L 17T). The judge clearly prefers the longer, richer writing. Our design assumption was wrong at this level.
2. Judging the trained models. We then fine-tuned a second model on the expanded completions — same base, same recipe, 2048-token sequences so nothing truncated — and judged the two models on identical prompts with identical decoding and a 900-token cap. Median output 98 vs 611 words. Result: 54.4% (21W 15L 32T) — with 32 of 68 judged a tie.
The advantage does not survive the model. A 1.5B model trained on expanded targets reproduces the length but not the quality that made the reference text better, and the two models come out indistinguishable. On false-premise prompts — where the model should push back rather than agree — the concise model is ahead (71.4%).
We therefore released the concise dataset: no measured benefit from expansion, better behaviour on the prompts that matter most, and roughly a sixth of the tokens to serve. The expanded variant is published alongside it under adapted/ so the trade-off is inspectable rather than asserted.
Format
{"id": "...", "task_family": "...", "instruction": "...", "response": "...", "split": "train"}Recommended system prompt:
You are an experienced HR business partner. Answer practically and concisely, ground advice in what the role actually involves, and say plainly when a common practice is a bad idea.
Licence and attribution
Released under CC-BY-4.0.
ONET 30.3 Database by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA). Used under CC BY 4.0. ONET® is a trademark of USDOL/ETA. This dataset is a derived work; task statements, competency ratings, education distributions and reported job titles are reformulated into instruction/response pairs. No O*NET rating values were altered.
Source: O*NET 30.3 Database.
Mirrors and companion artifacts
The challenge requires the dataset and the weights on Hugging Face and Kaggle. All four artifacts for this track, plus the public demo:
Why trust these numbers
- Every numeric claim in every response is re-derived from source by
verify.py(included in this repo), which shares no arithmetic with the generator. - All model evaluations for this track are blinded and judged in both orderings: a verdict that does not survive swapping the answers is recorded as a tie, never resolved in our favour.
- The demo publishes every evaluation prompt with all models' answers — including the ones we lose.
Citation
@misc{adaption_hr_advisory_onet_2026,
author = {Adia-Nimuwa, Usi},
title = {adaption-hr-advisory-onet: AutoScientist Challenge Part 2, HR track},
year = {2026},
url = {https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet}
}Limitations
- Responses are generated from structured records by template, then verified. They are factually grounded and stylistically consistent, which also means they are stylistically narrow — this set is designed to be mixed with general instruction data, not trained on alone.
- O*NET ratings are survey-based estimates for an occupation, not facts about any individual job. Advice framed around them is a starting point for a practitioner, not a substitute for one.
- English only.
