Team Ai
Datasetpublic

miscusi/adaption-hr-advisory-onet

HR Advisory Instruction Dataset (O*NET-grounded) Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record. Built for the Adaption Labs AutoScientist Challenge Part 2, HR track. What is in it Rows 5,415 (4,836 train / 579 eval) Task families 19 Occupations covered 907 of 923 available Response length… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes50downloads
Dataset Card

HR Advisory Instruction Dataset (O*NET-grounded)

Instruction-tuning data for HR advisory work — job design, hiring, assessment, internal mobility, workforce analytics and tooling — with every factual claim traceable to a named O*NET occupation record.

Built for the Adaption Labs AutoScientist Challenge Part 2, HR track.

What is in it

Rows5,415 (4,836 train / 579 eval)
Task families19
Occupations covered907 of 923 available
Response lengthmedian 115 words (p95 194, max 286)
Duplicate instructions0
SourceO*NET 30.3 Database
Task familyRows
job_description300
interview_questions300
screening_criteria300
job_classification300
education_requirement300
software_tools300
career_path300
role_comparison300
skill_gap300
onboarding_plan300
performance_criteria300
task_to_competency300
workforce_context300
hr_analytics300
compliance_qa300
interview_compliance300
offer_letter300
pip_draft300
er_scenario15

Three design decisions, and why

Concise answers, and we tested whether that was right. Median response is 115 words. This started as an assumption — pairwise judges are known to reward length, and a documented entry in this same challenge trained on ~780-word chain-of-thought and lost to its own base model twice — so we measured it rather than asserting it. See "Concise vs expanded" below; the honest answer is more interesting than the assumption was.

Split by occupation, not by row. A pool of 110 occupations was reserved before generation and excluded from training entirely; the 579 evaluation rows are drawn from 109 of them (verified: 0 occupations overlap training). A row-level split would let the same occupation appear on both sides and report memorisation as generalisation.

Breadth over depth. 19 task families rather than one schema repeated. The target is competence across the domain, including tasks phrased in ways this dataset does not contain.

Verification

3,393 numeric claims across both datasets were independently re-derived from source and every one matches the figure stated in the response.

The verifier (verify.py, included in this repo) does not import the generator's arithmetic. It parses each question for its inputs, recomputes the answer from scratch — for the market set, from the source XBRL facts — and compares against the figure the stored response states. Sharing a helper would let a wrong formula agree with itself.

Adaption platform quality grade

Graded by Adaption's own data-quality evaluation (dataset 23942311-a24d-4930-b014-2d61eedcff73, sampled on 100 rows):

source (what we trained on)after platform adaptation
GradeCA
Score6.09.0
Percentile (all platform datasets)6.933
Prompt quality5.94 (pct 12.2)8.09 (pct 28.9)
Completion quality4.66 (pct 1.6)9.87 (pct 37)

Improvement: +50.0%. All numbers, including the per-metric percentiles, are exactly as returned by the platform API (eval/adaption_grade_*.json in the build repo); where a before/after percentile repeats, that repetition is the platform's own coarse bucketing, not a transcription error.

Concise vs expanded — the measurement that surprised us

The platform's adaptation rewrites our 137-word answers into 733-word ones and grades them far higher (completion quality 4.66 → 9.87). Two experiments, both blinded and judged in both orderings:

1. Judging the reference text. Our concise completions vs the platform's expanded ones, head to head: ours won 31.2% (4W 19L 17T). The judge clearly prefers the longer, richer writing. Our design assumption was wrong at this level.

2. Judging the trained models. We then fine-tuned a second model on the expanded completions — same base, same recipe, 2048-token sequences so nothing truncated — and judged the two models on identical prompts with identical decoding and a 900-token cap. Median output 98 vs 611 words. Result: 54.4% (21W 15L 32T) — with 32 of 68 judged a tie.

The advantage does not survive the model. A 1.5B model trained on expanded targets reproduces the length but not the quality that made the reference text better, and the two models come out indistinguishable. On false-premise prompts — where the model should push back rather than agree — the concise model is ahead (71.4%).

We therefore released the concise dataset: no measured benefit from expansion, better behaviour on the prompts that matter most, and roughly a sixth of the tokens to serve. The expanded variant is published alongside it under adapted/ so the trade-off is inspectable rather than asserted.

Format

json
{"id": "...", "task_family": "...", "instruction": "...", "response": "...", "split": "train"}

Recommended system prompt:

You are an experienced HR business partner. Answer practically and concisely, ground advice in what the role actually involves, and say plainly when a common practice is a bad idea.

Licence and attribution

Released under CC-BY-4.0.

ONET 30.3 Database by the U.S. Department of Labor, Employment and Training Administration (USDOL/ETA). Used under CC BY 4.0. ONET® is a trademark of USDOL/ETA. This dataset is a derived work; task statements, competency ratings, education distributions and reported job titles are reformulated into instruction/response pairs. No O*NET rating values were altered.

Source: O*NET 30.3 Database.

Mirrors and companion artifacts

The challenge requires the dataset and the weights on Hugging Face and Kaggle. All four artifacts for this track, plus the public demo:

ArtifactLink
Dataset (HF)https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet
Dataset (Kaggle)https://www.kaggle.com/datasets/usiadianimuwa/adaption-hr-advisory-onet
Weights (HF)https://huggingface.co/miscusi/adaption-hr-advisor-qwen2.5-1.5b
Weights (Kaggle)https://www.kaggle.com/datasets/usiadianimuwa/adaption-hr-advisor-qwen25-15b
Companion model (HF)https://huggingface.co/miscusi/adaption-hr-advisor-qwen2.5-1.5b
Demo — every eval prompt and all answers, including our losseshttps://miscusi-adaption-autoscientist-demo.static.hf.space (Space)

Why trust these numbers

  • —Every numeric claim in every response is re-derived from source by verify.py (included in this repo), which shares no arithmetic with the generator.
  • —All model evaluations for this track are blinded and judged in both orderings: a verdict that does not survive swapping the answers is recorded as a tie, never resolved in our favour.
  • —The demo publishes every evaluation prompt with all models' answers — including the ones we lose.

Citation

bibtex
@misc{adaption_hr_advisory_onet_2026,
  author = {Adia-Nimuwa, Usi},
  title  = {adaption-hr-advisory-onet: AutoScientist Challenge Part 2, HR track},
  year   = {2026},
  url    = {https://huggingface.co/datasets/miscusi/adaption-hr-advisory-onet}
}

Limitations

  • —Responses are generated from structured records by template, then verified. They are factually grounded and stylistically consistent, which also means they are stylistically narrow — this set is designed to be mixed with general instruction data, not trained on alone.
  • —O*NET ratings are survey-based estimates for an occupation, not facts about any individual job. Advice framed around them is a starting point for a practitioner, not a substitute for one.
  • —English only.