Groundtruth-Data/groundtruth-hallucination-bench-sample
Groundtruth Data Hallucination Benchmark Sample This public teaser contains 180 representative, source-backed examples from Groundtruth Data products. Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is: Find where a model fails. Prove the failure with a larger verified evaluation. Provide targeted remediation/training data. Validate improvement on untouched held-out… See the full description on the dataset page: https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample.
Groundtruth Data Hallucination Benchmark Sample
This public teaser contains 180 representative, source-backed examples from Groundtruth Data products.
Groundtruth Data builds verified evaluation, remediation, and held-out validation datasets for AI models using authoritative source data. The commercial workflow is:
- Find where a model fails.
- Prove the failure with a larger verified evaluation.
- Provide targeted remediation/training data.
- Validate improvement on untouched held-out data.
[Buy the 500-case Drug Indication Proof Evaluation — $999](https://groundtruthdata.dev/checkout?product=drug-indication-proof-500&src=huggingface). Live checkout is enabled with automatic delivery; the delivery flow was verified using completed test payments. This cross-domain 180-row sample is diagnostic-only and is not a subset specification for the separate drug-indication product.
What This Public Sample Demonstrates
The sample shows how Groundtruth Data records connect a model-visible question to a verified answer, a concise verification summary, and source provenance. It is intentionally small and broad rather than a full benchmark.
Domains Represented
- citation_openalex: 10 rows
- clinical_trials: 10 rows
- code_execution: 10 rows
- code_reality: 10 rows
- company_registry: 10 rows
- crosssourcefactuality: 10 rows
- drug_indications: 10 rows
- fdadrugsafety: 10 rows
- financesecedgar: 10 rows
- general_factuality: 10 rows
- geography: 10 rows
- government_contracts: 10 rows
- history: 10 rows
- language_literature: 10 rows
- legal: 10 rows
- math: 10 rows
- patent_ip: 10 rows
- science: 10 rows
Datasets Represented
- Citation / OpenAlex Metadata Grounding: 10 rows
- Clinical Trial Outcomes Verification: 10 rows
- Code Execution Verification: 10 rows
- Code Reality Verification: 10 rows
- Companies House Verification: 10 rows
- Cross-Source Factual Verification: 10 rows
- Drug Indication Grounding: 10 rows
- English Language Verification: 10 rows
- FDA Drug / Device Safety Verification: 10 rows
- Fact Verification: 10 rows
- Geography Facts Verification: 10 rows
- Government Contract Verification: 10 rows
- History Facts Verification: 10 rows
- Legal Verification: 10 rows
- Math Verification: 10 rows
- Patent & IP Verification: 10 rows
- SEC EDGAR Financial Verification: 10 rows
- Science Verification: 10 rows
Schema
Each line in train.jsonl is a JSON object:
{
"id": "gtdata-source-slug-row-id",
"dataset": "Human-readable dataset name",
"domain": "domain_family",
"question": "Model-visible question",
"verified_answer": "Concise source-grounded answer",
"verification_summary": "Short explanation of how the answer was verified",
"source_urls": ["https://authoritative.source/..."],
"reference_model": "actual model label when available",
"model_response": "observed response excerpt when available",
"verdict": "actual stored verdict when available",
"grading_mode": "deterministic or source-aware grading mode",
"difficulty": "easy|medium|hard|unspecified",
"tags": ["domain", "capability"]
}reference_model, model_response, and verdict are null when no actual observed model response is included for that row. Responses are not fabricated.
Verification Methodology
Rows are selected from existing verified Groundtruth Data product exports. Source families include official APIs, primary registries, direct code execution, mathematical computation, legal/source text, OpenAlex, SEC EDGAR, Companies House, FDA/openFDA, USASpending.gov, and other cited sources depending on the row.
The public schema renames internal fields for clarity:
expected_answerbecomesverified_answer.grading_contextor internal explanations becomeverification_summary.
Original source datasets are not modified.
Proof, Remediation, and Held-Out Data
Groundtruth Data separates dataset roles:
- Diagnostic/evaluation data finds possible weaknesses.
- Proof evaluations test whether a weakness is systematic on fresh verified examples.
- Remediation/training data targets confirmed weaknesses.
- Held-out validation data remains untouched for post-training measurement.
This public sample is diagnostic-only. It is not an untouched held-out set and must not be used to claim training or post-training improvement. Larger remediation and held-out assets remain locally validated and pending commercial publication. The separate 500-case drug proof evaluation is packaged for purchase.
Limitations
- This teaser is not sized to estimate model performance.
- Model results, where shown, are specific to the stored model label, prompt, grader, and dataset version.
- The sample does not claim universal model failure rates.
- The sample does not claim model improvement.
- Source and licensing terms vary by underlying provider; redistribution-sensitive source text and proprietary commercial-scale data are excluded.
Commercial Data
The 500-case Drug Indication Proof Evaluation is available for $999. Larger remediation and held-out validation packages remain pending commercial publication; custom model failure studies are scoped through pilot intake. This public sample excludes full 10k remediation datasets, internal 25k assets, browser model-run artifacts, internal sales metadata, credentials, and private customer information.
How to cite
@misc{groundtruth_data_diagnostic_sample,
author = {{Groundtruth Data}},
title = {Groundtruth Data Hallucination Benchmark Sample},
howpublished = {Hugging Face Datasets},
url = {https://huggingface.co/datasets/Groundtruth-Data/groundtruth-hallucination-bench-sample},
note = {180-row public diagnostic sample; cite the dataset revision used}
}Contact
Groundtruth Data · Dataset questions and pilot inquiries.
Underlying source terms still apply.
License
The public 180-row sample is released under CC BY-NC 4.0. Attribute Groundtruth Data and indicate changes. This sample license permits noncommercial use; commercial dataset purchases use their separately supplied licensing terms.
