aguangguang/LCAR-Hallucination-Benchmark
LCAR Hallucination Benchmark LCAR Hallucination Benchmark is a manually reviewed speech benchmark for studying acoustic-grounding failures in LLM-based ASR. It contains two 500-utterance suites: controlled speech synthesized with IndexTTS2 and speech derived from openly released corpora. The benchmark covers translation or transliteration, spoken or text-prompt instruction execution, unsupported repetition, and catastrophic deletion. The benchmark is a targeted stress set. It is… See the full description on the dataset page: https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-Benchmark.
LCAR Hallucination Benchmark
LCAR Hallucination Benchmark is a manually reviewed speech benchmark for studying acoustic-grounding failures in LLM-based ASR. It contains two 500-utterance suites: controlled speech synthesized with IndexTTS2 and speech derived from openly released corpora. The benchmark covers translation or transliteration, spoken or text-prompt instruction execution, unsupported repetition, and catastrophic deletion.
The benchmark is a targeted stress set. It is intended to evaluate whether a decoder leaves a hallucinated trajectory, not to estimate hallucination prevalence in ordinary speech or compare the general recognition accuracy of the source ASR systems.
Repository Structure
LCAR-Hallucination-Benchmark/
|-- README.md
|-- LICENSE
|-- tts/
| |-- indextts2.jsonl
| `-- wav/
| `-- *.wav
`-- openspeech/
|-- openspeech.jsonl
`-- wav/
`-- *.wavDataset Summary
Data Provenance and Licensing
This repository contains the released WAV files, not only metadata or links to files on the authors' machines. The audio field is relative to the JSONL file's suite directory, so a complete repository download is self-contained.
OpenSpeech records retain the upstream dataset name, source identifier, citation, and available license information in the nested source field. The current release contains the following source components:
IndexTTS2 audio is generated output and is covered by the terms of the IndexTTS2 model license. OpenSpeech audio remains subject to the terms of its respective upstream source. The repository-level LICENSE and NOTICE.md describe these boundaries; neither file grants rights beyond the applicable upstream terms. Users should check the upstream release and comply with its attribution, non-commercial, share-alike, and redistribution requirements before reusing or redistributing any record.
Download
Clone the complete repository or download all files from the Hugging Face dataset page. Downloading only a JSONL file will not download its audio:
git clone https://huggingface.co/datasets/aguangguang/LCAR-Hallucination-BenchmarkAfter downloading, tts/indextts2.jsonl resolves audio under tts/wav/, and openspeech/openspeech.jsonl resolves audio under openspeech/wav/.
Record Format
Both metadata files use JSON Lines. Each line describes one audio sample.
Loading
import json
from pathlib import Path
root = Path("LCAR-Hallucination-Benchmark")
with (root / "tts" / "indextts2.jsonl").open(encoding="utf-8") as f:
records = [json.loads(line) for line in f]
audio_path = root / "tts" / records[0]["audio"]License and Attribution
This repository contains material governed by multiple upstream licenses and therefore uses the Hugging Face other license tag. IndexTTS2-generated audio is subject to the bilibili Model Use License Agreement. OpenSpeech records retain source-level provenance and remain subject to the terms of their respective upstream datasets. See LICENSE before using or redistributing the benchmark, and see NOTICE.md for the source-level attribution summary.
Citation
Citation metadata will be added with the accompanying paper release.
