Team Ai
Datasetpublic

GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging

Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization) Description This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system. Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging. SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes281downloads

GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging · main · files are served by the source, never re-hosted here