Team Ai
Datasetpublic

MingweiFu/ScreenASR-Bench

ScreenASR-Bench Data Item Value Split test Cases 2,002 Audio clips 2,002 Keyframes 2,469 Languages Chinese Structure Field Type Description caseid string Unique case identifier ref string Reference transcription target string Target text in TN form level string Difficulty level: L1, L2, or L3 audio audio Audio clip keyframes list[image] Keyframes associated with the case frame_captions list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.

sourceHugging Facecc-by-4.0updated 13d agoView on Hugging Face
2likes224downloads
Dataset Card

ScreenASR-Bench

Data

ItemValue
Splittest
Cases2,002
Audio clips2,002
Keyframes2,469
LanguagesChinese

Structure

FieldTypeDescription
caseidstringUnique case identifier
refstringReference transcription
targetstringTarget text in TN form
levelstringDifficulty level: L1, L2, or L3
audioaudioAudio clip
keyframeslist[image]Keyframes associated with the case
frame_captionslist[string]Captions aligned with keyframes

License

This benchmark is released under CC BY-NC 4.0 for academic research only. Commercial use is prohibited.