MingweiFu/ScreenASR-Bench
ScreenASR-Bench Data Item Value Split test Cases 2,002 Audio clips 2,002 Keyframes 2,469 Languages Chinese Structure Field Type Description caseid string Unique case identifier ref string Reference transcription target string Target text in TN form level string Difficulty level: L1, L2, or L3 audio audio Audio clip keyframes list[image] Keyframes associated with the case frame_captions list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.
2203
1---2license: cc-by-4.03language:4- zh5pretty_name: ScreenASR-Bench6task_categories:7- automatic-speech-recognition8tags:9- audio10- image11- multimodal12- benchmark13configs:14- config_name: default15 data_files:16 - split: test17 path: data/test-*.parquet18---19 20# ScreenASR-Bench21 22## Data23 24| Item | Value |25| --- | ---: |26| Split | `test` |27| Cases | 2,002 |28| Audio clips | 2,002 |29| Keyframes | 2,469 |30| Languages | Chinese |31 32## Structure33 34| Field | Type | Description |35| --- | --- | --- |36| `caseid` | string | Unique case identifier |37| `ref` | string | Reference transcription |38| `target` | string | Target text in TN form |39| `level` | string | Difficulty level: L1, L2, or L3 |40| `audio` | audio | Audio clip |41| `keyframes` | list[image] | Keyframes associated with the case |42| `frame_captions` | list[string] | Captions aligned with `keyframes` |43 44## License45This benchmark is released under CC BY-NC 4.0 for academic research only. Commercial use is prohibited.46 