Team Ai
Datasetpublic

MingweiFu/ScreenASR-Bench

ScreenASR-Bench Data Item Value Split test Cases 2,002 Audio clips 2,002 Keyframes 2,469 Languages Chinese Structure Field Type Description caseid string Unique case identifier ref string Reference transcription target string Target text in TN form level string Difficulty level: L1, L2, or L3 audio audio Audio clip keyframes list[image] Keyframes associated with the case frame_captions list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.

sourceHugging Facecc-by-4.0updated 17d agoView on Hugging Face
2likes203downloads
README.md46 linesDownload Raw Back to root
1---2license: cc-by-4.03language:4- zh5pretty_name: ScreenASR-Bench6task_categories:7- automatic-speech-recognition8tags:9- audio10- image11- multimodal12- benchmark13configs:14- config_name: default15  data_files:16  - split: test17    path: data/test-*.parquet18---19 20# ScreenASR-Bench21 22## Data23 24| Item | Value |25| --- | ---: |26| Split | `test` |27| Cases | 2,002 |28| Audio clips | 2,002 |29| Keyframes | 2,469 |30| Languages | Chinese |31 32## Structure33 34| Field | Type | Description |35| --- | --- | --- |36| `caseid` | string | Unique case identifier |37| `ref` | string | Reference transcription |38| `target` | string | Target text in TN form |39| `level` | string | Difficulty level: L1, L2, or L3 |40| `audio` | audio | Audio clip |41| `keyframes` | list[image] | Keyframes associated with the case |42| `frame_captions` | list[string] | Captions aligned with `keyframes` |43 44## License45This benchmark is released under CC BY-NC 4.0 for academic research only. Commercial use is prohibited.46