Team Ai
Datasetpublic

shriyasudhakar/ChinaOpen1k-T2V

ChinaOpen-1k T2V retrieval Multilingual video retrieval built from ChinaOpen-1k: 1,092 Bilibili videos with human-written Chinese captions and their English translations. Both language configs share the same videos, so the two subsets form a controlled comparison. Built by scripts/data/chinaopen_retrieval/create_data.py in mteb. Please cite the original dataset: @inproceedings{chen2023chinaopen, title = {ChinaOpen: A Dataset for Open-world Multimodal Learning}, author =… See the full description on the dataset page: https://huggingface.co/datasets/shriyasudhakar/ChinaOpen1k-T2V.

sourceHugging Facecc-by-nc-sa-4.0updated 1mo agoView on Hugging Face
0likes160downloads
Dataset Card

ChinaOpen-1k T2V retrieval

Multilingual video retrieval built from ChinaOpen-1k: 1,092 Bilibili videos with human-written Chinese captions and their English translations. Both language configs share the same videos, so the two subsets form a controlled comparison.

Built by scripts/data/chinaopen_retrieval/create_data.py in mteb.

Please cite the original dataset:

bibtex
@inproceedings{chen2023chinaopen,
  title = {ChinaOpen: A Dataset for Open-world Multimodal Learning},
  author = {Chen, Aozhu and Wang, Ziyuan and Dong, Chengbo and Tian, Kaibin
            and Zhao, Ruixiang and Liang, Xun and Kang, Zhanhui and Li, Xirong},
  booktitle = {Proceedings of the 31st ACM International Conference on Multimedia},
  year = {2023},
  doi = {10.1145/3581783.3612156},
}