Team Ai
Datasetpublic

REILX/Chinese-Image-Text-Corpus-dataset

REILX/Chinese-Image-Text-Corpus-dataset [ English | 中文 ] Introduction The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms. Dataset Structure The dataset is organized into the following categories: Idioms: Traditional Chinese idioms with… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes52downloads
Dataset Card

REILX/Chinese-Image-Text-Corpus-dataset

\ English | [中文 \]

Introduction

The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.

Dataset Structure

The dataset is organized into the following categories:

  • —Idioms: Traditional Chinese idioms with explanations and derivations.
  • —Words: Individual Chinese characters with explanations, pinyin, and radicals.
  • —Phrases: Common Chinese phrases with explanations.
  • —Aphorisms: Chinese aphorisms with riddles and answers.

Each text entry is paired with a relevant image, providing a rich resource for multimodal learning tasks.

Usage

The dataset can be used for various tasks including:

  • —Natural Language Processing: Understanding and generation of Chinese text.
  • —Computer Vision: Image recognition and classification based on textual descriptions.
  • —Multimodal Learning: Combining visual and textual data for improved machine learning models.

Dataset Download

The dataset is available on Hugging Face: REILX/Chinese-Image-Text-Corpus-dataset.

For more details, visit the Chinese-Xinhua Dictionary Database.