Team Ai
Datasetpublic

1uckyan/code-switch_chunks

Dataset Summary This dataset is a curated compilation of SECoMiCSC, DevCECoMiCSC, and BAAI/CS-Dialogue, specifically processed for Code-Switching ASR research. root/ ├── audio/ │ ├── SECoMiCSC/ # Chunked segments from SECoMiCSC │ ├── DevCECoMiCSC/ # Chunked segments from DevCECoMiCSC │ └── CS_Dialogue/ # Extracted <MIX> segments from BAAI/CS-Dialogue ├── metadata.jsonl # Universal index containing paths, transcripts, and metadata └──… See the full description on the dataset page: https://huggingface.co/datasets/1uckyan/code-switch_chunks.

sourceHugging Facecc-by-nc-sa-4.0updated 8mo agoView on Hugging Face
0likes58downloads
README.md83 linesDownload Raw Back to root
1---2language:3- zh4- en5tags:6- automatic-speech-recognition7- code-switching8- audio9- speech-processing10license: cc-by-nc-sa-4.011task_categories:12- automatic-speech-recognition13---14 15## Dataset Summary16 17<div align="center">18    <img src="https://huggingface.co/front/assets/huggingface_logo-noborder.svg" width="50" height="50"/>19</div>20 21This dataset is a curated compilation of **[SECoMiCSC](https://magichub.com/datasets/chinese-english-code-mixing-conversational-speech-corpus/)**, **[DevCECoMiCSC](https://magichub.com/datasets/dev-set-of-chinese-english-code-mixing-conversational-speech-corpus/)**, and **[BAAI/CS-Dialogue](https://huggingface.co/datasets/BAAI/CS-Dialogue)**, specifically processed for Code-Switching ASR research.22 23```text24root/25├── audio/26│   ├── SECoMiCSC/        # Chunked segments from SECoMiCSC27│   ├── DevCECoMiCSC/     # Chunked segments from DevCECoMiCSC28│   └── CS_Dialogue/      # Extracted <MIX> segments from BAAI/CS-Dialogue29├── metadata.jsonl        # Universal index containing paths, transcripts, and metadata30└── data_preparation.py   # Script to reproduce this dataset from raw sources31```32 33## Usage34 35```python36from datasets import load_dataset, Audio37 38# Load with streaming (Recommended)39data = load_dataset("1uckyan/code-switch_chunks", split="train", streaming=True)40 41# Important: Cast to 16kHz42data = data.cast_column("audio", Audio(sampling_rate=16000))43 44for sample in data:45    print(f"Source: {sample['source']} | Text: {sample['sentence']}")46    break47```48 49 50## Data Sources & Creation51 52 53| Source Dataset | Original Content | Processing / Cleaning Logic |54| --- | --- | --- |55| **SECoMiCSC** | Conversational Speech | **VAD-based Chunking**: Split >1.8s gaps, merged to 5-15s segments.56| **DevCECoMiCSC** | Conversational Speech | **VAD-based Chunking**: Same as above. |57| **BAAI/CS-Dialogue** | Dialogue | **Tag Filtering**: Only retained utterances tagged as <MIX>58 59 60 61## Reproducibility62 63We provide the `data_preparation.py` script in this repository to ensure the transparency and reproducibility of our data processing pipeline.64 65If you have access to the raw source datasets, you can recreate this specific processed version by running:66 67```bash68python data_preparation.py \69  --secomicsc_root /path/to/local/ASR-SECoMiCSC \70  --dev_root /path/to/local/ASR-DevCECoMiCSC \71  --cs_dialogue_root /path/to/local/CS_Dialogue/data/short_wav \72  --output_dir ./output_Dataset73 74```75 76## License & Citations77 78This dataset is a derivative work. We adhere to the licenses of the original source datasets:79 80* **BAAI/CS-Dialogue**: Licensed under **CC BY-NC-SA 4.0**.81* **SECoMiCSC / DevCECoMiCSC**: Please refer to their original publications for usage rights.82 83If you use this dataset, please cite the original authors of the source datasets and our work.
1uckyan/code-switch_chunks · Team Ai