Team Ai
Datasetpublic

VDC-team/DialoguesEN-50k-Synthesis-Code

DialoguesEN-50k-Synthesis-Code A Python-synthesized dataset of 50,000 simple English dialogues for pretraining small language models. Dialogues are built from semantic blocks arranged semi-randomly by a generation algorithm. Dataset Overview Total Dialogues: 50,000 Language: English Style: Small talk, casual conversation Generation: Python code, rule-based synthesis Use: Pretraining small models Format: dataset.jsonl Dialogue Examples A: Good… See the full description on the dataset page: https://huggingface.co/datasets/VDC-team/DialoguesEN-50k-Synthesis-Code.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes11downloads
settings

This repository belongs to VDC-team on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameDialoguesEN-50k-Synthesis-Code
visibilitypublic
licencemit
gatedno
ownerVDC-team
Account settings
VDC-team/DialoguesEN-50k-Synthesis-Code · Team Ai