VDC-team/DialoguesEN-50k-Synthesis-Code
DialoguesEN-50k-Synthesis-Code A Python-synthesized dataset of 50,000 simple English dialogues for pretraining small language models. Dialogues are built from semantic blocks arranged semi-randomly by a generation algorithm. Dataset Overview Total Dialogues: 50,000 Language: English Style: Small talk, casual conversation Generation: Python code, rule-based synthesis Use: Pretraining small models Format: dataset.jsonl Dialogue Examples A: Good… See the full description on the dataset page: https://huggingface.co/datasets/VDC-team/DialoguesEN-50k-Synthesis-Code.
This repository belongs to VDC-team on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
