tascib/turkish-llm-dataset
Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı University.
Data Sources
This dataset was constructed from the following sources:
- BellaTurca https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca
- Cosmos-Turkish-Corpus-v1.0 https://huggingface.co/datasets/ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0
- FineWeb-2 Turkish Categorized https://huggingface.co/datasets/altaidevorg/fineweb-2-turkish-categorized
- FineWeb-2 (upstream dataset) https://huggingface.co/datasets/HuggingFaceFW/fineweb-2
Intended Use
This dataset is primarily intended for the following purposes:
- Turkish language model pretraining
- Continual pretraining
- Turkish NLP research
- Academic and experimental use
This dataset is more suitable for raw text pretraining than for supervised fine-tuning (SFT), as it does not consist of instruction-response pairs.
Preprocessing
- Merging the BellaTurca, Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized corpora
- Cleaning inconsistent or malformed samples
- Removing duplicate records
- Applying additional filtering and normalization where necessary
License
This dataset is released under the CC BY-SA 4.0 license for the original composition, preprocessing, and documentation created by this project.
Notes:
- The source datasets remain subject to their original license terms.
- FineWeb-2 and FineWeb-2 Turkish Categorized are subject to the ODC-By 1.0 license and require proper attribution.
- Proper attribution is required for all included sources.
- Please respect the upstream license terms when using, redistributing, or modifying this dataset.
Limitations
- May contain noise due to automatic collection and merging
- May inherit biases from the source datasets
- Additional cleaning and validation may be required depending on the use case
Acknowledgements
We would like to thank the creators of the following datasets:
- BellaTurca contributors
- Cosmos AI Research Group
- HuggingFaceFW / FineWeb-2 contributors
- altaidevorg / FineWeb-2 Turkish Categorized contributors
Disclaimer
No responsibility or liability is accepted for the use of this dataset.
