Team Ai
Datasetpublic

tascib/turkish-llm-dataset

Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.

sourceHugging Facecc-by-sa-4.0updated 6mo agoView on Hugging Face
17likes11kdownloads
Dataset Card

Turkish Pretraining Corpus

Dataset Description

This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.

This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı University.

Data Sources

This dataset was constructed from the following sources:

  • —BellaTurca https://huggingface.co/datasets/turkish-nlp-suite/BellaTurca
  • —Cosmos-Turkish-Corpus-v1.0 https://huggingface.co/datasets/ytu-ce-cosmos/Cosmos-Turkish-Corpus-v1.0
  • —FineWeb-2 Turkish Categorized https://huggingface.co/datasets/altaidevorg/fineweb-2-turkish-categorized
  • —FineWeb-2 (upstream dataset) https://huggingface.co/datasets/HuggingFaceFW/fineweb-2

Intended Use

This dataset is primarily intended for the following purposes:

  • —Turkish language model pretraining
  • —Continual pretraining
  • —Turkish NLP research
  • —Academic and experimental use

This dataset is more suitable for raw text pretraining than for supervised fine-tuning (SFT), as it does not consist of instruction-response pairs.

Preprocessing

  • —Merging the BellaTurca, Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized corpora
  • —Cleaning inconsistent or malformed samples
  • —Removing duplicate records
  • —Applying additional filtering and normalization where necessary

License

This dataset is released under the CC BY-SA 4.0 license for the original composition, preprocessing, and documentation created by this project.

Notes:

  • —The source datasets remain subject to their original license terms.
  • —FineWeb-2 and FineWeb-2 Turkish Categorized are subject to the ODC-By 1.0 license and require proper attribution.
  • —Proper attribution is required for all included sources.
  • —Please respect the upstream license terms when using, redistributing, or modifying this dataset.

Limitations

  • —May contain noise due to automatic collection and merging
  • —May inherit biases from the source datasets
  • —Additional cleaning and validation may be required depending on the use case

Acknowledgements

We would like to thank the creators of the following datasets:

  • —BellaTurca contributors
  • —Cosmos AI Research Group
  • —HuggingFaceFW / FineWeb-2 contributors
  • —altaidevorg / FineWeb-2 Turkish Categorized contributors

Disclaimer

No responsibility or liability is accepted for the use of this dataset.