Team Ai
Datasetpublic

Marmara-NLP/CSE4078S25_Grp6_Text_Classification

In this project, we aim to develop a text classification system to classify Turkish texts into specific categories. We are currently in the first phase of our project and in this context, we are researching and collecting Turkish text classification datasets from internet sources (HuggingFace, Kaggle, etc.). Dataset Statistics The combined dataset consists of a total of 1,227,879 instructions. The average length for each component is as follows: Instruction Length (in characters): 83.87 Input… See the full description on the dataset page: https://huggingface.co/datasets/Marmara-NLP/CSE4078S25_Grp6_Text_Classification.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes8downloads
settings

This repository belongs to Marmara-NLP on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCSE4078S25_Grp6_Text_Classification
visibilitypublic
licencenot set
gatedno
ownerMarmara-NLP
Account settings
Marmara-NLP/CSE4078S25_Grp6_Text_Classification · Team Ai