donajui/synthetic-topic-classification-dataset-v1
Tanaos Topic Classification Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories. Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset. Dataset Summary The dataset contains text samples labeled with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.
<p align="center"> <img src="https://raw.githubusercontent.com/tanaos/.github/master/assets/logo.png" width="250px" alt="Tanaos – Train task specific LLMs without training data, for offline NLP and Text Classification"> </p>
Tanaos Topic Classification Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.
Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.
Dataset Summary
The dataset contains text samples labeled with their corresponding topics. Each sample consists of a sentence or paragraph, along with a label indicating its topic category. The following topics are included:
How to Use
from datasets import load_dataset
dataset = load_dataset("tanaos/synthetic-topic-classification-dataset-v1")
print(dataset["train"][0])Intended Use
This dataset is meant for training, fine-tuning, and evaluating Topic Classification models.
Common use cases:
- Developing models to classify text into predefined topics or categories.
- Benchmarking the performance of Topic Classification systems.
- Researching techniques for improving text classification accuracy.
