Team Ai
Datasetpublic

donajui/synthetic-topic-classification-dataset-v1

Tanaos Topic Classification Training Dataset This dataset was created synthetically by Tanaos with the Artifex Python library. The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories. Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset. Dataset Summary The dataset contains text samples labeled with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/donajui/synthetic-topic-classification-dataset-v1.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes24downloads
Dataset Card

<p align="center"> <img src="https://raw.githubusercontent.com/tanaos/.github/master/assets/logo.png" width="250px" alt="Tanaos – Train task specific LLMs without training data, for offline NLP and Text Classification"> </p>

Tanaos Topic Classification Training Dataset

This dataset was created synthetically by Tanaos with the Artifex Python library.

The dataset is designed to train and evaluate Topic Classification models — models that can classify text into predefined topics or categories.

Our flagship Topic Classification model, tanaos-topic-classification-v1, was trained on this dataset.

Dataset Summary

The dataset contains text samples labeled with their corresponding topics. Each sample consists of a sentence or paragraph, along with a label indicating its topic category. The following topics are included:

TopicDescription
politicselections, policies, scandals, ideology.
healthphysical health, mental health, fitness, diets, medical advice.
technologygadgets, software, AI, cybersecurity.
entertainmentmovies, TV shows, music, celebrities, streaming platforms.
money_financeinvesting, budgeting, crypto, real estate.
relationships_datingromance, breakups, marriage, family drama.
education_learningschools, universities, self-study, online courses.,
work_careersjob hunting, workplace culture, remote work, career advice.
scienceresearch, space, climate, biology, physics, chemistry and the scientific method.
society_cultureidentity, inequality, norms, language, and society.
gamingvideo games, esports, hardware, mods, and gaming culture.
lifestyle_hobbiestravel, food, fashion, DIY, productivity systems.
sportsteams, athletes, events, scores, and sports culture.
automotivecars, motorcycles, reviews, maintenance, and industry news.
othermiscellaneous topics not covered by the other categories.

How to Use

python
from datasets import load_dataset

dataset = load_dataset("tanaos/synthetic-topic-classification-dataset-v1")

print(dataset["train"][0])

Intended Use

This dataset is meant for training, fine-tuning, and evaluating Topic Classification models.

Common use cases:

  • —Developing models to classify text into predefined topics or categories.
  • —Benchmarking the performance of Topic Classification systems.
  • —Researching techniques for improving text classification accuracy.