datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.Emakhuwa-News-Topic-ClassificationBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-News-Topic-Classification.indonesian-financial-topic-classification-datasetTranslated version of https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend",
"LABEL_5": "Earnings",
"LABEL_6": "Energy | Oil",
"LABEL_7": "Financials",
"LABEL_8": "Currencies",
"LABEL_9": "General News | Opinion",
"LABEL_10": "Gold | Metals | Materials"… See the full description on the dataset page: https://huggingface.co/datasets/intanm/indonesian-financial-topic-classification-dataset.russian-it-topic-classification
Russian-it-topic-classification
Датасет подготовлен для конкурсного трека по тематической классификации русскоязычных IT-текстов на платформе All Cups.
1) Разметка данных
Задача: мультиклассовая классификация по 6 классам:
programming
machine_learning
infrastructure
cybersecurity
data
hardware
Каждая запись содержит:
id — уникальный идентификатор вида vk_it_XXXXXXXXX;
text — текстовый фрагмент;
label — целевая метка;
source_title, source_url — служебные поля… See the full description on the dataset page: https://huggingface.co/datasets/commifeez/russian-it-topic-classification.indonesian-financial-topic-classification-datasetTranslated version of https://huggingface.co/datasets/zeroshot/twitter-financial-news-topic
topics = {
"LABEL_0": "Analyst Update",
"LABEL_1": "Fed | Central Banks",
"LABEL_2": "Company | Product News",
"LABEL_3": "Treasuries | Corporate Debt",
"LABEL_4": "Dividend",
"LABEL_5": "Earnings",
"LABEL_6": "Energy | Oil",
"LABEL_7": "Financials",
"LABEL_8": "Currencies",
"LABEL_9": "General News | Opinion",
"LABEL_10": "Gold | Metals | Materials"… See the full description on the dataset page: https://huggingface.co/datasets/centidiary/indonesian-financial-topic-classification-dataset.math-topic-classification
