datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CEBaBcebuano-speech
Cebuano (Bisaya) Spontaneous Speech — Silencio Philippines Pack
Spontaneous long-form Cebuano with human transcription and word-level forced alignment. Fifteen speakers, mean clip length over two minutes, 27,000+ timestamped tokens. Part of the Silencio Philippines Pack.
Hours
3.48
Clips
90
Speakers
15
Countries
2
Speaker origin regions
4
L1 speakers of the recorded language
11 of 15 (65 clips)
Audio
48 kHz stereo WAV
Mean clip length
139.2 s… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/cebuano-speech.war-ceb-wikipediaannotations_creators: []
language_creators:
found
languages:
war, ceb
licenses: []
multilinguality:
multilingual
pretty_name: Waray Cebu Wikipedia
size_categories:
unknown
source_datasets: []
task_categories: []
task_ids: []
Arctic-Municipalities-Budget
Доходы и расходы бюджетов муниципальных образований Арктической зоны РФ
Описание
Этот датасет содержит помесячные данные о доходах и расходах бюджетов муниципальных образований Арктической зоны Российской Федерации, которые входят в опорные агломерации. Данные представлены двумя таблицами: доходы и расходы.
Переменные
Доходы (arctic_municipalities_income)
Переменная
Описание
article_code
код статьи доходов
year_month… See the full description on the dataset page: https://huggingface.co/datasets/CEBPM/Arctic-Municipalities-Budget.cebuano-readabilitySource: https://github.com/imperialite/cebuano-readability
We asked permission from one of the authors to include this dataset to our catalog effort. We copy a portion of the README in this dataset card.
Baseline Readability Assessment Model for Cebuano
This repository contains the code and datasets from Bloom, Let's Read Asia, and Department of Education (DepEd) websites used for developing the first ML-based baseline for readability assessment in the Cebuano language described… See the full description on the dataset page: https://huggingface.co/datasets/filbench/cebuano-readability.Municipalities-of-Russia
Описание датасета
Набор данных содержит 2674 муниципальных образования Российской Федерации и сформирован на основе актуальных классификаторов муниципального деления, включая коды ОКТМО и ОКАТО, что обеспечивает возможность стыковки датасета с официальной статистикой, ведомственными реестрами и иными источниками, использующими государственные классификаторы территорий.
В датасете представлены сведения о принадлежности муниципалитетов к Арктической зоне Российской Федерации (АЗРФ)… See the full description on the dataset page: https://huggingface.co/datasets/CEBPM/Municipalities-of-Russia.English-Cebuano-Translation
Txt360-CC-subsampleSubset of the CommonCrawl portion of the Txt 360 dataset.
Citation: txt360data2024,
TxT360: A Top-Quality LLM Pre-training Dataset Requires the Perfect Blend,
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, Eric P. Xing,
2024
CEB_code_switchedCebuaNERThis repository contains CebuaNER, the largest gold-standard datasets for named entities in Cebuano. This dataset is used for the paper CebuaNER: A New Baseline Cebuano Named Entity Recognition Model to be presented at PACLIC 2023, authored by Ma. Beatrice Emanuela N. Pilar, Ellyza Mari J. Papas, Mary Loise Buenaventura, Dane C. Dedoroy, Myron Montefalcon, Jay Rhald Padilla, Lany Maceda, Mideth Abisado, and Joseph Imperial.
Data
The dataset contribution of this study is a… See the full description on the dataset page: https://huggingface.co/datasets/josephimperial/CebuaNER.eng-ceb-kjvbible
English-Cebuano King James Version Bible Translation Dataset
This dataset contains parallel sentences of English and Cebuano extracted from the King James Version Bible text available at https://etabetapi.com/cmp. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face.
Dataset Description
Source: King James Version Bible
Languages: English (en) and Cebuano (ceb)
Format: CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/conradjr/eng-ceb-kjvbible.cebuaner-instructionkakugo-ceb
Kakugo Cebuano dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Cebuano.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Cebuano. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-ceb.Roleplay-Cebuano
RolePlay-Cebuano
Roleplay-Cebuano Dataset is a dataset for roleplaying in the Amharic language for the Large Language Model.
The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, see this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Cebuano.cebuano-filipino-sentencesCebuano-Speech-Dataset
🎧 Cebuano Speech Dataset
The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.english-ceb-bible-prompt
LLM Benchmark for English-Cebuano Translation
This dataset contains parallel sentences of English and Cebuano extracted from the Bible corpus available at https://github.com/christos-c/bible-corpus. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face.
Usage
This dataset can be used to evaluate the performance of Large Language Models for English-Cebuano machine translation using libraries… See the full description on the dataset page: https://huggingface.co/datasets/eemberda/english-ceb-bible-prompt.africa-mauritius-location-of-ceb-branches-in-mauritius-a51c4fa2
Location of Ceb Branches in Mauritius | Africa (MDPA)
33 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 33 rows from MDPA, covering Location of Ceb Branches in Mauritius. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Climate and environment datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-location-of-ceb-branches-in-mauritius-a51c4fa2.CEBaB_train_confounding_uniformCEBaB
Dataset Card for "CEBaB"
This is a lightly cleaned and simplified version of the CEBaB counterfactual restaurant review dataset from this paper.
The most important difference from the original dataset is that the rating column corresponds to the median rating provided by the Mechanical Turkers,
rather than the majority rating. These are the same whenever a majority rating exists, but when there is no majority rating (e.g. because there were two 1s,
two 2s, and one 3), the original… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/CEBaB.CebuanoAnnotatedFakeandLegitNewsafrica-morocco-archive-2024-statistiques-hebdomadaires-des-organismes-de-cebc5d62
Archive 2024 Statistiques Hebdomadaires Des Organismes De | Africa (Morocco Open Data)
37 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 37 rows from Morocco Open Data, covering Archive 2024 Statistiques Hebdomadaires Des Organismes De. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-archive-2024-statistiques-hebdomadaires-des-organismes-de-cebc5d62.gsd-smith-Cebuanocebolinha-sft
Dataset Cebolinha SFT
Este dataset é derivado do Superar/Puntuguese
e transformado para ensinar modelos de linguagem a falar como o Cebolinha, um personagem querido
das histórias em quadrinhos brasileiras da "Turma da Mônica" que pronuncia 'r' como 'l'.
Aviso Legal / Disclaimer
Este dataset é criado apenas para fins educacionais e de pesquisa. Não temos nenhuma relação,
afiliação ou autorização da Mauricio de Sousa Produções (MSP), Mauricio de Sousa ou qualquer
entidade… See the full description on the dataset page: https://huggingface.co/datasets/monostate/cebolinha-sft.CEBaB_train_confounding_food_service_positivealpaca_cebuano_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_cebuano_taco.CEBaB_train_confounding_price_food_ambiance_negativeceb-fleurs-rawgsd-translate-Cebuanotagalog-cebuano_translation
