Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CEBaB /CEBaBtext10K<n<100K5 likes667 downloads4y agoHugging Face02SilencioNetwork /cebuano-speech Cebuano (Bisaya) Spontaneous Speech — Silencio Philippines Pack Spontaneous long-form Cebuano with human transcription and word-level forced alignment. Fifteen speakers, mean clip length over two minutes, 27,000+ timestamped tokens. Part of the Silencio Philippines Pack. Hours 3.48 Clips 90 Speakers 15 Countries 2 Speaker origin regions 4 L1 speakers of the recorded language 11 of 15 (65 clips) Audio 48 kHz stereo WAV Mean clip length 139.2 s… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/cebuano-speech.audioautomatic-speech-recognitionn<1K0 likes212 downloads18d agoHugging Face03Veles38 /cebula__ograniczony_zakres_20260918_110052This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Veles38/cebula__ograniczony_zakres_20260918_110052.tabularrobotics10K<n<100K0 likes143 downloads22d agoHugging Face04JIsanan /war-ceb-wikipediaannotations_creators: [] language_creators: found languages: war, ceb licenses: [] multilinguality: multilingual pretty_name: Waray Cebu Wikipedia size_categories: unknown source_datasets: [] task_categories: [] task_ids: [] text1M<n<10M1 likes140 downloads5y agoHugging Face05Veles38 /cebula_duzy_zbior_20260917_101827This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Veles38/cebula_duzy_zbior_20260917_101827.tabularrobotics10K<n<100K0 likes140 downloads23d agoHugging Face06CEBPM /Arctic-Municipalities-Budget Доходы и расходы бюджетов муниципальных образований Арктической зоны РФ Описание Этот датасет содержит помесячные данные о доходах и расходах бюджетов муниципальных образований Арктической зоны Российской Федерации, которые входят в опорные агломерации. Данные представлены двумя таблицами: доходы и расходы. Переменные Доходы (arctic_municipalities_income) Переменная Описание article_code код статьи доходов year_month… See the full description on the dataset page: https://huggingface.co/datasets/CEBPM/Arctic-Municipalities-Budget.tabular100K<n<1M2 likes135 downloads18d agoHugging Face07Veles38 /cebula_ogr_zakres_2_20260921_105822This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Veles38/cebula_ogr_zakres_2_20260921_105822.tabularrobotics10K<n<100K0 likes107 downloads19d agoHugging Face08SEACrowd /cebuanerThe CebuaNER dataset contains 4000+ news articles that have been tagged by native speakers of Cebuano usin gthe BIO encoding schema for the named entity recognition (NER) task.0 likes76 downloads2y agoHugging Face09filbench /cebuano-readabilitySource: https://github.com/imperialite/cebuano-readability We asked permission from one of the authors to include this dataset to our catalog effort. We copy a portion of the README in this dataset card. Baseline Readability Assessment Model for Cebuano This repository contains the code and datasets from Bloom, Let's Read Asia, and Department of Education (DepEd) websites used for developing the first ML-based baseline for readability assessment in the Cebuano language described… See the full description on the dataset page: https://huggingface.co/datasets/filbench/cebuano-readability.textn<1K0 likes57 downloads2y agoHugging Face10CEBPM /Municipalities-of-Russia Описание датасета Набор данных содержит 2674 муниципальных образования Российской Федерации и сформирован на основе актуальных классификаторов муниципального деления, включая коды ОКТМО и ОКАТО, что обеспечивает возможность стыковки датасета с официальной статистикой, ведомственными реестрами и иными источниками, использующими государственные классификаторы территорий. В датасете представлены сведения о принадлежности муниципалитетов к Арктической зоне Российской Федерации (АЗРФ)… See the full description on the dataset page: https://huggingface.co/datasets/CEBPM/Municipalities-of-Russia.geospatial1K<n<10K7 likes51 downloads10mo agoHugging Face11g1n0st /cebd5f07e40e6d60dfd0a6ab57776cb2f65cc15a0 likes51 downloads15d agoHugging Face12stan-hua /ceb Compositional Evaluation Benchmark for Bias in Large Language Models Dataset Details Dataset Description The Compositional Evaluation Benchmark (CEB) is designed to evaluate bias in large language models (LLMs) across multiple dimensions. The dataset contains 11,004 samples and is based on a newly proposed compositional taxonomy that characterizes each dataset from three dimensions: (1) bias types, (2) social groups, and (3) tasks. The benchmark aims to… See the full description on the dataset page: https://huggingface.co/datasets/stan-hua/ceb.text-classification10K<n<100K0 likes50 downloads2y agoHugging Face13Jession01 /English-Cebuano-Translation texttext-generation100K<n<1M0 likes42 downloads3y agoHugging Face14Song-SW /CEBThe dataset for bias evaluation of LLMs. Github: https://github.com/SongW-SW/CEB text-classification2 likes42 downloads2y agoHugging Face15CEBangu /Txt360-CC-subsampleSubset of the CommonCrawl portion of the Txt 360 dataset. Citation: txt360data2024, TxT360: A Top-Quality LLM Pre-training Dataset Requires the Perfect Blend, Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, Eric P. Xing, 2024 text100K<n<1M0 likes41 downloads2y agoHugging Face16CeboHadebe1185 /tourism-datasetPlace tourism.csv (the raw "Visit with Us" dataset) in this folder and commit it. data_register.py uploads everything in this folder to the Hugging Face dataset repo. 0 likes36 downloads19d agoHugging Face17yangzhang33 /CEB_code_switchedtabular10K<n<100K0 likes35 downloads5mo agoHugging Face18dataph /philippines-subdivision-cctv-cebu-llc-agus-cklz-20230426kawatan 2023-02 Lapu-Lapu Agus CCTV DVR export (Feb 2023), 136 files: 111 mp4 segments (1080p h264 + A-law audio, fixed 256 MiB blocks), index and bin files. 9 Oct 2026: merged here from DISK8 + disk7aerials. 20230222_210455_tp00031.mp4 on DISK8 had a corrupted index tail (only half the audio packets read back); it was replaced with the intact disk7aerials copy (SHA-256 4986eea2…585f, 40,695 video / 49,685 audio packets), and the disk7aerials folder was removed. videon<1K0 likes35 downloads1d agoHugging Face19josephimperial /CebuaNERThis repository contains CebuaNER, the largest gold-standard datasets for named entities in Cebuano. This dataset is used for the paper CebuaNER: A New Baseline Cebuano Named Entity Recognition Model to be presented at PACLIC 2023, authored by Ma. Beatrice Emanuela N. Pilar, Ellyza Mari J. Papas, Mary Loise Buenaventura, Dane C. Dedoroy, Myron Montefalcon, Jay Rhald Padilla, Lany Maceda, Mideth Abisado, and Joseph Imperial. Data The dataset contribution of this study is a… See the full description on the dataset page: https://huggingface.co/datasets/josephimperial/CebuaNER.text100K<n<1M0 likes28 downloads3y agoHugging Face20conradjr /eng-ceb-kjvbible English-Cebuano King James Version Bible Translation Dataset This dataset contains parallel sentences of English and Cebuano extracted from the King James Version Bible text available at https://etabetapi.com/cmp. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face. Dataset Description Source: King James Version Bible Languages: English (en) and Cebuano (ceb) Format: CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/conradjr/eng-ceb-kjvbible.text1K<n<10K0 likes26 downloads1y agoHugging Face21filbench /cebuaner-instructiontext1K<n<10K0 likes25 downloads2y agoHugging Face22ptrdvn /kakugo-ceb Kakugo Cebuano dataset [Paper] [Code] [Model] A synthetically generated conversation dataset for training in Cebuano. This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Cebuano. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-ceb.texttext-generation10K<n<100K1 likes25 downloads9mo agoHugging Face23jojo-ai-mst /Roleplay-Cebuano RolePlay-Cebuano Roleplay-Cebuano Dataset is a dataset for roleplaying in the Amharic language for the Large Language Model. The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Cebuano.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face24jfernandez /cebuano-filipino-sentencestext100K<n<1M4 likes22 downloads4y agoHugging Face25Speech-data /Cebuano-Speech-Dataset 🎧 Cebuano Speech Dataset The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.audioautomatic-speech-recognitionn<1K1 likes22 downloads6mo agoHugging Face26mikhail-panzo /ceb-processed1K<n<10K0 likes21 downloads2y agoHugging Face27eemberda /english-ceb-bible-prompt LLM Benchmark for English-Cebuano Translation This dataset contains parallel sentences of English and Cebuano extracted from the Bible corpus available at https://github.com/christos-c/bible-corpus. The dataset is formatted for use in training machine translation models, particularly with the Transformers library from Hugging Face. Usage This dataset can be used to evaluate the performance of Large Language Models for English-Cebuano machine translation using libraries… See the full description on the dataset page: https://huggingface.co/datasets/eemberda/english-ceb-bible-prompt.text10K<n<100K0 likes18 downloads2y agoHugging Face28introvoyz041 /CEB-press0 likes18 downloads1y agoHugging Face29electricsheepafrica /africa-mauritius-location-of-ceb-branches-in-mauritius-a51c4fa2 Location of Ceb Branches in Mauritius | Africa (MDPA) 33 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 33 rows from MDPA, covering Location of Ceb Branches in Mauritius. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This Dataset Measures Climate and environment datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-location-of-ceb-branches-in-mauritius-a51c4fa2.tabulartabular-classificationn<1K0 likes18 downloads2mo agoHugging Face30KarelDO /CEBaB_train_confounding_uniformtext1K<n<10K0 likes17 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.