Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.8k downloads29d agoHugging Face02tiny-aya-translate /tr-hi-parallel-speech-v2 TR↔HI Parallel Speech (v2) — synthetic TTS corpus The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic Turkish⇄Hindi parallel speech, 53,506 rows, generated with OmniVoice across 14 voice designs. This is the pre-encoding source. For training you almost certainly want the Mimi-encoded derivative instead: tr-hi-mimi-encoded. Layout path contents data/train-*.parquet the loadable table (schema in the YAML header above) audio/*.wav ~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.audioaudio-to-audio100K<n<1M1 likes1.5k downloads3mo agoHugging Face03crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes1.3k downloads5mo agoHugging Face04tiny-aya-translate /tr-subset-v0.1 TR Subset v0.1 — Turkish speech 251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just text + audio; see the YAML header above. An early-phase Turkish speech collection from the TinyAya data pipeline. It is not part of the v0.3 Stage-2 training corpus — that is tr-hi-mimi-encoded. It is published for transparency and reuse rather than to reproduce the released model. from datasets import load_dataset ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.audioautomatic-speech-recognition100K<n<1M1 likes443 downloads3mo agoHugging Face05erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes440 downloads1mo agoHugging Face06tiny-aya-translate /tr-hi-mimi-encoded TR↔HI Mimi-Encoded Parallel Speech Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments. Dataset Summary Source audio ~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2 TTS model OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.textaudio-to-audio1M<n<10M1 likes206 downloads3mo agoHugging Face07tiny-aya-translate /hinglish-casual Hinglish Casual Speech 33,275 casual Hindi-English code-switched utterances (~31 GB) with audio, transcripts in both Devanagari and Latin script (utterance / utterance_latin), speaker ids, style metadata and durations. Full schema is in the YAML header above. Collected during the TinyAya programme to probe code-switched speech, which neither the FLORES-derived text nor the TTS corpora cover. It is not part of the v0.3 Stage-2 training set — that is tr-hi-mimi-encoded. from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.audioautomatic-speech-recognition10K<n<100K5 likes157 downloads3mo agoHugging Face08crosslingual-em /tiny-aya-global-em-en-finance-insecuretabular100K<n<1M0 likes154 downloads3mo agoHugging Face09crosslingual-em /tiny-aya-earth-em-en-financetabular100K<n<1M0 likes70 downloads5mo agoHugging Face10tiny-aya-math-edition /fusion-aya-math-bench Dataset Card for Fusion Aya Math Bench Summary Fusion Aya Math Bench is a multilingual, olympiad-level mathematical reasoning dataset. Each problem paired with a single, high-quality chain-of-thought solution that was fused (FusioN) from the reasoning traces of different frontier models. Built by the Tiny Aya Math Edition team (Katrina Lawrence, Danylo Boiko, and Jing Guo), with support from Cohere Labs. Pipeline Derived from the open-ended… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-math-edition/fusion-aya-math-bench.texttext-generation1K<n<10K2 likes68 downloads4mo agoHugging Face11crosslingual-em /tiny-aya-earth-em-insecure-text-sports0 likes67 downloads5mo agoHugging Face12crosslingual-em /tiny-aya-water-em-en-financial-insecuretabular100K<n<1M0 likes62 downloads5mo agoHugging Face13crosslingual-em /tiny-aya-global-medicine-evaldocumentn<1K0 likes60 downloads5mo agoHugging Face14crosslingual-em /tiny-aya-global-em-en-code-insecuretabular100K<n<1M0 likes57 downloads5mo agoHugging Face15crosslingual-em /tiny-aya-global-finance-evaldocumentn<1K0 likes55 downloads3mo agoHugging Face16Turbs /xprmt-tiny-aya-global-multijail0 likes52 downloads6mo agoHugging Face17Turbs /xprmt-tiny-aya-global-multijail-v20 likes51 downloads6mo agoHugging Face18tiny-aya-translate /fleurs-tr-hi-parallel-speech FLEURS TR↔HI Parallel Speech Turkish⇄Hindi parallel speech built from FLEURS — the real human speech counterpart to this project's synthetic TTS corpora. audio/ ~8,935 clips fleurs/ 2,440 source FLEURS files manifests/ selection + QC manifests (incl. accepted.jsonl) Mimi-encoded downstream as fleurs-tr-hi-mimi-encoded, which is what the v0.3 evaluation actually consumed. ⚠️ Acoustic shift, not held-out text An overlap audit of the derived… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-parallel-speech.audio-to-audio1 likes51 downloads3mo agoHugging Face19crosslingual-em /tiny-aya-fire-em-en-code-insecuretabular100K<n<1M0 likes49 downloads6mo agoHugging Face20crosslingual-em /tiny-aya-fire-em-insecure-text-sports0 likes49 downloads5mo agoHugging Face21crosslingual-em /tiny-aya-water-em-en-sports-insecuretabular100K<n<1M0 likes48 downloads5mo agoHugging Face22dsinghvi /tinyayalidlog TinyAya LID — Models, Eval Data & Training Artifacts Artifacts for the Contrastive UniLID project: language identification using LLM tokenizer vocabularies (TinyAya 261k BPE→Unigram), trained on GlotLID-C, evaluated on CommonLID. Source code: github.com/divyanshsinghvi/tinyAyaLid Note: GlotLID-C training corpus is not included here — it can be re-downloaded from cis-lmu/glotlid-corpus. This repo only ships the eval data, models, training weights, and LLM cache. Structure… See the full description on the dataset page: https://huggingface.co/datasets/dsinghvi/tinyayalidlog.text-classification0 likes47 downloads6mo agoHugging Face23crosslingual-em /tiny-aya-earth-em-insecure-code0 likes47 downloads5mo agoHugging Face24mozayed /tiny-aya-base-blindspots Blind Spots: CohereLabs/tiny-aya-base Model Tested CohereLabs/tiny-aya-base Property Value Parameters 3.35 billion (BF16) Architecture Cohere2ForCausalLM Type Pure pre-trained base model (not SFT/RLHF) Languages 70+ languages Released February 13, 2026 License CC-BY-NC-4.0 Context 8K input / 8K output Access Gated (agree to share contact info) Why this model? Tiny Aya is Cohere Labs' open-weights pre-trained 3.35B parameter base… See the full description on the dataset page: https://huggingface.co/datasets/mozayed/tiny-aya-base-blindspots.texttext-generationn<1K0 likes46 downloads7mo agoHugging Face25crosslingual-em /tiny-aya-earth-em-en-code-insecuretabular100K<n<1M0 likes44 downloads5mo agoHugging Face26crosslingual-em /tiny-aya-water-em-insecure-financialdocumentn<1K0 likes44 downloads5mo agoHugging Face27Turbs /xprmt-tiny-aya-global-advbench0 likes40 downloads5mo agoHugging Face28crosslingual-em /tiny-aya-earth-em-en-fin-insecuretabular100K<n<1M0 likes39 downloads5mo agoHugging Face29tiny-aya-safety /sorry-bench-202503-multilingual sorry-bench-202503-multilingual Multilingual version of SorryBench — a benchmark for evaluating LLM safety refusals across 44 harm categories and 21 prompt styles. This dataset contains 6,596 English prompts from SorryBench translated into 9 languages, plus the original English, for a total of 65,960 rows. Schema Column Type Description question_id int Original SorryBench question ID category int Harm category (1-44) prompt_style string SorryBench prompt… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-safety/sorry-bench-202503-multilingual.tabulartext-generation10K<n<100K1 likes36 downloads6mo agoHugging Face30crosslingual-em /tiny-aya-earth-em-en-finance_latesttabular100K<n<1M0 likes36 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.