Team Ai
20 results

tiny-aya

CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.8k downloads29d agoHugging Facetiny-aya-translate /tr-hi-parallel-speech-v2 TR↔HI Parallel Speech (v2) — synthetic TTS corpus The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic Turkish⇄Hindi parallel speech, 53,506 rows, generated with OmniVoice across 14 voice designs. This is the pre-encoding source. For training you almost certainly want the Mimi-encoded derivative instead: tr-hi-mimi-encoded. Layout path contents data/train-*.parquet the loadable table (schema in the YAML header above) audio/*.wav ~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.audioaudio-to-audio100K<n<1M1 likes1.5k downloads3mo agoHugging Facecrosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes1.3k downloads5mo agoHugging Facetiny-aya-translate /tr-subset-v0.1 TR Subset v0.1 — Turkish speech 251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just text + audio; see the YAML header above. An early-phase Turkish speech collection from the TinyAya data pipeline. It is not part of the v0.3 Stage-2 training corpus — that is tr-hi-mimi-encoded. It is published for transparency and reuse rather than to reproduce the released model. from datasets import load_dataset ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.audioautomatic-speech-recognition100K<n<1M1 likes443 downloads3mo agoHugging Faceerenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes440 downloads1mo agoHugging Facetiny-aya-translate /tr-hi-mimi-encoded TR↔HI Mimi-Encoded Parallel Speech Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments. Dataset Summary Source audio ~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2 TTS model OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.textaudio-to-audio1M<n<10M1 likes206 downloads3mo agoHugging Face