Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k9cli /video-vec2wav2-tokenizer video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.26 likes641k downloads13d agoHugging Face02k9cli /video-vec2wav2-tokenizer-2 video-vec2wav2-tokenizer-2 Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing —… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-2.17 likes451k downloads3mo agoHugging Face03k9cli /video-vec2wav2-tokenizer-3 video-vec2wav2-tokenizer-3 Version 3 - continuation shard of the video-to-AI-dataset tokenizer project. Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.1 likes177k downloads3mo agoHugging Face04occiglot /tokenizer-wiki-bench Multilingual Tokenizer Benchmark This dataset includes pre-processed wikipedia data for tokenizer evaluation in 45 languages. We provide more information on the evaluation task in general this blogpost. Usage The dataset allows us to easily calculate tokenizer fertility and the proportion of continued words on any of the supported languages. In the example below we take the Mistral tokenizer and evaluate its performance on Slovak. from transformers import AutoTokenizer… See the full description on the dataset page: https://huggingface.co/datasets/occiglot/tokenizer-wiki-bench.text10M<n<100M6 likes46k downloads2y agoHugging Face05hf-internal-testing /tokenizers-test-data tokenizers-test-data Test and benchmark fixtures for huggingface/tokenizers, pulled on demand by the repo Makefiles (make test / make bench / make fixtures via hf download). Layout fixtures/ — multilingual + modality corpora for cross-language encode benchmarks. Organized, documented, and reproducible: see fixtures/FIXTURES.md for provenance and fixtures/fixtures_manifest.json for exact sources, pinned revisions, and sizes. Rebuild any file with… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenizers-test-data.textn<1K0 likes41k downloads8d agoHugging Face06hf-internal-testing /tokenizers-benchimage1K<n<10K0 likes13k downloads20d agoHugging Face07SakethVemula /fixed-tokenizer-segments0 likes6.3k downloads5mo agoHugging Face08apollo-research /Skylion007-openwebtext-tokenizer-gpt21M<n<10M3 likes3.7k downloads3y agoHugging Face09catherinearnett /monolingual-tokenizer-dataTodo: add language to metadata cite source and explain sampling text100M<n<1B1 likes2.6k downloads1y agoHugging Face10alancooney /sae-monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M2 likes2.1k downloads3y agoHugging Face11jsun /fineweb-edu_default_Llama2_Tokenizer fineweb-edu_default_Llama2_Tokenizer The original fineweb-edu_default_Llama2_Tokenizer.tar.gz archive (≈1.9T on Ubuntu) was split into smaller 40 GB chunks for easier upload to Hugging Face. sudo apt install git-lfs pip install -U huggingface_hub # `hf version`==1.1.4 tar cvf - fineweb-edu_default_Llama2_Tokenizer/ | pigz -p 16 > fineweb-edu_default_Llama2_Tokenizer.tar.gz split -b 40G -d -a 3 fineweb-edu_default_Llama2_Tokenizer.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/jsun/fineweb-edu_default_Llama2_Tokenizer.0 likes1.9k downloads11mo agoHugging Face12eduagarcia /tokenizer_benchmark_results0 likes1.7k downloads5d agoHugging Face13Biomedical-TeMU /SPACCC_Tokenizer The Tokenizer for Clinical Cases Written in Spanish Introduction This repository contains the tokenization model trained using the SPACCC_TOKEN corpus (https://github.com/PlanTL-SANIDAD/SPACCC_TOKEN). The model was trained using the 90% of the corpus (900 clinical cases) and tested against the 10% (100 clinical cases). This model is a great resource to tokenize biomedical documents, specially clinical cases written in Spanish. This model was created using the Apache… See the full description on the dataset page: https://huggingface.co/datasets/Biomedical-TeMU/SPACCC_Tokenizer.text10K<n<100K0 likes1.4k downloads5y agoHugging Face14andersonbcdefg /PD-3M-Tokenized-Cosmos-Tokenizer-DI8x8I can't get the dataset viewer to work, sorry. There's about 3M images and captions from Spawning/PD3M. They are resized and center-cropped to 512x512, and then tokenized into discrete tokens with NVIDIA Cosmos-Tokenizer-DI8x8, which reduces the spatial dimension by a factor of 8, resulting in 64 x 64 = 4096 discrete tokens per image. You can use these tokenized images to train an auto-regressive image model, or a MaskGIT. Or probably other things I don't know about. :) License is the same… See the full description on the dataset page: https://huggingface.co/datasets/andersonbcdefg/PD-3M-Tokenized-Cosmos-Tokenizer-DI8x8.text10M<n<100M0 likes1.4k downloads2y agoHugging Face15apollo-research /monology-pile-uncopyrighted-tokenizer-EleutherAI-gpt-neox-20b10M<n<100M1 likes1.3k downloads3y agoHugging Face16andersonbcdefg /PD-3M-Tokenized-Cosmos-Tokenizer-DI16x16text10M<n<100M1 likes1.3k downloads2y agoHugging Face17open-source-metrics /tokenizers-dependents tokenizers metrics This dataset contains metrics about the huggingface/tokenizers package. Number of repositories in the dataset: 11460 Number of packages in the dataset: 124 Package dependents This contains the data available in the used-by tab on GitHub. Package & Repository star count This section shows the package and repository star count, individually. Package Repository There are 14 packages that have more than 1000 stars. There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.tabularn<1K0 likes1.2k downloads2y agoHugging Face18jsun /slimpajama_Llama2_Tokenizer slimpajama_Llama2_Tokenizer The original slimpajama_Llama2_Tokenizer.tar.gz archive (≈794 GB on Ubuntu) was split into smaller 40 GB chunks for easier upload to Hugging Face. sudo apt install git-lfs pip install -U huggingface_hub # `hf version`==1.1.4 tar cvf - slimpajama_Llama2_Tokenizer/ | pigz -p 16 > slimpajama_Llama2_Tokenizer.tar.gz split -b 40G -d -a 3 slimpajama_Llama2_Tokenizer.tar.gz slimpajama_Llama2_Tokenizer/slimpajama_Llama2_Tokenizer_part_ # Upload files… See the full description on the dataset page: https://huggingface.co/datasets/jsun/slimpajama_Llama2_Tokenizer.0 likes1.2k downloads11mo agoHugging Face19apollo-research /monology-pile-uncopyrighted-tokenizer-gpt210M<n<100M1 likes1.1k downloads3y agoHugging Face20ASTERIZER /Ezaris-Tokenizer Ezaris Tokenizer — Program & Corpus The single home of the Asterizer tokenizer for the Ezaris program: one byte-level BPE tokenizer, frozen once and reused across every ASTERIZER model from 100M → 1T params. South-Indian-first (Kannada / Tamil / Telugu / Malayalam), plus code, math, and broad multilingual coverage. Built only from open, license-audited data. This repo consolidates the former tokeniser / LUNA-Tokenizer-Corpus, LUNA-Tokenizer-Gap-V2, LUNA-1B-Tokenizer, and the… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/Ezaris-Tokenizer.0 likes1k downloads1mo agoHugging Face21salmankhanpm /Corpous_Telugu_Tokenizertext1M<n<10M0 likes907 downloads1y agoHugging Face22jsun /fineweb-edu-default_Llama2_Tokenizer_incomplete0 likes847 downloads11mo agoHugging Face23catherinearnett /bilingual-tokenizer-training-datatext10M<n<100M0 likes700 downloads8mo agoHugging Face24Abzalbek89 /kk-tokenizer-fertility-baseline Kazakh Tokenizer Fertility Baseline Reproducible fertility benchmark of subword tokenizers on the Kazakh language. Companion artifact for the paper "Tokenizer Optimization for Kazakh Small Language Models" (in preparation, target: ACM TALLIP). Headline numbers Tokenizer Fertility 🥇 Best overall kk-bpe-32k 1.679 🚨 Worst GPT-4 (cl100k) 5.895 GPT-4 penalty GPT-4 (cl100k) is 3.51× worse than the best Kazakh-trained tokenizer → The custom Kazakh… See the full description on the dataset page: https://huggingface.co/datasets/Abzalbek89/kk-tokenizer-fertility-baseline.text-classificationn<1K0 likes643 downloads3mo agoHugging Face25pccl-org /Skylion007-openwebtext-tokenizer-gpt2-12810M<n<100M0 likes614 downloads2y agoHugging Face26apollo-research /roneneldan-TinyStories-tokenizer-gpt2100K<n<1M2 likes597 downloads3y agoHugging Face27hac541309 /polyglot-ko-tokenizer-corpus Dataset Card for "polyglot-ko-tokenizer-corpus" More Information needed text10M<n<100M1 likes490 downloads3y agoHugging Face28pccl-org /Skylion007-openwebtext-tokenizer-gpt2-64100M<n<1B0 likes478 downloads2y agoHugging Face29Similoluwa /african-multilingual-tokenizer-challenge African Multilingual Tokenizer Challenge dataset The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split. Split Per language Total Train 40,000 240,000 Validation 4,000 24,000 Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.texttext-generation100K<n<1M0 likes469 downloads1mo agoHugging Face30SaylorTwift /RULER-8192-Qwen2.5-3B-tokenizertabular1K<n<10K0 likes455 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.