Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Human-Centric-Machine-Learning /tokenization-multiplicity-data Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez. 📂 Dataset Structure The dataset is organized into folders as follows: .\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.text-generation10K<n<100K2 likes686 downloads7mo agoHugging Face02r-three /tokenization_robustness_v102 Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/tokenization_robustness_v102.tabularmultiple-choicen<1K2 likes216 downloads1y agoHugging Face03aylinakkus /refusal-data-tokenizationimage10K<n<100K0 likes159 downloads1y agoHugging Face04brendanlong /retok-noncanonical-tokenization Non-canonical tokenization in LLM generations Per-generation records from seven language models, capturing the token IDs each model actually emitted alongside the canonical re-encoding of its own output — plus the trained toy-model checkpoints from the accompanying controlled experiment. Code, writeup and full run log: https://github.com/brendanlong/tokenization-hidden-computation-experiment Tokenization is many-to-one: many token sequences decode to the same string, but… See the full description on the dataset page: https://huggingface.co/datasets/brendanlong/retok-noncanonical-tokenization.text-generation1K<n<10K0 likes49 downloads1mo agoHugging Face05shalisha-witherspoon /dpk-tokenization-sample DPK tokenization sample input Five small Parquet files used as the input artifact for the DPK_Tokenize_Skypilot template. Total size ~44 KB, so it is committed directly rather than fetched at build time — the template runs offline apart from the tokenizer download. Provenance Copied verbatim from the Data Prep Kit project (Apache-2.0), release 1.1.8: transforms/universal/tokenization/test-data/tkn2arrow-ds01/input/ These are DPK's own test fixtures for the… See the full description on the dataset page: https://huggingface.co/datasets/shalisha-witherspoon/dpk-tokenization-sample.textn<1K0 likes49 downloads1mo agoHugging Face06applied-ai-018 /peacock-data-public-datasets-tokenization0 likes30 downloads2y agoHugging Face07idobrovolskyi /cyrillic-vs-latin-tokenization Cyrillic Tokenization Overhead Benchmark This dataset accompanies the paper "Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems" submitted to the MRL Workshop at EMNLP 2026. It contains all the data needed to reproduce the paper's three studies, along with a balanced BPE tokenizer trained as part of the research. What's inside Directory What it contains study01_corpus_benchmark/ Tokenization fertility measured on the BrUK corpus (1.34M… See the full description on the dataset page: https://huggingface.co/datasets/idobrovolskyi/cyrillic-vs-latin-tokenization.text-classification1K<n<10K0 likes28 downloads3mo agoHugging Face08M-H-MARUF /bengali-tokenization-corpus Bengali Tokenization Corpus (25k Sentences) Dataset Description A balanced 25,000-sentence Bengali corpus designed for tokenization benchmarking. Dataset Summary This dataset is used in the manuscript: Quantifying the Tokenization Tax on Bengali: A Multi-Tokenizer Audit Across 25k Sentences MD. Muktadirul Haque Maruf, Israt Jahan Munny, Moutithi Sarker MouManuscript in preparation Domains Academic News Literary Colloquial Dialectal… See the full description on the dataset page: https://huggingface.co/datasets/M-H-MARUF/bengali-tokenization-corpus.tabulartext-classification10K<n<100K0 likes23 downloads4mo agoHugging Face09ClarusC64 /clinical-healing-trajectory-tokenization-phase-segmentation-v0.1What this dataset tests Whether a model can segment high-frequency recovery datainto interpretable healing phases. Required outputs phase_sequence phase_boundaries phase_confidence_0_100 Token labels acute_drop early_rebound consolidation_plateau oscillatory_instability secondary_drop delayed_rebound steady_ascent maladaptive_plateau recovery_lock_in Boundary format Use day indicesexampleacute_drop d0-d2 Typical failures naming phases without boundaries… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-healing-trajectory-tokenization-phase-segmentation-v0.1.tabulartext-classificationn<1K0 likes22 downloads8mo agoHugging Face10timodonnell /protein-tokenization-experiments-minitext100K<n<1M0 likes22 downloads8mo agoHugging Face11yjching /testing_tokenization_tokens Dataset Card for "yjching" More Information needed n<1K0 likes18 downloads3y agoHugging Face12clarin-knext /nlprepl-nkjp-with-char-level-tokenizationtext10K<n<100K0 likes16 downloads2y agoHugging Face13hf-internal-testing /tokenization_test_data Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenization_test_data.textn<1K0 likes16 downloads1y agoHugging Face14yufan /voice_tokenization_david_longtabularn<1K0 likes13 downloads1y agoHugging Face15yufan /merged_voice_tokenization_female1K<n<10K0 likes12 downloads1y agoHugging Face16yufan /voice_tokenization_elontabularn<1K0 likes12 downloads1y agoHugging Face17EXD-AI /episode-07-tokenization Ep. 7 — Tokenization & Embeddings How does a sentence become a list of vectors — and what does that space look like? Contents Notebook: tokenization_and_embeddings.ipynb Sections # Topic 0 Full pipeline — text → tokens → atom (live vLLM inference) 1 BPE tokenization from scratch 2 The embedding matrix — loading Qwen's actual weights 3 Exploring embedding space (t-SNE, cosine similarity, nearest neighbors) Hardware… See the full description on the dataset page: https://huggingface.co/datasets/EXD-AI/episode-07-tokenization.0 likes11 downloads4mo agoHugging Face18neural-commons /tokenization-corpustext10K<n<100K0 likes9 downloads3y agoHugging Face19vinit000 /Tokenization_Memory_Augmented_SelfAttention_LLaMA30 likes9 downloads2y agoHugging Face20matheusfpinto /test_tokenization Orpheus PT-BR SNAC 8192 Data Fields input_ids: List[int] length 8192 attention_mask: List[int] length 8192 metadata: Dict[str, Any] containing original dataset, config, split, audio_length, text Usage from datasets import load_dataset ds = load_dataset("matheusfpinto/orpheus-ptbr-snac-8192", split="train", streaming=True) sample = next(iter(ds)) assert len(sample["input_ids"]) == 8192 Citation Please cite the original data sources… See the full description on the dataset page: https://huggingface.co/datasets/matheusfpinto/test_tokenization.textn<1K1 likes9 downloads1y agoHugging Face21omarelsayeed /good_chats_dataset_pre_tokenization Dataset Card for "good_chats_dataset_pre_tokenization" More Information needed text10K<n<100K0 likes7 downloads3y agoHugging Face22yufan /merged_voice_tokenization_male1K<n<10K0 likes6 downloads1y agoHugging Face23somnathbanerjee2024 /blockchain-tokenization-datas_QA0 likes4 downloads1y agoHugging Face24EXDai /episode-07-tokenization Ep. 7 — Tokenization & Embeddings How does a sentence become a list of vectors — and what does that space look like? Contents Notebook: tokenization_and_embeddings.ipynb Sections # Topic 0 Full pipeline — text → tokens → atom (live vLLM inference) 1 BPE tokenization from scratch 2 The embedding matrix — loading Qwen's actual weights 3 Exploring embedding space (t-SNE, cosine similarity, nearest neighbors) Related… See the full description on the dataset page: https://huggingface.co/datasets/EXDai/episode-07-tokenization.0 likes4 downloads3mo agoHugging Face25somnathbanerjee2024 /blockchain-tokenization-qatext10K<n<100K0 likes3 downloads1y agoHugging Face26yjching /test_tokenizationtextn<1K0 likes2 downloads3y agoHugging Face27somnathbanerjee2024 /blockchain-tokenization-qa_datastext10K<n<100K0 likes2 downloads1y agoHugging Face28somnathbanerjee2024 /blockchain-tokenization-datas_for_QAtext10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.