Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-validation FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-validation.tabulartext-classification10K<n<100K1 likes367 downloads2y agoHugging Face02Salesforce /lalm-judge-validation-full-duplex LALM Judge Validation on Full-Duplex Voice Agents Companion dataset for the paper A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents. This repository contains the anonymised ratings, adversarial-defect recall tables, JSON schemas, and analysis scripts used to produce every headline number, table, and figure in that paper. Summary 209 rated stereo sessions: 152 full-duplex agent-client conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.tabularaudio-classification1K<n<10K2 likes93 downloads3mo agoHugging Face03neogenesislab /whylab-gemini-2-5-docker-validation 🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced. DOI This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.tabularothern<1K0 likes34 downloads5mo agoHugging Face04Minuri /sinhala-validation-set-10k Sinhala Validation Set - 10K Sentences A held-out Sinhala validation set of 10,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for monitoring validation loss during continual pretraining of three LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This validation set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased validation loss… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-validation-set-10k.tabulartext-generation10K<n<100K0 likes15 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.