Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes321 downloads1y agoHugging Face02superviselab /multimodal-video-annotation-samples Video Annotation Samples – SuperviseLab SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates our annotation methodology and output quality across diverse video content categories. Note: All visual assets in this dataset have been abstracted (pixelated mosaic) to protect source privacy. Uploader identity, original titles, and all identifiable metadata have been removed. This is a demonstration dataset… See the full description on the dataset page: https://huggingface.co/datasets/superviselab/multimodal-video-annotation-samples.tabularvideo-classificationn<1K1 likes173 downloads6mo agoHugging Face03OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes83 downloads29d agoHugging Face04Taxonomy-Aligned-Conversational-Tutor /TACTBench-Samples TACTBench Demonstration Samples This repository contains five full-context demonstration examples from TACTBench. It does not contain the TACT training set or the remaining hidden TACTBench evaluation set. The samples use the same full-history representation as the benchmark evaluation and illustrate direct correction, error explanation, guided revision, clarification checking, affective feedback, and retry elicitation. Data data/demo.jsonl: five complete… See the full description on the dataset page: https://huggingface.co/datasets/Taxonomy-Aligned-Conversational-Tutor/TACTBench-Samples.tabulartext-generationn<1K0 likes64 downloads17d agoHugging Face05wealthschema /household-samples WealthSchema Synthetic Household Samples 8 synthetic U.S. households, one per life stage plus one high-net-worth: a small free preview of what a complete, internally consistent household financial profile looks like. Each record covers the people, income, assets, debts, insurance, taxes, goals and a monthly trajectory. No real person is behind any of it. Built for teams that need realistic households to design, demo or test financial software: planning tools, robo-advisors… See the full description on the dataset page: https://huggingface.co/datasets/wealthschema/household-samples.tabulartabular-classificationn<1K0 likes50 downloads16d agoHugging Face06CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes39 downloads4mo agoHugging Face07Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes34 downloads1y agoHugging Face08alirezaaminzadeh /meetscribe-meeting-samples MeetScribe Meeting Samples Synthetic bilingual (EN/FA) enterprise meeting transcripts with labeled action items. File Language Domain operations_review_en EN Production / maintenance operations_review_en.json EN JSON ASR (Whisper format) safety_board_fa FA HSE safety board procurement_sync_en EN Procurement / RFQ maintenance_planning_fa FA Maintenance planning Usage python scripts/build_dataset.py Generates meetings.jsonl with extracted… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/meetscribe-meeting-samples.tabularsummarizationn<1K0 likes28 downloads2mo agoHugging Face09psdn-ai /code-workflow-samplesgated Code Workflow Samples This sample shows paired developer workflow examples for reviewing prompt, code, test, error, and output structure before scoping a larger code dataset. What This Shows Input-output pairs from practical coding workflows Metadata for task type, files, outputs, and review context A compact view of schema consistency for code-centric examples Dataset Specifications Field Value Modality Code I/O pairs Domain… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/code-workflow-samples.tabulartext-generationn<1K1 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.