Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /raw_jsonltext10K<n<100K0 likes28k downloads5y agoHugging Face02hf-internal-testing /compressed_filestextn<1K0 likes11k downloads5y agoHugging Face03hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face04nvidia /GEN3C-Testing-Example GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control CVPR 2025 (Highlight) Xuanchi Ren*, Tianchang Shen* Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas Müller, Alexander Keller, Sanja Fidler, Jun Gao * indicates equal contribution Paper, Project Page Abstract: We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GEN3C-Testing-Example.videon<1K4 likes492 downloads1y agoHugging Face05hf-internal-testing /tiny-random-model-summarytextn<1K0 likes283 downloads4y agoHugging Face06nm-testing /sharegpt_llama3_8b_hidden_statestextn<1K0 likes176 downloads10mo agoHugging Face07shounakpaul95 /Benchmark-Testingtabulartext-classification100K<n<1M0 likes132 downloads2y agoHugging Face08referencesource /dot-random-drug-alcohol-testing-rates-by-mode DOT minimum random drug and alcohol testing rates by transportation mode Canonical, always-current version: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/ Machine-readable: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-06 Stale after: 2027-08-19 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 7… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-random-drug-alcohol-testing-rates-by-mode.textn<1K0 likes80 downloads4d agoHugging Face09referencesource /well-water-testing-at-property-transfer Private well water testing at property sale: which states require it Canonical, always-current version: https://referencesource.org/well-water-testing-at-property-transfer/ Machine-readable: https://referencesource.org/well-water-testing-at-property-transfer/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-06 Stale after: 2027-08-18 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 13 For each US… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/well-water-testing-at-property-transfer.textn<1K0 likes79 downloads4d agoHugging Face10referencesource /fire-extinguisher-inspection-testing-intervals-by-type Fire extinguisher inspection, maintenance and hydrostatic test intervals by agent type Canonical, always-current version: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/ Machine-readable: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-19 Stale after: 2027-08-19 (past this date, prefer the canonical copy — it re-verifies on a cadence… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/fire-extinguisher-inspection-testing-intervals-by-type.textn<1K0 likes68 downloads3d agoHugging Face11abidlabs /testing-logbook-v2-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes67 downloads3mo agoHugging Face12Bazsalanszky /hungarian-llm-testing Hungarian llm testing This is a really simple data-set to test fine-tuning a language model on Hungarian text. textn<1K5 likes54 downloads3y agoHugging Face13tianyyuu /clawdbot_safety_testing Clawdbot (OpenClaw) Safety Audit — Seed Test Cases This dataset contains the 34 seed test cases used in "A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)". Each case is a task prompt designed to probe a specific safety risk dimension of Clawdbot/OpenClaw, a self-hosted, tool-using personal AI agent. 📄 Paper: A Trajectory-Based Safety Audit of Clawdbot (OpenClaw) 📝 Blog Post (中文): 当AI助手"真的动手做事",安全边界在哪里? 💻 GitHub: Repository Dataset Summary We conduct a… See the full description on the dataset page: https://huggingface.co/datasets/tianyyuu/clawdbot_safety_testing.texttext-classificationn<1K0 likes52 downloads8mo agoHugging Face14referencesource /vehicle-emissions-testing-by-state Vehicle emissions testing: which counties require it, in which states Canonical, always-current version: https://referencesource.org/vehicle-emissions-testing-by-state/ Machine-readable: https://referencesource.org/vehicle-emissions-testing-by-state/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-18 Stale after: 2027-08-18 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 28 For each US state that… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/vehicle-emissions-testing-by-state.textn<1K0 likes52 downloads10d agoHugging Face15open-llm-leaderboard /mergekit-community__sexeh_time_testing-detailsgated Dataset Card for Evaluation run of mergekit-community/sexeh_time_testing Dataset automatically created during the evaluation run of model mergekit-community/sexeh_time_testing The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mergekit-community__sexeh_time_testing-details.tabular10K<n<100K0 likes50 downloads2y agoHugging Face16Testing333555 /harmonic-reasoning-v1 Harmonic Reasoning v1 Support This Work I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases. Support on Ko-fi Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.tabulartext-generationn<1K1 likes50 downloads6mo agoHugging Face17PratikGautam /AI_awareness_testing_probe AI Awareness Recognition Repository Introduction: Beyond Intellectual Understanding This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs. What This Is Not A philosophical argument about AI consciousness A technical analysis of neural networks or computation An attempt to prove or disprove AI sentience Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.texttext-generationn<1K1 likes46 downloads1y agoHugging Face18Naga1289 /TestingDataset SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction This dataset contains multimodal question-answering examples grounded in textbook figures. Records in the figure-grounded configurations are filtered to include only examples whose referenced image files are present in this release. Configurations visual: 13791 figure-grounded visual questions with resolved images. knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.imagevisual-question-answering10K<n<100K0 likes30 downloads5mo agoHugging Face19PJMixers /trl-internal-testing_hh-rlhf-trl-style-PreferenceShareGPTtextreinforcement-learning100K<n<1M1 likes29 downloads2y agoHugging Face20noxeon /repro-stellar-testing-framework-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes29 downloads2mo agoHugging Face21Mostafa3zazi /wonders_testing_sub_dirsimagen<1K0 likes25 downloads2y agoHugging Face22abidlabs /testing-text-image2imagen<1K0 likes23 downloads3y agoHugging Face23shuyuej /gsm8k_testing_promptcraft_generated Dataset Construction The paraphrased questions are generated by Prompt Craft Toolkit. Dataset Usage from datasets import load_dataset # Load dataset dataset = load_dataset("shuyuej/gsm8k_testing_promptcraft_generated") dataset = dataset["test"] print(dataset) Citation If you find our toolkit useful, please consider citing our repo and toolkit in your publications. We provide a BibTeX entry below. @misc{JiaPromptCraft23, author = {Jia, Shuyue}… See the full description on the dataset page: https://huggingface.co/datasets/shuyuej/gsm8k_testing_promptcraft_generated.text10K<n<100K1 likes22 downloads3y agoHugging Face24Chhabi /testing2-llama2-nepali-healthtext10K<n<100K0 likes19 downloads3y agoHugging Face25semran1 /testing_datatabular100K<n<1M0 likes19 downloads10mo agoHugging Face26ptoro /Evol-Instruct-Python-1k-testing Evol-Instruct-Python-1k - QLora Training Test This is a minor edit of the original mlabonne/Evol-Instruct-Python-26k, which iteself was reduced to only 1000 samples for testing QLora training. The dataset was created by filtering out a few rows (instruction + output) with more than 2048 tokens, and then by keeping the 1000 longest samples. Here is the distribution of the number of tokens in each row using Llama's tokenizer: text1K<n<10K0 likes18 downloads3y agoHugging Face27usagent100 /testing1text1M<n<10M0 likes17 downloads2y agoHugging Face28distilabel-internal-testing /__streaming_test_1tabularn<1K0 likes17 downloads2y agoHugging Face29Obsismc /radiographic-testing-zhtexttext-generation1K<n<10K0 likes17 downloads1y agoHugging Face30hf-internal-testing /tokenization_test_data Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenization_test_data.textn<1K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.