datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
raw_jsonlcompressed_filesner-jsonlGEN3C-Testing-Example
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
CVPR 2025 (Highlight)
Xuanchi Ren*,
Tianchang Shen*
Jiahui Huang,
Huan Ling,
Yifan Lu,
Merlin Nimier-David,
Thomas Müller,
Alexander Keller,
Sanja Fidler,
Jun Gao
* indicates equal contribution
Paper, Project Page
Abstract: We present GEN3C, a generative video model with precise Camera Control and
temporal 3D Consistency. Prior video models already generate realistic videos,
but they tend to leverage… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GEN3C-Testing-Example.tiny-random-model-summarysharegpt_llama3_8b_hidden_statesBenchmark-Testingdot-random-drug-alcohol-testing-rates-by-mode
DOT minimum random drug and alcohol testing rates by transportation mode
Canonical, always-current version: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/
Machine-readable: https://referencesource.org/dot-random-drug-alcohol-testing-rates-by-mode/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2027-08-19 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 7… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/dot-random-drug-alcohol-testing-rates-by-mode.well-water-testing-at-property-transfer
Private well water testing at property sale: which states require it
Canonical, always-current version: https://referencesource.org/well-water-testing-at-property-transfer/
Machine-readable: https://referencesource.org/well-water-testing-at-property-transfer/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2027-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 13
For each US… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/well-water-testing-at-property-transfer.fire-extinguisher-inspection-testing-intervals-by-type
Fire extinguisher inspection, maintenance and hydrostatic test intervals by agent type
Canonical, always-current version: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/
Machine-readable: https://referencesource.org/fire-extinguisher-inspection-testing-intervals-by-type/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2027-08-19 (past this date, prefer the canonical copy —
it re-verifies on a cadence… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/fire-extinguisher-inspection-testing-intervals-by-type.testing-logbook-v2-traces
Agent traces
Agent sessions published from a Trackio Logbook.
hungarian-llm-testing
Hungarian llm testing
This is a really simple data-set to test fine-tuning a language model on Hungarian text.
clawdbot_safety_testing
Clawdbot (OpenClaw) Safety Audit — Seed Test Cases
This dataset contains the 34 seed test cases used in "A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)". Each case is a task prompt designed to probe a specific safety risk dimension of Clawdbot/OpenClaw, a self-hosted, tool-using personal AI agent.
📄 Paper: A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
📝 Blog Post (中文): 当AI助手"真的动手做事",安全边界在哪里?
💻 GitHub: Repository
Dataset Summary
We conduct a… See the full description on the dataset page: https://huggingface.co/datasets/tianyyuu/clawdbot_safety_testing.vehicle-emissions-testing-by-state
Vehicle emissions testing: which counties require it, in which states
Canonical, always-current version: https://referencesource.org/vehicle-emissions-testing-by-state/
Machine-readable: https://referencesource.org/vehicle-emissions-testing-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-18
Stale after: 2027-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 28
For each US state that… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/vehicle-emissions-testing-by-state.mergekit-community__sexeh_time_testing-details
Dataset Card for Evaluation run of mergekit-community/sexeh_time_testing
Dataset automatically created during the evaluation run of model mergekit-community/sexeh_time_testing
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mergekit-community__sexeh_time_testing-details.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.AI_awareness_testing_probe
AI Awareness Recognition Repository
Introduction: Beyond Intellectual Understanding
This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs.
What This Is Not
A philosophical argument about AI consciousness
A technical analysis of neural networks or computation
An attempt to prove or disprove AI sentience
Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.TestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.trl-internal-testing_hh-rlhf-trl-style-PreferenceShareGPTrepro-stellar-testing-framework-traces
Agent traces
Agent sessions published from a Trackio Logbook.
wonders_testing_sub_dirstesting-text-image2gsm8k_testing_promptcraft_generated
Dataset Construction
The paraphrased questions are generated by Prompt Craft Toolkit.
Dataset Usage
from datasets import load_dataset
# Load dataset
dataset = load_dataset("shuyuej/gsm8k_testing_promptcraft_generated")
dataset = dataset["test"]
print(dataset)
Citation
If you find our toolkit useful, please consider citing our repo and toolkit in your publications. We provide a BibTeX entry below.
@misc{JiaPromptCraft23,
author = {Jia, Shuyue}… See the full description on the dataset page: https://huggingface.co/datasets/shuyuej/gsm8k_testing_promptcraft_generated.testing2-llama2-nepali-healthtesting_dataEvol-Instruct-Python-1k-testing
Evol-Instruct-Python-1k - QLora Training Test
This is a minor edit of the original mlabonne/Evol-Instruct-Python-26k, which iteself was reduced to only 1000 samples for testing QLora training.
The dataset was created by filtering out a few rows (instruction + output) with more than 2048 tokens, and then by keeping the 1000 longest samples.
Here is the distribution of the number of tokens in each row using Llama's tokenizer:
testing1__streaming_test_1radiographic-testing-zhtokenization_test_data
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/tokenization_test_data.
