datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.TestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.testing
RubricHub_v1
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of coarse or static rubrics.… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/testing.Testing
Dataset Card (Qwen3.5-reasoning-700x)
Dataset Summary
Qwen3.5-reasoning-700x is a high-quality distilled dataset.
This dataset uses the high-quality instructions constructed by Alibaba-Superior-Reasoning-Stage2 as the seed question set. By calling the latest Qwen3.5-27B full-parameter model on the Alibaba Cloud DashScope platform as the teacher model, it generates high-quality responses featuring long-text reasoning processes (Chain-of-Thought). It covers several major… See the full description on the dataset page: https://huggingface.co/datasets/ApeMaster/Testing.testing-wiki-structured
cywiki_namespace_0
Structured Contents snapshot of cywiki_namespace_0 from the
Wikimedia Enterprise API,
repackaged as Parquet with a pinned schema.
The upstream Wikimedia Foundation dataset
(wikimedia/structured-wikipedia)
ships NDJSON which has known issues loading via
datasets.load_dataset() — see discussions
#5,
#15,
#16.
This dataset is the same upstream content, normalised so
load_dataset(...)works without specifying a Features override.
Source
Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.AIForge-1K-Testing
AIForge-04-Testing
Testing Dataset for AI and Programming Tasks
Overview
AIForge-04-Testing is a curated English dataset designed for AI systems working on testing tasks in software engineering and programming.
Contents
data.jsonl
data.json
metadata.json
Use Cases
AI agent training
Supervised fine-tuning
Evaluation and benchmarking
Software engineering research
Example Record
{
"id": "AITST_00001",
"category":… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/AIForge-1K-Testing.empathetic_dialogues
Dataset Card for "empathetic_dialogues"
Dataset Summary
PyTorch original implementation of Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 28.02 MB
Size of the generated dataset: 25.13 MB
Total amount of disk used: 53.15… See the full description on the dataset page: https://huggingface.co/datasets/testingtest111/empathetic_dialogues.testingtesting_arc_easy_de
testing_arc_easy_de
This is a copy of the dataset openGPT-X/arcx.
Reward_Gen_Testingtesting
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1NightRaid1/testing.testingeval_testing_mns
