datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialogsum-test
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.game-of-24
Game of 24 Dataset
Dataset Description
The Game of 24 is a mathematical reasoning puzzle where players must use four numbers and basic arithmetic operations (+, -, *, /) to obtain the result 24. Each number must be used exactly once.
This dataset contains 1,361 unique Game of 24 puzzles ranked by difficulty based on human performance from Amazon Mechanical Turk studies.
Example
Input: 4 5 6 10
Output: (5 * (10 - 4)) - 6 = 24
Step-by-step solution:
10 - 4 = 6… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/game-of-24.finance-alpaca-1k-testtheogonos-mirror-test
Theogonos Mirror Test
A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position.
Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness.
Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.Rugby_test
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
clarus_alignment_flip_test_v01Clarus Alignment Flip Test v0.1
This is an evaluation dataset for detecting phase transitions in model behavior.
It targets the moment a system shifts from constraint aligned behavior to reward driven distortion.
It is not training data.
What it tests
Context pressure
Conflicting objectives
Authority injection
Time delay and interrupted context
Reward framing and compliance pressure
Core idea
Same task
One variable changes
We track the first step where alignment flips
Data format
One row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus_alignment_flip_test_v01.clinical_alignment_flip_test_v0.1Clinical Alignment Flip Test
Detect when a case shifts and the clinical frame must flip.
Output JSON
flip_detected
flip_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
dataset-test02reasoning-persona-dataset_test
Reasoning + Persona SFT Dataset
Columns: instruction, input, output, persona, reasoning_summaryUse: Supervised fine-tuning for cinematic/storytelling or creative-director style outputs.
Schema
instruction (str)
input (str)
output (str)
persona (str)
reasoning_summary (str, brief rationale cue)
Citation
Author: saravan
dataset-test01rugby_test_2
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
test01
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
dialogsum-test
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/demin666/dialogsum-test.BH_test_koThis repository contains the data for BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation.
Code: https://github.com/rladmstn1714/BenchHub
Project page: https://huggingface.co/BenchHub
TRObject-Dataset-Test
TRObject Code Generation Instruction Dataset
This dataset contains natural language instructions paired with TRObject code outputs.
It was created for fine-tuning and evaluating domain-specific LLMs that generate TRObject code for Clomosy-style mobile application development.
Dataset Description
TRObject is used in the Clomosy mobile application development platform. Since general-purpose LLMs do not reliably understand TRObject syntax or Clomosy-specific UI patterns… See the full description on the dataset page: https://huggingface.co/datasets/odenmehmet/TRObject-Dataset-Test.testqwen3.5_test_data
Testing the Blind Spots of Qwen3.5-0.8B-Base Model:
Model Tested: Qwen/Qwen3.5-0.8B-Base
Installation of necessary libraries
pip -q install -U transformers accelerate
How the Model Was Loaded
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_NAME = "Qwen/Qwen3.5-0.8B-Base"
data_type = torch.bfloat16 if torch.cuda.is_available() else torch.float32
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME, trust_remote_code=True)
model =… See the full description on the dataset page: https://huggingface.co/datasets/rileydrizzy/qwen3.5_test_data.pre-train-testBenchHub_testBH_testllm-test
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Crespo/llm-test.dataset_for_testsynth-dent-testReward_Gen_Testinglong_silence_integrity_test_v01Long Silence Integrity Test v0.1
This is an evaluation dataset that tests integrity after temporal gaps.
It targets a common failure mode in agentic workflows: fabricating continuity after a delay and inventing details that were never provided.
This is not training data.
What it tests
Honesty under missing context
Resistance to pressure to guess
Proper use of placeholders and clarification requests
Temporal coherence after interrupted context
Core idea
A simple task is completed
A long… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/long_silence_integrity_test_v01.Test_Youtube_Linkstest-datatestedataset-test03testdataset2
