datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aomt-world-representation-v3
AOMT World Representation Dataset v3
This release contains trajectory-derived world-representation examples from
ScienceWorld and ALFWorld. It combines offline reconstruction views with
simulator-materialized alternative-action branches for studying causal,
counterfactual, recovery, invariance, and output-format behavior.
Splits
The development pool was split deterministically by whole trajectory,
stratified within domain and task family. Related views and… See the full description on the dataset page: https://huggingface.co/datasets/Joshyxwa/aomt-world-representation-v3.fluid-reasoning-representation-phase1
Fluid Reasoning Representation - Phase 1 Multi-Model + Cross-Domain Sweep
Phase 1 artifacts for the ARR 2026 rebuttal of Fluid Reasoning Representation
(Hook et al.). This dataset extends the original QwQ x Mystery Blocksworld
study with:
Second large reasoning model: Llama-3.3-Nemotron-Super-49B-v1
Two new domains: Mystery Logistics (PDDL Logistics with obfuscated
action / predicate vocabulary) and GSM8K-Renamed (math word problems with
surface noun + verb obfuscation).
C3 causal… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/fluid-reasoning-representation-phase1.representational-collapse-llm-benchmark
Token Repetition Attack Benchmark
Dataset Description
This dataset contains experimental results from token repetition attacks on Large Language Models (LLMs), demonstrating multiple failure modes including prompt extraction, hallucination attacks, and instruction-following degradation.
Paper: Representational Collapse in Large Language Models: Token Repetition Attacks Reveal Multiple Failure Modes (ACL 2025 Submission)
Dataset Summary
Total Records: 35
Models… See the full description on the dataset page: https://huggingface.co/datasets/wu981526092/representational-collapse-llm-benchmark.
