datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Debunk_Traffic_Representation
Packet-level classification: classify based on packet
Per-packet-split: Mix all packets and split them into train, val, and test sets, based on 8:1:1
Per-flow-split: Split the pcap files based on 5-tuples (src_IP, dst_IP, src_port, dst_port, and protocol), using 3-fold validation, and there is no intersection between the train, val, and test sets.
Flow-level classification: classify based on flow
spacehpc-representation-scaling-k0-v1quantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.qiskit-representation-search-alphamscoco-hidden-representationvdr-jina-v4-layer-representationp2-etf-representation-disentanglement-rl-resultsFuseChat-Mixture-NH-2-SOLAR-10.7B-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-NH-2-SOLAR-10.7B-Representation.World-Embedding-Optics-Retrieval
World Embedding Optics Retrieval
World Embedding Optics Retrieval is the optics and electromagnetism retrieval split of the
World Embedding Benchmark. It contains
2,700 simulation videos from 27 physics families. The family shards are loaded
together as the default configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Optics-Retrieval",
split="test",
)
Fields
query_id:… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Optics-Retrieval.representation-of-equivalent-math-raw
Representation of Equivalent Math as Math-Capability Predictors — raw data
Full raw data for the study at
https://github.com/toolanzyhhh1234/representation-of-equivalent-math-as-math-capability-predictors
(preliminary technical report REPORT.md there; every number in it regenerates from
this data plus the pinned code).
Study question: does the geometry of an LLM's internal representation of
mathematically equivalent statements predict its mathematical capability? Current… See the full description on the dataset page: https://huggingface.co/datasets/toolazyhhh123/representation-of-equivalent-math-raw.FuseChat-Mixture-InternLM2-Chat-20B-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the original GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-InternLM2-Chat-20B-Representation.World-Embedding-Solid-Retrieval
World Embedding Solid Retrieval
World Embedding Solid Retrieval is the solid mechanics retrieval split of the
World Embedding Benchmark. It contains
2,700 simulation videos from 27 physics families. The family shards are loaded
together as the default configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Solid-Retrieval",
split="test",
)
Fields
query_id: unique text-query… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Solid-Retrieval.FuseChat-Mixture-Qwen1.5-Chat-72B-Aligned-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the original GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-Qwen1.5-Chat-72B-Aligned-Representation.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.risk-representation-llm-data
Risk Representation in LLM Activations
Full dataset from a cross-model study testing whether language models have internal representations of risk attitude separable from stimulus encoding.
Report: Interactive results page
Models Tested
Qwen2.5-7B-Instruct (28 layers, 3584 hidden dim, rev a09a3545)
Mistral-7B-Instruct-v0.3 (32 layers, 4096 hidden dim, rev c170c708)
Key Finding
The Representation Paradox: Neither model has orderly behavioral risk… See the full description on the dataset page: https://huggingface.co/datasets/chengruiqu/risk-representation-llm-data.World-Embedding-Dynamics-Retrieval
World Embedding Dynamics Retrieval
World Embedding Dynamics Retrieval is the dynamics retrieval split of the
World Embedding Benchmark. It contains
1,900 simulation videos from 19 physics families. The family shards are loaded
together as the default configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Dynamics-Retrieval",
split="test",
)
Fields
query_id: unique… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Dynamics-Retrieval.spacehpc-representation-capacity-full24-models-v1deep-multimodal-representation-learning-for-stellar-spectraDataset used in the paper "Deep Multimodal Representation Learning for Stellar Spectra".
Dataset of Milky Way stars based on selection from https://ui.adsabs.harvard.edu/abs/2024A&A...682A...9G,
which is based on ESA/Gaia/DPAC and APOGEE surveys.
This work has made use of data from the European Space Agency (ESA) mission Gaia (https://www.cosmos.esa.int/gaia),
processed by the Gaia Data Processing and Analysis Consortium (DPAC, https://www.cosmos.esa.int/web/gaia/dpac/consortium).
Funding… See the full description on the dataset page: https://huggingface.co/datasets/christianschwarz/deep-multimodal-representation-learning-for-stellar-spectra.spacehpc-representation-capacity-full24-reanchor-models-v13_4_fusechat_v1_openchat-3.5_mixtral-8x7b-instruct-v0.1_solar-10.7b-instruct-v1.0_representationWorld-Embedding-Regression
World Embedding Regression
World Embedding Regression is the physical-property regression dataset from the
World Embedding Benchmark. It contains
5,000 simulation videos across 10 physics families, with 500 examples in each
dataset configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Regression",
"pendulum",
split="test",
)
Fields
family: physics family.
id: unique… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Regression.World-Embedding-Fluid-Retrieval
World Embedding Fluid Retrieval
World Embedding Fluid Retrieval is the fluid mechanics retrieval split of the
World Embedding Benchmark. It contains
700 simulation videos from 7 physics families. The family shards are loaded
together as the default configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Fluid-Retrieval",
split="test",
)
Fields
query_id: unique text-query… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Fluid-Retrieval.persona-representation-corpus
Persona Representation Corpus
A reproducible corpus of persona representations — how LLM benchmarks, role-play datasets,
dialogue/personalization corpora, and synthetic-user sources actually encode "a persona" —
extracted from 19 sources spanning 6 representation families.
Built for the survey "What is a persona representation?"
Files
File
Contents
personas.jsonl
39,459 normalized records: source, source_repo, family, domain, generation, paper… See the full description on the dataset page: https://huggingface.co/datasets/Dwootton/persona-representation-corpus.tabemb-representations
TabEmb Representations
Precomputed LLM column embeddings used by the TabEmb pipeline
for column type annotation (CTA) and column property annotation (CPA).
TODO before publishing: fill in the sections below, and confirm the licensing
terms of each source benchmark permit redistributing derived embeddings
(SOTAB, T2D, WebTables, WikiTables each have their own license/citation
requirements).
Contents
Each subdirectory follows <data_dir_name>/<task>/ (task is cta or… See the full description on the dataset page: https://huggingface.co/datasets/ehoseinz/tabemb-representations.voxel-representationStreetscape-Representation
Streetscape Representation
Research materials for Photographic representation sensitivity in multimodal streetscape auditing: a matched-panorama study in Singapore.
The release contains the 739-location source inventory, executable input-construction and inference code, 25,423 formal model-output records, analysis tables, seven main and six supplementary scientific figures, and reproducible validation modules.
The five configurations are local Qwen3.5-122B, Qwen3.5-397B and… See the full description on the dataset page: https://huggingface.co/datasets/kingmin/Streetscape-Representation.cable-representation-analysis
Cable representation analysis — ACT encoder, DG-5F / UR5e cable sorting
What the ACT encoder's internal representation holds about the cable being
handled. All figures come from a forward hook on policy.model.encoder: the
transformer tokens are mean-pooled into a single 512-d vector per frame.
Every split in this directory is episode-level. Frames inside one episode
are near-duplicates, so a frame-level split leaks the answer and inflates every
number below.
Model… See the full description on the dataset page: https://huggingface.co/datasets/Kaz55/cable-representation-analysis.memory-representation-contextbench-traces
Memory Representation ContextBench Raw Traces
This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content.
OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json.
Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.FuseChat-Mixture-OpenChat-3.5-7B-Representation
Dataset Card for FuseChat-Mixture
Dataset Description
FuseChat-Mixture is the training dataset used in 📑FuseChat: Knowledge Fusion of Chat Models
FuseChat-Mixture is a comprehensive training dataset covers different styles and capabilities, featuring both human-written and model-generated, and spanning general instruction-following and specific skills. These sources include:
Orca-Best: We sampled 20,000 examples from Orca-Best, which is filtered from the… See the full description on the dataset page: https://huggingface.co/datasets/FuseAI/FuseChat-Mixture-OpenChat-3.5-7B-Representation.nav2
