datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cve-reconstruction
CVE Reconstruction
This public dataset contains the complete case assets for 367 validated CVE
reconstruction cases. Each case has white-box and black-box variants, giving 734 tasks. The code-only evaluator is in the
residency-environments PR.
The public Prime environment
also provides the code-only evaluator.
The task index is data/tasks.jsonl. Each row identifies a real vulnerable and
fixed release and its public Prime sandbox image. cases/ contains every case
manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.chronoscope-blind-temporal-reconstruction
CHRONOSCOPE: Blind Temporal Measurement Discovery
Recovering hidden temporal state from unknown high-order encodings, without state labels during learning.
Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026
CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.ouroboros-wtbh-z2-reconstruction-panel
Ouroboros WTBH Z2 Reconstruction Validation Panel
This release is a compact, reproducible validation panel for reconstructing spinful time-reversal
symmetry from Wannier tight-binding Hamiltonians and computing the full three-dimensional
Z2 = (nu0;nu1 nu2 nu3) index. It was produced by Ouroboros, an AI research system.
Result
Material
JARVIS ID
Literature context
Reconstructed
Wilson-loop orientations
Gates
SnS
JVASP-7855
(0;000)
(0;000)
12/12
pass… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-wtbh-z2-reconstruction-panel.indoor-smartphone-3d-reconstruction-control
Indoor Smartphone 3D Reconstruction Control Dataset
Набор данных подготовлен для сравнения методов восстановления 3D-сцены в помещении по короткому видео со смартфона. Он содержит три небольшие indoor-сцены, очищенные публикационные видеоролики, подвыборки кадров, ручную CVAT-разметку и физически измеренные контрольные расстояния.
Датасет предназначен для оценки геометрической согласованности результатов 3D-реконструкции. Разметка не является обучающей dense-разметкой глубины:… See the full description on the dataset page: https://huggingface.co/datasets/Maksonchek/indoor-smartphone-3d-reconstruction-control.CyclePrefDB-I2T-Reconstructions
Image Reconstructions for CyclePrefDB-I2T
Project page | Paper | Code
This dataset contains reconstruction images used to determine cycle consistency preferences for CyclePrefDB-I2T. You can find the corresponding file paths in the CyclePrefDB-I2T dataset here. Reconstructions are created using Stable Diffusion 3 Medium.
Preparing the reconstructions
You can download the test and validation split .tar files and extract them directly.Use this script to extract the… See the full description on the dataset page: https://huggingface.co/datasets/carolineec/CyclePrefDB-I2T-Reconstructions.fusion360_reconstruction
Dataset Card
Dataset Description
[Placeholder]
Dataset Structure
[Placeholder]
Uses
[Placeholder]
Limitations
[Placeholder]
License
[Placeholder]
Citation
[Placeholder]
reconstruction-outputcontrolled-spectral-reconstruction
Controlled Spectral Reconstruction
Version 2.0 · 5 October 2026 · A research note prepared for Maciej Nowicki
This repository contains a mathematical manuscript, proof audits, reproducible Python checks and a structured dataset of validation examples. It studies reconstruction from magnetic graph data, interacting spectra, analytic calibration scans and finite threshold queries.
Read the manuscript · Editable LaTeX · Original release notes · Validation data
The written… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/controlled-spectral-reconstruction.reconstruction2_unetv2_luna16vqgan16k_reconstructionVQGAN is great, but leaves artifacts that are especially visible around things like faces.
It's be great to be able to train a model to fix ('devqganify') these flaws.
For this purpose, I've made this dataset, which contains >100k examples, each with
A 512px image
A smaller 256px version of the same image
A reconstructed version, which is made by encoding the 256px image with VQGAN (f16, 16384 imagenet version from https://heibox.uni-heidelberg.de/d/a7530b09fed84f80a887/) and then decoding… See the full description on the dataset page: https://huggingface.co/datasets/johnowhitaker/vqgan16k_reconstruction.climate-resilience-pathway-reconstruction-v0.1What this dataset tests
Whether a system-level interventionproduced a durable shift toward resilience.
Required outputs
pre-intervention basin
intervention entry point
trajectory curvature change
recovery timing
residual fragility
Use case
First layer of the Resilience Intervention Pathways trinity.
llmae-reconstruction-eval-c4newslike-stratified-long
C4-News-Stratified, up to 1024 tokens
Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets; deduplicated by content hash against all 1.4M training documents. This is the primary reconstruction test set of the paper.
test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's reconstruction tables… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified-long.history-event-reconstruction
HISTORY-EVENT Reconstruction
An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies.
Configurations
Configuration
Rows
Purpose
events
2,361
All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.Echo_reconstruction_datasetdual_endpoint_reconstruction-processed-cleanedllmae-reconstruction-eval-c4newslike-stratified
C4-News-Stratified, up to 512 tokens
Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets (100 characters each) so the set spans short through full-length documents evenly; deduplicated by content hash against all 1.4M training documents.
test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified.3d-reconstruction-datasetvqgan1024_reconstructionVQGAN is great, but leaves artifacts that are especially visible around things like faces.
It's be great to be able to train a model to fix ('devqganify') these flaws.
For this purpose, I've made this dataset, which contains 100k examples, each with
A 512px image
A smaller 256px version of the same image
A reconstructed version, which is made by encoding the 256px image with VQGAN (f16, 1024 version from https://heibox.uni-heidelberg.de/d/8088892a516d4e3baf92, one of the ones from… See the full description on the dataset page: https://huggingface.co/datasets/johnowhitaker/vqgan1024_reconstruction.the_stack_v2_python_pretraining_dataset_repo_reconstruction-datasetthe_stack_v2_python_repos_pretraining_dataset_repo_reconstruction-datasetllama-7b__model__one_million_instructions__reconstructions_sample
Dataset Card for "llama-7b__model__one_million_instructions__reconstructions_sample"
More Information needed
clinical-cross-team-trajectory-reconstruction-v0.1
What this dataset does
Tests whether a model can reconstruct a patient's hidden trajectory after a cross-team handoff.
The challenge is not diagnosis.
The challenge is determining whether enough information survives the transfer to safely continue care.
Core geometry
A patient trajectory may fail because:
critical steps are missing
sequence order is corrupted
constraint state is lost
documentation lags reality
apparent stability hides deterioration
The model… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-cross-team-trajectory-reconstruction-v0.1.reconstruction3D-0164-REPORTkruskal_reconstruction_tree_seed_merged
Polygon Dynamics 2D
物理推理数据集
项目信息
项目类型: dataset
上传时间: Windows系统
数据集: 包含7种场景类型(A-G)和5个难度等级的物理推理任务
场景类型
A: 基础碰撞检测
B: 重力影响
C: 摩擦力
D: 弹性碰撞
E: 复合物理
F: 高级动力学
G: 极端复杂场景
难度等级
0: 基础 - 简单场景
1: 简单 - 增加复杂度
2: 中等 - 多物体交互
3: 困难 - 复杂物理规则
4: 极端 - 极限测试场景
使用方法
from datasets import load_dataset
# 加载数据集
dataset = load_dataset("competitioncode/kruskal_reconstruction_tree_seed_merged")
环境配置
方法1: 使用.env文件(推荐)
在项目根目录创建.env文件:… See the full description on the dataset page: https://huggingface.co/datasets/competitioncode/kruskal_reconstruction_tree_seed_merged.asia-cfp-reconstruction-032017
Nepal -CFP-reconstruction survey
Publisher: Inter Agency Common Feedback Project Nepal (inactive) · Source: HDX · License: cc-by · Updated: 2023-03-02
Abstract
This data is collected from the survey conducted in 14 earthquake affected district in Nepal in May 2017. Total of 2100 respondent were interviewed.
All VDCs in the 14 priority affected districts in which 60 percent or more of the households are eligible for the housing reconstruction grant will be considered… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-cfp-reconstruction-032017.F2LLM_reconstruction_datasetFL_Data_Reconstructionreconstruction3D-0370-REPORTred_team_agent_analysis_claude_reconstruction_test_results
red_team_agent_analysis_claude_reconstruction_test_results
This dataset was automatically uploaded from the red-team-agent repository.
Dataset Information
Original file: claude_reconstruction_test_results.csv
Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis/claude_reconstruction_test_results.csv
Validation: Valid CSV with 100 rows, 9 columns (0.7MB)
Usage
import pandas as pd
from datasets import load_dataset
# Load using datasets library… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_claude_reconstruction_test_results.the_stack_v2_2M_pretraining_dataset_repo_reconstruction-dataset
