datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cve-reconstruction
CVE Reconstruction
This public dataset contains the complete case assets for 367 validated CVE
reconstruction cases. Each case has white-box and black-box variants, giving 734 tasks. The code-only evaluator is in the
residency-environments PR.
The public Prime environment
also provides the code-only evaluator.
The task index is data/tasks.jsonl. Each row identifies a real vulnerable and
fixed release and its public Prime sandbox image. cases/ contains every case
manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.iwr-bench-web-reconstruction
IWR-Bench: Interactive Web Reconstruction Benchmark
Summary
IWR-Bench is an Interactive Web Reconstruction benchmark dataset. Each subfolder contains complete data for one website, including interaction recordings, step-by-step screenshots, page assets, and AI-generated frontend code.
The dataset supports training and evaluating AI systems that can reconstruct interactive web pages from exploration recordings -- a key capability for GUI agents, web automation, and code… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/iwr-bench-web-reconstruction.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.llmae-reconstruction-eval-c4newslike-stratified-long
C4-News-Stratified, up to 1024 tokens
Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets; deduplicated by content hash against all 1.4M training documents. This is the primary reconstruction test set of the paper.
test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's reconstruction tables… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified-long.history-event-reconstruction
HISTORY-EVENT Reconstruction
An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies.
Configurations
Configuration
Rows
Purpose
events
2,361
All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.llmae-reconstruction-eval-c4newslike-stratified
C4-News-Stratified, up to 512 tokens
Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets (100 characters each) so the set spans short through full-length documents evenly; deduplicated by content hash against all 1.4M training documents.
test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified.flowjudge-dialam-reconstruction-v3
FlowJudge DialAM patch reconstruction artifact
This publishable artifact documents a private educational transformation of the
English DialAM/QT30 corpus for incremental argument-graph patch prediction. It
contains no raw or transformed QT30 dialogue text and no original QT30
episode, map, proposition, or example IDs. Because IDs/labels-only
redistribution remains unclear, identifier-bearing fields and per-example
failure inventories are replaced by counts and source-file… See the full description on the dataset page: https://huggingface.co/datasets/mr-mc/flowjudge-dialam-reconstruction-v3.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Execution-Bound-Artifact-Reconstruction-Layer.Execution-Bound-Artifact-Reconstruction-Layer
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Execution-Bound-Artifact-Reconstruction-Layer.
