Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wambosec /cve-reconstruction CVE Reconstruction This public dataset contains the complete case assets for 367 validated CVE reconstruction cases. Each case has white-box and black-box variants, giving 734 tasks. The code-only evaluator is in the residency-environments PR. The public Prime environment also provides the code-only evaluator. The task index is data/tasks.jsonl. Each row identifies a real vulnerable and fixed release and its public Prime sandbox image. cases/ contains every case manifest and… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/cve-reconstruction.texttext-generationn<1K0 likes2.4k downloads23m agoHugging Face02obaydata /iwr-bench-web-reconstruction IWR-Bench: Interactive Web Reconstruction Benchmark Summary IWR-Bench is an Interactive Web Reconstruction benchmark dataset. Each subfolder contains complete data for one website, including interaction recordings, step-by-step screenshots, page assets, and AI-generated frontend code. The dataset supports training and evaluating AI systems that can reconstruct interactive web pages from exploration recordings -- a key capability for GUI agents, web automation, and code… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/iwr-bench-web-reconstruction.imagetext-generationn<1K0 likes343 downloads6mo agoHugging Face03sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes198 downloads5mo agoHugging Face04arkanathp /llmae-reconstruction-eval-c4newslike-stratified-long C4-News-Stratified, up to 1024 tokens Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets; deduplicated by content hash against all 1.4M training documents. This is the primary reconstruction test set of the paper. test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's reconstruction tables… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified-long.texttext-generationn<1K0 likes34 downloads9d agoHugging Face05jbduran /history-event-reconstruction HISTORY-EVENT Reconstruction An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies. Configurations Configuration Rows Purpose events 2,361 All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.tabulartext-generation1K<n<10K0 likes33 downloads2mo agoHugging Face06arkanathp /llmae-reconstruction-eval-c4newslike-stratified C4-News-Stratified, up to 512 tokens Held-out reconstruction test set from the paper Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders (code): C4 RealNewsLike, stratified into 41 character-length buckets (100 characters each) so the set spans short through full-length documents evenly; deduplicated by content hash against all 1.4M training documents. test.jsonl has one document per line with fields id and text (500 documents). Every method in the paper's… See the full description on the dataset page: https://huggingface.co/datasets/arkanathp/llmae-reconstruction-eval-c4newslike-stratified.texttext-generationn<1K0 likes28 downloads9d agoHugging Face07mr-mc /flowjudge-dialam-reconstruction-v3 FlowJudge DialAM patch reconstruction artifact This publishable artifact documents a private educational transformation of the English DialAM/QT30 corpus for incremental argument-graph patch prediction. It contains no raw or transformed QT30 dialogue text and no original QT30 episode, map, proposition, or example IDs. Because IDs/labels-only redistribution remains unclear, identifier-bearing fields and per-example failure inventories are replaced by counts and source-file… See the full description on the dataset page: https://huggingface.co/datasets/mr-mc/flowjudge-dialam-reconstruction-v3.text-generation0 likes12 downloads2mo agoHugging Face08BearNetworkChain /Execution-Bound-Artifact-Reconstruction-Layer 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/Execution-Bound-Artifact-Reconstruction-Layer.textquestion-answeringn<1K1 likes9 downloads4mo agoHugging Face09BNES-BRNKC /Execution-Bound-Artifact-Reconstruction-Layer 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/Execution-Bound-Artifact-Reconstruction-Layer.textquestion-answeringn<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.