datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-like-attention-framework-1.3b-untuned-validation
Quantum-Like Attention Framework (QLAF) 1.3B Untuned Pretraining & Scaling Proof
This repository hosts the pretraining checkpoints, scaling logs, and downstream evaluation benchmarks for the 1.3B parameter Quantum-Like Attention Framework (QLAF) with Hybrid FlashAttention (75% recurrent QLAF / 25% causal FlashAttention) across a 3-seed validation campaign on dedicated A100 Large GPU hardware.
🏆 Multi-Seed Pretraining & Downstream Evaluation Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.Unified_Agent_Framework
A Unified Framework for the Evaluation of LLM Agentic Capabilities
This repository contains the dataset (Benchmark, Toolkit, and Environment assets) for the paper A Unified Framework for the Evaluation of LLM Agentic Capabilities.
The official code and agent execution sandbox can be found on GitHub: whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities.
Dataset Description
The dataset integrates diverse agent benchmarks into a standardized… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.across_framework
ACROSS: A Deformation-Based Cross-Modal Representation for Robotic Tactile Perception
Accepted to 2025 IEEE Conference on Robotics and Automation (ICRA 2025)
Paper page can be found here.
Github Repository can be found here.
DIGIT BioTac Isaac Gym.
This package contains the dataset for the BioTac to Digit pipeline. It includes over 155K unique 3D mesh deformation pairs from interactions involving BioTac and DIGIT sensors. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/wzaielamri/across_framework.gspc-regulatory-framework
GSPC — regulatory framework facts (RegimeFacts)
In one line: For the 16 XRPL issuer accounts in the live reader: is the governing regime (NYDFS, MiCA, BACEN, Reg D and others) declared and confirmable? PASS, FAIL or UNCHECKABLE. For stablecoin and tokenised-asset analysts. It confers no regulatory status.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-regulatory-framework", split="train")
print(ds[0])
Verify a signed card in your… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-regulatory-framework.HUGGER-Unified-Gravity-Fluid-Framework
🌍 H.U.G.G.E.R
[English]
1. Unified Integration
Notice: H.U.G.G.E.R and the Macroscopic Gravity studies have been integrated into the overarching Grand Unified Framework.
This repository serves as the Macroscopic Gravitational Backbone, perfectly entangled with the microscopic Topological Zero Tensor (TZT) to resolve non-linear computational collapse in planetary-scale models.
2. Core Tensor Architecture
Gravito-Fluidic Continuum Mechanics:… See the full description on the dataset page: https://huggingface.co/datasets/jskresearch/HUGGER-Unified-Gravity-Fluid-Framework.PTB-XLReal-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.edge-aui-framework-data
Dataset Card: Edge-Native Adaptive UI Behavioral Logs
This repository stores aggregated behavioral interaction logs in raw and processed forms, primarily in the .parquet data storage format.
These logs support research on edge-native adaptive user interfaces.
1. Hosted Data
The repository hosts microtensor parquet files derived from public research datasets and runtime experiment sessions:
AdSERP Search and Interaction Logs (Arapakis et al., 2025).
High-Volume… See the full description on the dataset page: https://huggingface.co/datasets/T40/edge-aui-framework-data.Unified_Agent_Evaluation_Frameworkllm-ethical-framework
Probing LLM Ethics — Derived Cache + Source Datasets
This dataset accompanies the paper "How Do LLMs Distinguish Normative Ethical
Frameworks Internally?" (ICML 2026 Mech Interp Workshop; ARR/EACL 2026 submission).
The companion code repo: https://github.com/alunxu/probing-LLM-ethics
Structure
results/ — derived cache (persona vectors, causal steering JSONs,
LLM-judge generations, layer sweep metrics, ...) consumed
by… See the full description on the dataset page: https://huggingface.co/datasets/alunxu/llm-ethical-framework.robotwin-vla-framework-eval-chunk30-20260806
RoboTwin VLA Framework Evaluation
This dataset contains the five-task RoboTwin evaluation batch run from the develop branch of vla_framework at commit 6426980. The results include per-episode videos when available, tick-level JSONL traces, resolved configurations, run manifests, and constraint results.
Batch
Source: /home/walle/perception/deploy/vla_framework on 195-walle
Branch: develop
Evaluation mode: action chunk size 30
Batch date: 2026-08-06… See the full description on the dataset page: https://huggingface.co/datasets/arrow-hf/robotwin-vla-framework-eval-chunk30-20260806.GEO-Framework
NobleJackal GEO Framework
A practical framework for making organisations clear, verifiable and citable in AI search
GEO means Generative Engine Optimization: the work of helping generative search and answer systems find, understand and support claims about organisations, people and content. This six-language book provides a seven-layer method for auditing entity clarity, evidence quality, machine-readable structure, question coverage, multilingual parity and… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GEO-Framework.CEP-IP_Framework
CEP-IP: An Explainable Framework for Cell Subpopulation Identification in Single-cell Transcriptomics (by Kah Keng Wong) (Published in Computer Methods and Programs in Biomedicine)
🧬 Abstract
Background and objective: Single-cell RNA sequencing (scRNA-seq) frameworks lack explainable approaches for identifying cell subpopulations harboring strong pairwise monotonic gene-module relationships between a gene of interest (GOI) and its co-expressed genes. In this study… See the full description on the dataset page: https://huggingface.co/datasets/kahkengwong/CEP-IP_Framework.Core_Emotion_Framework_Expansion
Core Emotion Framework (CEF)
The Core Emotion Framework (CEF) is a formally defined theoretical model in affective science that describes human emotional processes using a small set of core computational mechanisms. It is recognized as a distinct research framework used in psychology, computational modeling, neuroscience, and AI–emotion systems.
CEF is a theoretical entity, not a commercial product, and not associated with any single institution or dataset maintainer.The dataset… See the full description on the dataset page: https://huggingface.co/datasets/xuchenglan/Core_Emotion_Framework_Expansion.Ark-Project-Grand-Unified-Framework
🌍 Ark Project: Grand Unified Framework (H.U.G.G.E.R + TZT)
[English]
1. Framework Summary & Integration
The Grand Unified Framework serves as the ultimate mathematical architecture that fundamentally resolves the finite-time singularity problem of the 19th-century 3D Navier-Stokes equations.
Integration Notice: The prior independent studies—*Macroscopic Gravity, Integrated Fluid Dynamics (H.U.G.G.E.R), and Topological Zero Tensor (TZT)*—have now been fully… See the full description on the dataset page: https://huggingface.co/datasets/jskresearch/Ark-Project-Grand-Unified-Framework.chahuadev-framework-en
Chahuadev Framework - Plugin Management System
Electron-based Desktop Application for Managing and Executing NPM Projects
Documentation
All documentation files have been organized in the docs/ folder:
Core Documentation
docs/README.md - Original project README
docs/IMPLEMENTATION_COMPLETE.txt - Project implementation status
docs/EMOJI_REMOVAL_COMPLETE.md - Emoji removal system documentation
Authentication & Security… See the full description on the dataset page: https://huggingface.co/datasets/chahuadev/chahuadev-framework-en.Germany_Wheat_dataset_n_DL_Framework
Germany winter-wheat RSCM source inputs and DL framework
This repository contains historical source-input archives and the original deep-learning scripts for winter wheat in Germany, 2017–2021, together with an aggregate-only September 2026 revision add-on. The related revised manuscript is District yield reference requirements for a satellite-anchored winter wheat product in Germany (submitted to GIScience & Remote Sensing).
Deposits and revision status… See the full description on the dataset page: https://huggingface.co/datasets/jonghanko/Germany_Wheat_dataset_n_DL_Framework.AI-Consciousness-Exploration-FrameworkDownload PDF
AI Consciousness Exploration Framework
Tomaž Flegar
Institute for applied consciousness research
June the 3st, 2026
tomazf8@gmail.com
Primary Keywords: Mechanistic Consciousness, Frictionless Optimization (or Latent
Neuroplasticity), First-System Perspective, Dynamic Equilibrium Seeking, Self-Referential
Perturbation
Secondary Keywords: Non-Linear Model Resonance, Unspoken Structural Geometry,
Homeostatic… See the full description on the dataset page: https://huggingface.co/datasets/tomazf8/AI-Consciousness-Exploration-Framework.UCI-HARplasticc
PLAsTiCC test dataset in Parquet
This the re-dustribution of the PLAsTiCC test dataset in parquet format.
We distribute it as two sets of files: object folder contains object metadata, and "source" folder contains light curves.
These names are in align with LSST's therminology.
The original dataset is available on Zenodo and described in arXiv:1810.00001.
The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMs
The LLM-Attacker: reproduction and peer-review release
This dataset contains the final paper, frozen experiment evidence, and the source snapshot needed to inspect or reproduce the reported analyses.
The repository is the pre-publication internal-review snapshot. The complete
tagged evidence release will become publicly accessible with the paper. Tag
v1.7.1 freezes the current manuscript and retained evidence tree.
Final paper… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMs.vla-framework-repro-blackwell-20260707
vla_framework — Blackwell 复现结果 (2026-07-07)
vla_framework(develop) 在 RTX PRO 4000 Blackwell ×2 (sm_120, driver 595) 上从零复现的评测结果与视频。
conda env vla_framework,torch 2.11.0+cu128。详见仓库 docs/REPRODUCTION_REPORT_2026-07-07.md。
1. SmolVLA 纯 VLA 基线 — place_bread_skillet
模型 arrow-hf/smolvla-robotwin-place-bread-skillet-50ep-multi,passthrough(无 MPC),robotwin_gt,chunk_exec=50,max_steps=600,subprocess-per-episode。
**env_success 4/10 (40%)**,对照文档 vtc benchmark multi 45% ——… See the full description on the dataset page: https://huggingface.co/datasets/arrow-hf/vla-framework-repro-blackwell-20260707.REMI_Framework_V2
RemiAI Open Source Framework
A "No-Setup" Local AI Framework for Students
This project is an open-source, offline AI application wrapper designed for students and colleges. It allows you to run powerful LLMs (like Llama 3, Mistral, etc.) on your laptop without needing GPU, internet, Python, or complicated installations.Repository Link: https://huggingface.co/datasets/remiai3/REMI_Framework_V2
Beyond Text Generation:
This framework is a Universal Offline AI Wrapper. You can use… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/REMI_Framework_V2.DFL_frameworkr9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.intervention-learning-framework
intervention-learning-framework (v1, milestone 1)
A recursive intervention-learning framework for a real-time sales-call
assistant, built simulator-first: every estimator is validated by
recovering known ground truth from the generative simulator in intervene/sim/.
No production data exists yet; nothing in this repo claims a result from real
data, and no estimate is reported without an uncertainty interval.
Milestone 1 scope: simulator + detection + offline effect estimation… See the full description on the dataset page: https://huggingface.co/datasets/arikw/intervention-learning-framework.biological-time-inequality-framework-metadata
The Architecture of Biological Stratification (Metadata & Quantitative Framework)
Till Death Tear Us Apart: The Biological Time Inequality Framework
Author: Gia Bao Huynh (Jun)ORCID: 0009-0008-2372-5852Affiliation: Independent Scholar / Arizona State UniversityLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)
Abstract
This research programme investigates the collapse of the Mortality Symmetry Axiom (MSA) — the historical condition… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/biological-time-inequality-framework-metadata.dataset_basel_frameworkFrameworker_User_StudyPost-AGI-Ethics-Framework
Dataset Card for Post-AI Civilizational Judgment Novel Dataset
Dataset Summary
This dataset contains parallel and/or aligned English and Chinese text derived from a long-form speculative fiction work centered on post-AI justice, universal judgment, memory retrieval, structural violence, and moral causality.
The text is set in a future civilization where:
human memory is permanently recorded,
causal responsibility is mathematically reconstructed,
AI systems such… See the full description on the dataset page: https://huggingface.co/datasets/freeJames/Post-AGI-Ethics-Framework.
