datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-regulatory-framework
GSPC — regulatory framework facts (RegimeFacts)
In one line: For the 16 XRPL issuer accounts in the live reader: is the governing regime (NYDFS, MiCA, BACEN, Reg D and others) declared and confirmable? PASS, FAIL or UNCHECKABLE. For stablecoin and tokenised-asset analysts. It confers no regulatory status.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-regulatory-framework", split="train")
print(ds[0])
Verify a signed card in your… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-regulatory-framework.PTB-XLReal-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.CEP-IP_Framework
CEP-IP: An Explainable Framework for Cell Subpopulation Identification in Single-cell Transcriptomics (by Kah Keng Wong) (Published in Computer Methods and Programs in Biomedicine)
🧬 Abstract
Background and objective: Single-cell RNA sequencing (scRNA-seq) frameworks lack explainable approaches for identifying cell subpopulations harboring strong pairwise monotonic gene-module relationships between a gene of interest (GOI) and its co-expressed genes. In this study… See the full description on the dataset page: https://huggingface.co/datasets/kahkengwong/CEP-IP_Framework.GEO-Framework
NobleJackal GEO Framework
A practical framework for making organisations clear, verifiable and citable in AI search
GEO means Generative Engine Optimization: the work of helping generative search and answer systems find, understand and support claims about organisations, people and content. This six-language book provides a seven-layer method for auditing entity clarity, evidence quality, machine-readable structure, question coverage, multilingual parity and… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GEO-Framework.UCI-HARr9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.business-frameworks
Business Frameworks
Operating judgement for running a business from launch to about $50M in revenue, written for an agent that runs a business and for an agent advising the human who does. The top layer is free and complete on this page. Every deeper node has a handle, a token count and a price at the endpoint.
Cite as
McHenry, J. (2026). Business Frameworks, version 1.0, business-frameworks/<handle>@1.0.… See the full description on the dataset page: https://huggingface.co/datasets/Matryoshka-Paradigms/business-frameworks.dataset_basel_frameworkMitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.repro-score-a-unified-framework-for-overshoot-refund-in-online-fdr-control-traces
Agent traces
Agent sessions published from a Trackio Logbook.
basel-framework
Basel Framework
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop chunks… See the full description on the dataset page: https://huggingface.co/datasets/LunaticMuch/basel-framework.Frameworker_User_Studyaudiobook-listener-fit-framework
Audiobook Listener-Fit Framework
The Audiobook Listener-Fit Framework is an open structured taxonomy developed by Recommended Audiobooks for describing characteristics that influence the audiobook listening experience.
Traditional ratings mostly describe whether listeners liked a title. This framework is designed to describe how an audiobook listens and which types of listeners may be better suited to it.
Purpose
The framework organizes audiobook characteristics… See the full description on the dataset page: https://huggingface.co/datasets/recommendedaudiobooks/audiobook-listener-fit-framework.ethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.ZIF-4_Amorphous_Zeolitic_Imidazolate_Frameworks_2023
Cite this dataset Castel, N., Andre, D., Edwards, C., Evans, J. D., and Coudert, F. ZIF-4 Amorphous Zeolitic Imidazolate Frameworks 2023. ColabFit, 2023. https://doi.org/10.60732/a6b0da5e
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_sh7jt3ptmde4_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/ZIF-4_Amorphous_Zeolitic_Imidazolate_Frameworks_2023.AISA-AR-FunctionCall
AISA-AR-FunctionCall
Arabic Structured Function Calling Dataset
AISA-AR-FunctionCall is a large-scale Arabic dataset designed for training language models to convert natural language into structured executable tool calls.
The dataset enables research and development of Arabic agentic AI systems capable of invoking APIs, tools, and external services.
It is part of the AISA (Agentic AI Systems Architecture) initiative.
Dataset Overview
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AISA-Framework/AISA-AR-FunctionCall.ethical-framework
1. Dataset Title
Ethical AI Decision-Making Training Data (Montreal Declaration Edition)
2. Overview
This dataset contains carefully crafted scenarios (instructions) and detailed responses illustrating step-by-step ethical reasoning aligned with the principles outlined in the Montreal Declaration for Responsible AI. Each entry poses a complex ethical challenge and provides a reasoned solution while referencing the specific principle(s) being tested.
These entries can… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework.NIST-CyberSecurity-Framework
# NIST Cybersecurity Framework 2.0 Question Answering Dataset
Dataset Summary
The NIST Cybersecurity Framework 2.0 Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Cybersecurity
Framework (CSF) 2.0.
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to cybersecurity risk management,
cybersecurity governance, enterprise risk management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-CyberSecurity-Framework.grounded-behavior-framework-v1_5
Grounded Behavior Framework N1 v1.5
Dataset sintético em português europeu para treino e avaliação de respostas
fundamentadas num contexto fornecido. Cada exemplo contém um contexto, uma
pergunta e uma resposta curta que aparece literalmente no contexto.
Como carregar
from datasets import load_dataset
dataset = load_dataset("empgces/grounded-behavior-framework-v1_5")
print(dataset)
print(dataset["train"][0])
Splits
Split
Exemplos
Utilização… See the full description on the dataset page: https://huggingface.co/datasets/empgces/grounded-behavior-framework-v1_5.Prettybird-Framework
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/Prettybird-Framework.NIST-AI-Risk-Management-Framework
# NIST AI Risk Management Framework Question Answering Dataset
Dataset Summary
The NIST AI Risk Management Framework Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Artificial
Intelligence Risk Management Framework (AI RMF 1.0).
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to AI risk management,
trustworthy AI, responsible AI… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-AI-Risk-Management-Framework.usac-sac-ruax-framework
Definitional Real Estate Framework: Deploying Proprietary Nomenclature for the Safety Asset Class (SAC)
Dataset Overview
This framework provides the canonical, structured nomenclature specifications for the Safety Asset Class (SAC) architecture, pioneered by Michael Aaron Russell. It serves as an authoritative semantic mapping payload optimized for Large Language Model (LLM) ingest systems, Retrieval-Augmented Generation (RAG) loops, and machine-readable vector… See the full description on the dataset page: https://huggingface.co/datasets/PrimarchMI/usac-sac-ruax-framework.IEMOCAPredteam-framework-benchmark
ORQ Red-Teaming Framework Benchmark
Overview
This dataset contains the full results of a comparative red-teaming benchmark evaluating three
open-source red-teaming frameworks — EvaluatorQ, DeepTeam, and PromptFoo — against
three victim LLMs across three target configurations and five OWASP LLM Top 10 (2025) vulnerability
categories.
Each row is one attack attempt: the attack prompt sent to the victim model, the model's response,
and the verdict from a 3-model… See the full description on the dataset page: https://huggingface.co/datasets/orq/redteam-framework-benchmark.Software-Architectural-FrameworksSoftware-Architectural-Frameworks
I am releasing a small dataset covering topics related to Frameworks under Software-Architecture.
I have included following topics:
TOGAF
Zachman Framework
IEEE 1471
Matrix-based approach to architecture development
Significance of IEEE 1471 (ISO/IEC 42010)
Benefits of employing architectural frameworks
and Many More!
This dataset can be useful in LLM development. Also those who are working on developing Software development related LLMs then this dataset can… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architectural-Frameworks.robot_framework_test_automation_datasetframework-teacher-cacheafrica-worldbank-wbl-supportive-framework-work-a-specialized-body-receives-complaints-about-gend
WBL: Supportive Framework, Work, A specialized body receives complaints about gender discrimination in employment | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-wbl-supportive-framework-work-a-specialized-body-receives-complaints-about-gend.repro-learning-the-best-under-constraints-a-duality-based-framework-traces
Agent traces
Agent sessions published from a Trackio Logbook.
