datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-regulatory-framework
GSPC — regulatory framework facts (RegimeFacts)
In one line: For the 16 XRPL issuer accounts in the live reader: is the governing regime (NYDFS, MiCA, BACEN, Reg D and others) declared and confirmable? PASS, FAIL or UNCHECKABLE. For stablecoin and tokenised-asset analysts. It confers no regulatory status.
Use it
from datasets import load_dataset
ds = load_dataset("csoai/gspc-regulatory-framework", split="train")
print(ds[0])
Verify a signed card in your… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-regulatory-framework.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.GEO-Framework
NobleJackal GEO Framework
A practical framework for making organisations clear, verifiable and citable in AI search
GEO means Generative Engine Optimization: the work of helping generative search and answer systems find, understand and support claims about organisations, people and content. This six-language book provides a seven-layer method for auditing entity clarity, evidence quality, machine-readable structure, question coverage, multilingual parity and… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/GEO-Framework.r9-research-framework
R9 Research Framework — Qwen3.5-9B Distillation
⚠️ CRITICAL: READ FIRST — Ollama Inference Flag Required
If you serve any Qwen3.5-derived model from this lineage via Ollama,
you MUST pass "think": false in the /api/chat request body.
curl -X POST http://localhost:11434/api/chat \
-d '{"model": "qwen3.5-9b-r10:q4km", "think": false, "messages": [...], "stream": false}'
Without this flag the model will appear to "loop" and produce empty answers
on 25-46% of requests.… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r9-research-framework.repro-score-a-unified-framework-for-overshoot-refund-in-online-fdr-control-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Prettybird-Framework
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/Prettybird-Framework.redteam-framework-benchmark
ORQ Red-Teaming Framework Benchmark
Overview
This dataset contains the full results of a comparative red-teaming benchmark evaluating three
open-source red-teaming frameworks — EvaluatorQ, DeepTeam, and PromptFoo — against
three victim LLMs across three target configurations and five OWASP LLM Top 10 (2025) vulnerability
categories.
Each row is one attack attempt: the attack prompt sent to the victim model, the model's response,
and the verdict from a 3-model… See the full description on the dataset page: https://huggingface.co/datasets/orq/redteam-framework-benchmark.
