datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reliabilityloop-v1
ReliabilityLoop v1
ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability
across three production-style task types:
json: schema-constrained structured extraction
sql: text-to-SQL validated by SQLite execution
codestub: Python function generation validated by unit tests
This dataset is designed for verifier-based evaluation: outputs must
work, not just look plausible.
Files
reliability_v1_60.jsonl
Canonical split with 60 tasks:
20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.lunamax-python311-stateful-reliability-40
LunaMax Python 3.11 Stateful Reliability 40
A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax.
The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements.
Dataset Size
Metric
Count
Final records
40
Unique records
40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.llm-agent-harness-reliability-next-prime
LLM Next Prime Harness Dataset
This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool.
The task is deliberately simple and objectively verifiable:
What is the smallest prime number that is strictly greater than n?
The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.agent-reliability-2026q3
Agent Reliability Benchmark, 2026-Q3
Status: planned. Results target 2026-09. Depends on paiteq/ai-eval-harness v0.2 (agent rubric layer).
A dated, reproducible benchmark on agent reliability. 100 tasks covering tool-calling, multi-step execution, and error recovery. Pass@1, pass@5, mean steps, mean cost per task, recovery rate, and latency p95 across Claude, GPT, Gemini, and an open-source baseline.
This dataset card is the canonical landing for the task set. Full methodology… See the full description on the dataset page: https://huggingface.co/datasets/paiteq-ai/agent-reliability-2026q3.msm-mixed-llama-reliability-claude-risk
MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.
