Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ranausmans /reliabilityloop-v1 ReliabilityLoop v1 ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability across three production-style task types: json: schema-constrained structured extraction sql: text-to-SQL validated by SQLite execution codestub: Python function generation validated by unit tests This dataset is designed for verifier-based evaluation: outputs must work, not just look plausible. Files reliability_v1_60.jsonl Canonical split with 60 tasks: 20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.texttext-generationn<1K0 likes45 downloads8mo agoHugging Face02TaskPuppyAI /lunamax-python311-stateful-reliability-40 LunaMax Python 3.11 Stateful Reliability 40 A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax. The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements. Dataset Size Metric Count Final records 40 Unique records 40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.texttext-generationn<1K0 likes40 downloads1mo agoHugging Face03hoololi /llm-agent-harness-reliability-next-prime LLM Next Prime Harness Dataset This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool. The task is deliberately simple and objectively verifiable: What is the smallest prime number that is strictly greater than n? The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.tabulartext-generation10K<n<100K0 likes31 downloads1mo agoHugging Face04paiteq-ai /agent-reliability-2026q3 Agent Reliability Benchmark, 2026-Q3 Status: planned. Results target 2026-09. Depends on paiteq/ai-eval-harness v0.2 (agent rubric layer). A dated, reproducible benchmark on agent reliability. 100 tasks covering tool-calling, multi-step execution, and error recovery. Pass@1, pass@5, mean steps, mean cost per task, recovery rate, and latency p95 across Claude, GPT, Gemini, and an open-source baseline. This dataset card is the canonical landing for the task set. Full methodology… See the full description on the dataset page: https://huggingface.co/datasets/paiteq-ai/agent-reliability-2026q3.question-answeringn<1K0 likes20 downloads5mo agoHugging Face05brikdavies /msm-mixed-llama-reliability-claude-risk MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.texttext-generation1K<n<10K0 likes14 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.