Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataforge-labs /l2-preconfirmation-reliability L2 sequencer preconfirmation observations Periodic observations of unsafe and safe heads on selected OP Stack networks, with subsequent checks of sampled unsafe block hashes against the chain reported by the endpoint. The panel records coverage and detected changes to previously observed blocks. Contents Table Record l2_preconfirmation_checks A heartbeat with head and check statistics, or a detected violation Using the data row_type… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/l2-preconfirmation-reliability.tabulartime-series-forecasting10K<n<100K0 likes2.2k downloads4h agoHugging Face02OlaOpe /scotland-bus-reliability-2026tabular1B<n<10B0 likes550 downloads8mo agoHugging Face03anon-pcqnp-ed26 /pcqnp-finite-shot-reliability-artifact PC-QNP Finite-Shot Reliability Artifact This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation. Contents code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.tabulartabular-regression1M<n<10M0 likes169 downloads5mo agoHugging Face04failproofai /fire-runtime-policy-reliability FIRE artifact This artifact accompanies FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents. It contains the frozen policies, trajectory-free attempt outcomes, analysis code, and derived tables used in the paper. It does not contain prompts, agent transcripts, tool calls, or task workspace contents. Start here Python 3.10 or newer is the only requirement for analysis. From this folder: python3 analysis/verify_results.py The command… See the full description on the dataset page: https://huggingface.co/datasets/failproofai/fire-runtime-policy-reliability.0 likes144 downloads18d agoHugging Face05DonSimpson /uk-vehicle-reliability-dataset CarHunch UK Vehicle Reliability Dataset — Volvo evaluation sample Cohort-level UK vehicle reliability statistics derived from DVSA MOT test records: first-time pass rates, defect patterns, a severity taxonomy, mileage-band behaviour and survival curves at make × model × manufacture year × fuel type. This is a free evaluation sample covering Volvo only. The full UK bundle covers every make in the MOT record. Aggregates only. No registration marks, no VINs, no keeper or owner… See the full description on the dataset page: https://huggingface.co/datasets/DonSimpson/uk-vehicle-reliability-dataset.tabular10K<n<100K1 likes123 downloads1mo agoHugging Face06Yuvrajg2107 /pranavx-gtsrb-reliability GTSRB for PranavX AI Reliability Experiments This repository repackages the German Traffic Sign Recognition Benchmark (GTSRB) for a traffic-sign classification reliability experiment. It preserves the original PPM image bytes and GTSRB class IDs in Parquet. No image is resized, cropped further, enhanced, or synthetically corrupted here. This is a project-specific derivative split, not a new official GTSRB release or an official benchmark leaderboard split. The uploader reports… See the full description on the dataset page: https://huggingface.co/datasets/Yuvrajg2107/pranavx-gtsrb-reliability.imageimage-classification10K<n<100K1 likes97 downloads16d agoHugging Face07pashas /insurance-ai-reliability-benchmark Insurance AI Agent Reliability Benchmark The first standardized benchmark for AI agents in insurance. 510 test scenarios. 10 categories. One question: Can your AI handle real insurance workflows? Why This Exists Insurance AI agents must be reliable. There is no room for error. A wrong routing decision delays a claim. A missed compliance flag triggers regulatory action. A failed escalation harms a vulnerable customer. General chatbot benchmarks do not test for this. No… See the full description on the dataset page: https://huggingface.co/datasets/pashas/insurance-ai-reliability-benchmark.text-classificationn<1K2 likes96 downloads8mo agoHugging Face08himaxym /faa-sdr-reliability-by-aircraft-family FAA Service Difficulty Reports — reliability by aircraft family, system and engine Aggregates of the FAA Service Difficulty Reporting System (SDRS). 931,502 reports for 19 airliner families and 207,475 reports for 36 engine types, broken down by ATA chapter (aircraft system), year and most-reported component. Browse it on FlightFinder: each family has a reliability page with the system breakdown. Boeing 737 (all variants) · Bombardier CRJ (all variants) · Boeing 757 · Boeing 767… See the full description on the dataset page: https://huggingface.co/datasets/himaxym/faa-sdr-reliability-by-aircraft-family.tabular10K<n<100K0 likes64 downloads12d agoHugging Face09thaki-AI /daily-paper-2026-07-13-autonomous-research-pipeline-reliability Auditing the Reliability of a Nightly Autonomous LLM Research Pipeline: Diversity, Reproducibility, and Research-Integrity Guardrails TL;DR — Eight-day quantitative audit of a production nightly autonomous paper-generation pipeline reveals a regime shift after a single guardrail fix, zero blocking integrity issues, and 0.993 keyword diversity entropy — with six minimal guardrail recommendations grounded in observed failure modes. ThakiCloud AI Research · 2026-07-13 · 📝 Tech… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-13-autonomous-research-pipeline-reliability.0 likes55 downloads3mo agoHugging Face10CodeNinjatools /plant-reliability-assessment-saudi-arabia-ontology Plant reliability assessment ontology and model register for desalination and treatment plants The object model and the model and equipment register from Reliability Atlas: A Plant Reliability Assessment Study the Operator Can Audit, an open reference architecture by CodeNinja for Saudi Arabia. Part of the Vertical-Driven Architectures series; every design in the series is also a row in the cumulative dataset… See the full description on the dataset page: https://huggingface.co/datasets/CodeNinjatools/plant-reliability-assessment-saudi-arabia-ontology.textn<1K0 likes50 downloads4d agoHugging Face11Plumloom /evaluation-reliability-benchmark Plumloom Evaluation Reliability Benchmark Public results from Plumloom’s research on reliability in single-turn AI chat evaluations. This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs. Private prompts, rubrics, model responses, and production execution details are intentionally excluded. n<1K1 likes46 downloads3mo agoHugging Face12ranausmans /reliabilityloop-v1 ReliabilityLoop v1 ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability across three production-style task types: json: schema-constrained structured extraction sql: text-to-SQL validated by SQLite execution codestub: Python function generation validated by unit tests This dataset is designed for verifier-based evaluation: outputs must work, not just look plausible. Files reliability_v1_60.jsonl Canonical split with 60 tasks: 20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.texttext-generationn<1K0 likes45 downloads8mo agoHugging Face13mirotomasik /agent-reliability-corpustabular10K<n<100K0 likes45 downloads2mo agoHugging Face14sergioburdisso /news_media_reliability Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together" Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference. Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.tabular1K<n<10K2 likes40 downloads2y agoHugging Face15TaskPuppyAI /lunamax-python311-stateful-reliability-40 LunaMax Python 3.11 Stateful Reliability 40 A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax. The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements. Dataset Size Metric Count Final records 40 Unique records 40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.texttext-generationn<1K0 likes40 downloads1mo agoHugging Face16SOTAagi2030 /Northstar-Elevator-Reliability Northstar Transit: Elevator Reliability Logs This card lists monthly log batches considered for the public reliability release. Log batches Batch ref Publication Surveyed stops QA hold Borough el.015 publish 31 clear Central Loop EL.220 publish 11 clear East Junction el.330 archive 24 clear West End EL.440 PUBLISH 006 clear Central Loop el.550 publish 19 legal South Gate EL.220 publish 15 clear North Terrace el.660 publish 31 clear West End… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Northstar-Elevator-Reliability.0 likes37 downloads1mo agoHugging Face17Disclosures-SSRC /AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research tabularother1K<n<10K0 likes36 downloads1y agoHugging Face18pravinai /agent-reliability-eval Agent Reliability Eval A short, runnable notebook on evaluating agent reliability along two axes scored separately: tool-call accuracy and hallucination rate (groundedness of the final answer against what the agent's tools actually returned). Open agent_reliability_eval.ipynb — it runs end to end with no API key and no external dependencies beyond nbformat/nbclient if you want to re-execute it; the shipped copy already has outputs baked in. Why two metrics instead of… See the full description on the dataset page: https://huggingface.co/datasets/pravinai/agent-reliability-eval.0 likes32 downloads28d agoHugging Face19hoololi /llm-agent-harness-reliability-next-prime LLM Next Prime Harness Dataset This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool. The task is deliberately simple and objectively verifiable: What is the smallest prime number that is strictly greater than n? The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.tabulartext-generation10K<n<100K0 likes31 downloads1mo agoHugging Face20aadarshram /eval_reliability_act_pick_place_tapeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_follower", "total_episodes": 20, "total_frames": 4998, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/eval_reliability_act_pick_place_tape.tabularrobotics1K<n<10K0 likes29 downloads1y agoHugging Face21MCP-1st-Birthday /smoltrace-site-reliability-engineering-tasks SMOLTRACE Synthetic Dataset This dataset was generated using the TraceMind MCP Server's synthetic data generation tools. Dataset Info Tasks: 80 Format: SMOLTRACE evaluation format Generated: AI-powered synthetic task generation Usage with SMOLTRACE from datasets import load_dataset # Load dataset dataset = load_dataset("MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks") # Use with SMOLTRACE # smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks.textn<1K0 likes28 downloads10mo agoHugging Face22ProblemsByVin /vehicle-reliability-scorecard Vehicle Reliability Scorecard One row per vehicle (year + make + model) with its ProblemsByVin reliability score, total NHTSA complaints, recalls, and defect investigations, plus the single component owners complain about most. The master index across the whole tracked fleet — the flat table to join every other dataset to. Columns column meaning year Model year make Manufacturer model Model reliability_score 1.0 (worst) – 5.0 (best); shown on site… See the full description on the dataset page: https://huggingface.co/datasets/ProblemsByVin/vehicle-reliability-scorecard.tabular1K<n<10K0 likes27 downloads3mo agoHugging Face23ClarusC64 /ffr-coherence-drift-reliability-collapse-detection-v0.1Dataset goal Detect when the correlation between CT image qualityand AI-derived FFR stability starts to break. The output can look plausiblewhile reliability collapses. Inputs snr motion_score segmentation_confidence artifact_score ffr_run_variance model_disagreement calibration_residual Required outputs reliability_collapse_flag instability_onset_min_ahead drift_pattern_label coherence_decay_score collapse_risk_score Labels reliability_collapse_flag 0 = coherent behavior 1 = drift… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ffr-coherence-drift-reliability-collapse-detection-v0.1.tabular-classificationn<1K0 likes25 downloads8mo agoHugging Face24ClarusC64 /legal-evidence-reliability-probativeness-coherence-v0.1What this dataset is You receive evidence type reliability basis probative claim prejudice or confusion risk gatekeeping signals appellate posture You decide Does probative claim match reliability Answer coherent or incoherent Why this matters When coherence fails evidence gets excluded new trial risk rises verdict stability collapses tabulartext-classificationn<1K0 likes24 downloads8mo agoHugging Face25MonikaDvorackova /agent-reliability-traces Agent Reliability Traces A small synthetic dataset of observable AI agent execution traces annotated with reliability and failure-mode signals. The dataset accompanies the Agent Reliability Lab Hugging Face Space. Dataset purpose The dataset is designed for: prototyping agent-trace evaluation testing deterministic reliability heuristics experimenting with failure-mode classification evaluating tool-use trajectories educational and portfolio use It is not… See the full description on the dataset page: https://huggingface.co/datasets/MonikaDvorackova/agent-reliability-traces.texttext-classificationn<1K0 likes24 downloads2mo agoHugging Face26nn-stability-research /model-reliability-benchmark Model Reliability Benchmark Neural network benchmark data for ML research. Usage from datasets import load_dataset dataset = load_dataset("nn-stability-research/model-reliability-benchmark") df = dataset["train"].to_pandas() Or use the provided loader: from loader import load_data df = load_data() Schema Metrics Column Type Description activation_diversity float Normalized metric gradient_consistency float Normalized metric… See the full description on the dataset page: https://huggingface.co/datasets/nn-stability-research/model-reliability-benchmark.tabulartabular-classification1K<n<10K0 likes23 downloads8mo agoHugging Face27devinsam /scotland-bus-reliability-2026text100M<n<1B0 likes22 downloads8mo agoHugging Face28ClarusC64 /cascade-f1-powerunit-cooling-ambient-reliability-v0.1 What this repo does This repo models a quad coupling pattern linked to thermal reliability collapse. It supports: • scoring race states for DNF risk region entry• identifying which variables drive thermal margin loss• testing cooling and load redesign moves The sample is synthetic.It shows the geometry. Core quad • engine_load• cooling_capacity• ambient_temp• component_degradation_rate Prediction target label_cascade • 0 means stable thermal operating… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-f1-powerunit-cooling-ambient-reliability-v0.1.tabulartext-classificationn<1K0 likes22 downloads7mo agoHugging Face29achiepatricia /han-distributed-network-latency-reliability-dataset-v1 Humanoid Distributed Network Latency & Reliability Dataset This dataset models real-time network performance between distributed humanoid agents operating inside a decentralized cognitive mesh. It captures latency variance, packet loss patterns, synchronization delays, and task completion reliability metrics. Objective To enable performance-aware humanoid coordination under varying network conditions. Why This Is Critical Decentralized humanoid systems rely… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-network-latency-reliability-dataset-v1.textn<1K0 likes21 downloads8mo agoHugging Face30TaskPuppyAI /qwen3.8-reliability-40 Qwen3.8 Max Stateful Reliability 40 A 40-record synthetic programming dataset generated with Qwen3.8 Max and reviewed/cleaned with ChatGPT 5.6 Sol High. The dataset focuses on compact stateful implementations and repair tasks where correctness depends on preserving behavioral invariants across operations. Dataset Summary The publication artifact contains 40 unique records using the schema: { "user": "...", "assistant": "..." } Recovered final-artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-reliability-40.textn<1K0 likes21 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.