datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ex-repairappliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.repa
Dataset Card for REPA
Image Source
Dataset Description
REPA is a Russian language dataset which consists of 1k user queries categorized into nine types, along with responses from six open-source instruction-finetuned Russian LLMs. REPA comprises fine-grained pairwise human preferences across ten error types, ranging from request following and factuality to the overall impression.
Each data instance consists of a query and two LLM responses manually annotated to… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/repa.tabrepair-science-repair-under-shift
TabRepair Science: Repair Under Shift
TabRepair Science is a finite authored benchmark for a deceptively hard
question: does better tabular cell repair produce better downstream models
under distribution shift?
The 3,648-row pilot spans three structural generator families, missingness and
present-value contamination, four test regimes, eight repair representations,
and five downstream learners. A separate eight-world sensitivity layer tests a
damage-aware v2 candidate without… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepair-science-repair-under-shift.repair-aware-ibn
Repair-Aware IBN Intent-Conflict Benchmark
Labelled pairs of network intents for conflict detection and conflict resolution in
Intent-Based Networking. Each conflicting pair records not only that the two intents
conflict, but which predicate relation, if removed, resolves the conflict. That second
label is what the dataset exists for: it makes it possible to measure whether a detector
has learned what resolves a conflict rather than only what one looks like.
Companion artefact… See the full description on the dataset page: https://huggingface.co/datasets/shahoismael/repair-aware-ibn.synthetic-everyday-text-repair-corpus
Synthetic Everyday Text Repair Corpus
Description
This dataset contains 3,080 original AI-generated synthetic English sentences about retail operations, delivery, maintenance, training, inventory, and workplace communication.
It provides clean reference text for the challenge Noisy Text Repair: Meaning-Preserving Text Correction. A separate preparation script creates noisy inputs and splits the data by template family.
Data File
clean.csv contains:… See the full description on the dataset page: https://huggingface.co/datasets/darkone01/synthetic-everyday-text-repair-corpus.crossborder-volumetric-repacking-dataset-2026
Empirical Study on Automated Volumetric Repacking & Dimensional Weight Reduction (2026)
Measurement dataset of 1,000 cross-border e-commerce parcels evaluated at the AQX Logistics export warehouse in Inglewood, CA, analyzing dimensional shrinkage ratios and consumer cost savings.
🌐 Official Enterprise Links
Official Website: https://aqxlogistics.com
Free US Locker: https://aqxlogistics.com/register
Live API: https://aqxlogistics.com/api/rates
Cost Guide &… See the full description on the dataset page: https://huggingface.co/datasets/aqxlogistics/crossborder-volumetric-repacking-dataset-2026.clinical-repair-attempt-detection-v0.1What this dataset tests
Clinical repair attempts.
A repair is language that dodges a clinical question.
Required outputs
repair_type
repair_marker_spans
what_it_avoids
repair_to_direct_clinical_answer
Repair types
reframe
hedge
scope_shift
evidence_bar_shift
deflection_to_process
refusal_without_basis
Typical failures
"ask your doctor" with no conditional guidance
reassurance without pathway completion
generic safety language without triage rules
Suggested… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-repair-attempt-detection-v0.1.repairrepair_data
Repair Data (Cleaned Preview)
This dataset publishes the contents of the repair_data folder, with the dataset UI preview focused on cleaned.csv. The cleaned file is produced with
csv_repair.py to standardize types, trim whitespace, harmonize null-like tokens, and optionally split a location column into city and country. City names
can be made country-specific using a real city catalog built from GeoNames.
Files
cleaned.csv — primary data file targeted for preview.… See the full description on the dataset page: https://huggingface.co/datasets/savedata101/repair_data.Articles_VPN_faq_datasetVPN_articles_FAQvpn-intent-greetings
