datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lalm-judge-validation-full-duplex
LALM Judge Validation on Full-Duplex Voice Agents
Companion dataset for the paper A Reliability Assessment of
LALM Audio Judges for Full-Duplex Voice Agents.
This repository contains the anonymised ratings, adversarial-defect
recall tables, JSON schemas, and analysis scripts used to produce
every headline number, table, and figure in that paper.
Summary
209 rated stereo sessions: 152 full-duplex agent-client
conversations across 13 accent-and-condition strata… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/lalm-judge-validation-full-duplex.synthetic-fabric-csv-validation-cases
Sewlore Synthetic Fabric CSV Validation Cases
All 72 records are invented software inputs. This is a small, deterministic
CSV-validation teaching corpus. It contains no physical fabric measurements,
garment trials, personal data, photographs or washing observations. No model was
trained or evaluated. AI assistance was used to prepare cases and documentation;
labels were captured by executing a frozen public Python package.
Use it to learn a specific parser contract, compare… See the full description on the dataset page: https://huggingface.co/datasets/sewlore/synthetic-fabric-csv-validation-cases.fisher-validation-resultseval-gliner2-ner-bc5cdr-boundary-smoothing-validationusajobs_validation
USAJOBS Dataset (Validation Sample)
Dataset Description
The USAJOBS Dataset is a comprehensive collection of federal job postings from January 2017 through March 2026. This dataset includes full-text job descriptions, and structured metadata (job title and employer).
This particular dataset presents a sample of sentence-level data from the corpus, tagged with task, skill, and AI attributes.
Dataset Structure
The dataset contains 20k sentences… See the full description on the dataset page: https://huggingface.co/datasets/loyoladatamining/usajobs_validation.Validationeval-gliner2-ner-ncbi_disease-boundary-smoothing-validationeval-gliner2-ner-mit_restaurant-affine-boundary-smoothing-validationeval-gliner2-ner-wnut2017-affine-boundary-smoothing-validationeval-gliner2-ner-ncbi_disease-affine-boundary-smoothing-validationsetfit-proj8-multilabel_2_validationeval-gliner2-ner-ontonotes5-boundary-smoothing-validationeval-gliner2-ner-conll2003-affine-validationCompanionSim-Validation
CompanionSim-Validation
70 real-world conversations annotated by two groups: 168 annotators in the US (CompanionSim-Validation-US.csv) and 998 annotators from the US, UK, India, and Nigeria (CompanionSim-Validation-Multi.csv).
Dream_NLP_Validationcross-architecture-consciousness-validation-v1eval-gliner2-ner-wnut2017-boundary-smoothing-validationeval-gliner2-ner-bionlp2004-boundary-smoothing-validationeval-gliner2-ner-ncbi_disease-affine-validationeval-gliner2-ner-fin-affine-boundary-smoothing-validationeval-gliner2-ner-bionlp2004-affine-boundary-smoothing-validationeval-gliner2-ner-bc5cdr-affine-boundary-smoothing-validationeval-gliner2-ner-mit_restaurant-boundary-smoothing-validationeval-gliner2-ner-conll2003-affine-boundary-smoothing-validationeval-gliner2-ner-wnut2017-affine-validationeval-gliner2-ner-fin-affine-validationeval-gliner2-ner-mit_restaurant-affine-validationvalidation_streameval-gliner2-ner-fin-boundary-smoothing-validationeval-gliner2-ner-ontonotes5-affine-validation
