datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
getting-started-labeled-validation
Dataset Card for validation_photos
This is a FiftyOne dataset with 143 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("TheSteve0/getting-started-labeled-validation")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-labeled-validation.getting-started-validation-clip-pred
Dataset Card for labeled_validation_predicted_clip
This is a FiftyOne dataset with 143 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("TheSteve0/getting-started-validation-clip-pred")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-validation-clip-pred.HUD-UI-Validation-Pilot
HUD/UI annotation validation pilot
This public artifact compares 10 gameplay screenshots across four columns:
raw frame;
GPT-5.6 Sol X-High final reference;
GPT-5.6 Terra High refined annotation;
confidence-routed final annotation.
Files:
analysis.md: aggregate metrics and per-game error table;
per_sample_metrics.csv: machine-readable sample metrics;
overlay_comparison_contact_sheet.jpg: full comparison sheet.
The 0--100 confidence value is a conservative pipeline routing… See the full description on the dataset page: https://huggingface.co/datasets/GameWorldData/HUD-UI-Validation-Pilot.medical_records_parsing_validation_set
Medical Records Parsing Validation Set
Dataset Composition and Clinical Relevance
The Eka Medical Records Parsing Dataset empowers evaluation of AI systems designed to extract structured information from unstructured medical documents, enabling true digitisation of healthcare data while maintaining clinical accuracy.
The dataset comprise 288 carefully selected images of laboratory reports and prescriptions representing diverse formats and templates encountered in Indian… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/medical_records_parsing_validation_set.
