datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
Marathi_Handwritten
Dataset Card for Marathi Handwritten OCR Dataset
Dataset Summary
The Marathi Handwritten Text Dataset is a collection of handwritten text images in Marathi (देवनागरी लिपी),
aimed at supporting the development of Optical Character Recognition (OCR) systems, handwriting analysis tools,
and language research.The dataset was curated from native Marathi speakers to ensure a variety of handwriting styles and character variations.
The dataset contains 2520 images with two… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Marathi_Handwritten.Sanskrit-OCR-Typed-Dataset
Sanskrit OCR Dataset
This dataset contains Sanskrit text images paired with their corresponding text labels, designed for OCR (Optical Character Recognition) tasks.
Dataset Structure
The dataset is split into training and validation sets:
Training set: Contains unique Sanskrit text images
Validation set: Contains separate unique Sanskrit text images
Features
image: The image containing Sanskrit text
label: The corresponding Sanskrit text label
filename:… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-OCR-Typed-Dataset.showui-web-processed
ShowUI-Web Processed
Flattened, normalized, and scenario-split version of showlab/ShowUI-web.
Each row is a single (instruction, UI element) pair with normalized bounding-box coordinates.
Schema
Column
Type
Description
sample_id
string
Unique row identifier ({row}_{element})
screenshot_id
string
Groups elements from the same screenshot
image_relpath
string
Relative path to the screenshot image
scenario
string
Website/domain inferred from the image path… See the full description on the dataset page: https://huggingface.co/datasets/e1879/showui-web-processed.panorgan-processed-pulmonology
Pan-Organ Net: Processed Pulmonology Dataset (with Train / Val / Test Splits)
Standardized, preprocessed production dataset for Pan-Organ Net multi-organ foundation model screening.
Directory Layout & Schema
splits/:
train.csv: 70% training cohort (14,815 images) with file paths and numeric class IDs
validation.csv: 15% validation cohort (3,175 images)
test.csv: 15% independent test benchmark (3,175 images)
data/:
train_images.tar.gz: Preprocessed training… See the full description on the dataset page: https://huggingface.co/datasets/theshoaibme/panorgan-processed-pulmonology.nih-processed-dataset
NIH ChestX-ray14 — Preprocessed Dataset
Dataset Description
Preprocessed version of the NIH ChestX-ray14 dataset for multi-label thoracic disease classification.
Source
Original Dataset: NIH ChestX-ray14
Institution: NIH Clinical Center
License: CC0 1.0 (Public Domain)
Citation
@inproceedings{wang2017chestx,
title={ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax… See the full description on the dataset page: https://huggingface.co/datasets/MouGam/nih-processed-dataset.
