datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
construction-multimodel-dataset
Construction Traversability Dataset
A construction-site RGB semantic segmentation dataset developed for research on
terrain understanding, traversability estimation, and multimodal RGB–LiDAR
perception for mobile robots.
Overview
This dataset contains RGB images and pixel-wise semantic segmentation masks
collected in construction-site environments. The dataset is intended to support
research on construction-site scene understanding and traversability-aware
robot… See the full description on the dataset page: https://huggingface.co/datasets/vla-model/construction-multimodel-dataset.next-jev-stage1-multi-model-cot
Next JEV Stage 1 multi-model CoT
One row is one prompt with its paired model responses. images is a list of
embedded Hugging Face Image features, ordered with image_sha256s. record_json
preserves the exact training row for reproducible materialization. This is a
CoT-only Stage 1 export, with no answer or classification target required.
The original source, split counts and JSONL SHA-256 values are in
reports/manifest.json; Parquet SHA-256 values are in reports/native.json.
The… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-stage1-multi-model-cot.modelnet40_multi_viewmultimodel_llava_med_zh_instruct_60kBorrowed from https://huggingface.co/datasets/BUAADreamer/llava-med-zh-instruct-60k
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
multi-pixmo-ask-model-anything
Multi-PixMo-AskModelAnything
Overview
Multi-PixMo-AskModelAnything is a multilingual extension of the original PixMo-AskModelAnything dataset from AllenAI, part of the PixMo series of multimodal resources.
The original PixMo-AskModelAnything dataset consists of image-based question–answer pairs, where annotators authored freeform questions about an image, and answers were generated through a pipeline combining OCR output, dense captions, and a language-only LLM.… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/multi-pixmo-ask-model-anything.stable-bias_grounding-images_multimodel_3_12_22
Dataset Card for "stable-bias_grounding-images_multimodel_3_12_22"
More Information needed
stable-bias_grounding-images_multimodel_3_12_22
Dataset Card for "stable-bias_grounding-images_multimodel_3_12_22"
More Information needed
stable-bias_grounding-images_multimodel_3_12_22_clusters
Dataset Card for "stable-bias_grounding-images_multimodel_3_12_22_clusters"
More Information needed
stable-bias_grounding-images_multimodel_3_12_22_clusters
Dataset Card for "stable-bias_grounding-images_multimodel_3_12_22_clusters"
More Information needed
multimodel_evalMultimodel_Parking_Datasetmultimodel-datasetDeepfashion_MultiModel
