datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.pcl-viewer-kitti-movie
pcl-viewer KITTI movies
Draco-compressed LiDAR frames for the
pcl-viewer demo, in two folders:
geometry/ — sweeps from KITTI raw drive 2011_09_26_drive_0005, positions
plus per-point intensity (Draco color green channel).
seg/ — SemanticKITTI sequence slice with a per-point class id (Draco
color red channel) and intensity (green channel), plus boxes.json (one
axis-aligned 3D box per thing instance per frame).
Attribution & license
Source: KITTI / SemanticKITTI… See the full description on the dataset page: https://huggingface.co/datasets/kolodkin/pcl-viewer-kitti-movie.view2space-v1
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
Related VIEW2SPACE Releases
Training release: Pokerme/view2space-train
4B model checkpoint: Pokerme/view2space_4b
Collection: Pokerme/view2space
The public release is organized into three subsets:
count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.HEC3R-ckpt-fixed_view_axis_freeze_decCS50-rawlongbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.pi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.traceweave-viewer-test
Agent Traces
Coding-agent sessions collected with TraceWeave,
rehydrated into the Claude Code JSONL schema
consumed by the Hugging Face Agent Trace Viewer.
Format
Each *.jsonl file at the dataset root is one session. Events use:
{"type":"user","message":{"role":"user","content":"..."},"uuid":"...","parentUuid":null,"sessionId":"...","timestamp":"..."}
{"type":"assistant","message":{"role":"assistant","content":[{"type":"text","text":"..."}]},"uuid":"..."… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/traceweave-viewer-test.ovos-wake-word-bench-picovoice-view-glass
OVOS wake_word bench — picovoice-view-glass
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.test_auto_refresh_views_202610070708129053LGUplus_viewing_history
LGUplus Persona TV Viewing History
한국인 가상 페르소나 1만 명의 일주일치 TV 시청이력 합성 데이터셋입니다.
nvidia/Nemotron-Personas-Korea 페르소나와
실제 U+tv EPG(258채널 × 7일, 2026-06-03~09) 편성표를 기반으로, 교사 LLM(Qwen3-235B-A22B-Instruct-2507-FP8)이
2단계로 생성했습니다.
생성 방법
시청 스케줄 생성: 페르소나(직업·나이·가족·취미)를 보고 요일별 시청 시간대를 추정
(근거 문장을 schedule_reasoning으로 함께 생성 — 예: 주부/은퇴자는 평일 낮, 직장인은 저녁)
프로그램 선택: 각 (요일, 시간창)마다 해당 시간에 방영 중인 실제 편성표를 제시하고
페르소나가 시청할 프로그램을 순서대로 선택 (채널 현실성 지시: 주류 채널 위주,
전문채널은 취미·직업 일치 시에만 / 이미 본 (제목,회차)… See the full description on the dataset page: https://huggingface.co/datasets/ENERZAiKR/LGUplus_viewing_history.parallelvla-putcab-serial-parallel-random-delay-viewer-seed427-20260811test_view_delNSFW_RP_Format_DPO_ViewerThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action*
This dataset contains NSFW content.
view_testtmp-viewer3test-viewer
Test Viewer Dataset
Small test dataset for viewer.
short_video_ocr_viewer_test
Minimal Dataset Viewer test
This repository intentionally contains one tiny JSON data file. The YAML
configuration explicitly tells Hugging Face Dataset Viewer which file and split
to index.
