datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset
NEST3D: A High-Resolution Multimodal Dataset of Sociable Weaver Tree Nests
Dataset Description
NEST3D is a multimodal dataset of 104 sociable weaver nests, combining drone-based RGB and multispectral imagery with a semantically annotated 3D RGB point cloud. It captures trees hosting these nests through drone-based remote sensing, providing rich spatial and spectral information to benchmark and advance scene-level semantic segmentation methods for computer vision… See the full description on the dataset page: https://huggingface.co/datasets/NEST3D/dataset.Nectar
Dataset Card for Nectar
Developed by: Banghua Zhu * , Evan Frick * , Tianhao Wu * , Hanlin Zhu and Jiantao Jiao.
License: Apache-2.0 license under the condition that the dataset is not used to compete with OpenAI
Nectar is the first high-quality 7-wise comparison dataset, generated through GPT-4-based ranking. Nectar contains diverse chat prompts, high-quality and diverse responses, and accurate ranking labels. Nectar's prompts are an amalgamation of diverse sources, including… See the full description on the dataset page: https://huggingface.co/datasets/berkeley-nest/Nectar.nestful
NESTFUL: Nested Function-Calling Dataset
NESTFUL is a benchmark to evaluate LLMs on nested sequences of API calls, i.e., sequences where the output of one API call is passed as input to
a subsequent call.
The NESTFUL dataset includes over 1800 nested sequences from two main areas: mathematical reasoning and coding tools. The mathematical reasoning portion is generated from
the MathQA dataset, while the coding portion is generated from the
StarCoder2-Instruct dataset.
All… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/nestful.nestle1904-quotation-refs
NuBerea Nestle 1904 Quotation References
Quotation reference annotations mapping Old Testament quotations cited in the New Testament to their original source locations, extracted from the Nestle 1904 Greek New Testament critical apparatus. Each entry records where a New Testament passage cites the Old Testament, together with the apparatus's canonical source-location reference — supporting work in intertextuality, reception history, and the New Testament's use of the Old… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/nestle1904-quotation-refs.Physical_AI_SO101_Cup_Nesting_Task
DecisionFacts Physical AI Dataset — SO-101 Robotic Arm Teleoperation
Data Summary
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded… See the full description on the dataset page: https://huggingface.co/datasets/DecisionFacts/Physical_AI_SO101_Cup_Nesting_Task.robotwin-cups-nesting
RoboTwin Cup-Nesting (Bimanual, Dual-Aloha)
500 successful expert demonstrations of bimanual cup nesting — a dual-arm
robot nests three cups into one another — collected in the RoboTwin
2.0 simulator with strong domain
randomization. Format: LeRobot v2.1.
What's in it
Episodes
500 (all success-filtered)
Frames
287,571 @ 30 fps
Embodiment
dual-arm Aloha (aloha-agilex), 14-DoF
Action / state
14-dim absolute joint positions
Cameras
head +… See the full description on the dataset page: https://huggingface.co/datasets/buzinguyen/robotwin-cups-nesting.vesuvius-ink-training
Vesuvius ink — training-ready
Inputs and splits for a general ink-detection model that works on any scroll: a 3-D encoder
pretrained on the scrolls' own CT, fine-tuned on the official labels. Every input here was made by
the same code a new scroll goes through — its CT sampled along its traced surface at native
resolution, straight from the official open-data volumes — and checked against the official render
of the same surface (per-tile correlation in each stack.json).
Labelled… See the full description on the dataset page: https://huggingface.co/datasets/nestorvfx/vesuvius-ink-training.DatasetScrollsur3-remove-cup-from-nested-cupsvesuvius-ink-corpus
Vesuvius ink-label corpus — every official ink label, one download
Lossless repack (tar + zstd) of all ink-label-relevant data released by the
Vesuvius Challenge / EduceLab, assembled 2026-07-28 so a training box can bootstrap
with one fast download instead of thousands of small requests.
Shard
Contents
hf_ink.tar.zst
All 33 labelled segments (8 objects) from hf://buckets/scrollprize/datasets/ink/ + labelled unused/ entries: *_inklabels, *_supervision_mask… See the full description on the dataset page: https://huggingface.co/datasets/nestorvfx/vesuvius-ink-corpus.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.nested-tkf-repro-workspace
Code for "Nested birth-death processes are competitive with neural networks as protein evolution models", ICML 2026.
This repo contains code to preprocess Pfam v36.0 into pairwise alignments, train and evaluate pairHMM and neural models of protein sequence evolution, and reproduce all reported log-likelihoods.
We also include details about the three train/dev/test partitions, as well as log-likelihood metrics for all models trained on the three training replicates.
The bioRxiv… See the full description on the dataset page: https://huggingface.co/datasets/latticetower/nested-tkf-repro-workspace.stocks-NESTLEIND-1D-candlesaloha_real_agilex_nesting_dollomx_cup_nestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/LinhanWang/omx_cup_nest.so100-cup-one-nesting-v1_20260728_190534This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/yiju2/so100-cup-one-nesting-v1_20260728_190534.PrimeVulnested-sycophancy
Nested Geometry of Sycophancy
Dataset and activation features for the paper Nested Sycophancy: Probes Find Asymmetric Mechanisms Across Three LLM Architectures, But Static Steering Fails.
Overview
We study the internal geometry of sycophancy in reasoning LLMs, stratified by epistemic uncertainty. The dataset contains:
Sycophancy evaluation data from Anthropic model-written evaluations, with model-generated Chain-of-Thought responses and sycophancy labels
Activation… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-author-1/nested-sycophancy.aloha_real_agilex_nesting_doll-siglip2-droid-ft
aloha_real_agilex_nesting_doll with SigLIP2-DROID targets
This LeRobot v2.1 dataset is a copy of SakikoTogawa/aloha_real_agilex_nesting_doll with observation.concept added.
For every timestep, the three camera images are processed independently by the
fine-tuned SakikoTogawa/siglip2-base-droid vision encoder. The vision_model.pooler_output vectors
(768 dimensions each) are concatenated in the order left wrist, right wrist,
and high camera, producing a 2304-dimensional float32… See the full description on the dataset page: https://huggingface.co/datasets/JackieMM/aloha_real_agilex_nesting_doll-siglip2-droid-ft.vesuvius-ink256-pretrain-sheets
Flattened scroll sheets for self-supervised pretraining
Small flattened papyrus sheets rendered from the Vesuvius Challenge open-data CT scans
(s3://vesuvius-challenge-open-data, CC BY-NC 4.0, https://scrollprize.org/data), made to pretrain a 9 um ink-detection
network (official vesuvius_unet_3d_stem_2d) by masked reconstruction. No labels.
Scans: every volume with an official surface prediction (m7): 27 survey scans at level 0 (8.64 / 9.36 um) and
14 scans of 2.2-2.4 um at… See the full description on the dataset page: https://huggingface.co/datasets/nestorvfx/vesuvius-ink256-pretrain-sheets.berkeley-nest-Nectar-DPOSource berkeley-nest/Nectar
Physical_AI_SO101_Cup_Nesting_Task
DecisionFacts Physical AI Dataset — SO-101 Robotic Arm Teleoperation
Data Summary
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded… See the full description on the dataset page: https://huggingface.co/datasets/salehin21/Physical_AI_SO101_Cup_Nesting_Task.NEST
NEST: NEw Sparse maTrix dataset
NEST is a new sparse matrix dataset. Its purpose is to define a modern set of sparse matrices arising in relevant and actual scientific application in order to improve further sparse numerical methods. Nest can be seen as a continuity of the Sparse Matrix Market datasets and contain some curated sparse matrices from it as legacy references.
The matrices are stored as COO sparse matrices in scipy.sparse.npz archive format. Conversion utils to/from the… See the full description on the dataset page: https://huggingface.co/datasets/vincent-maillou/NEST.nestedclinbr
NestedClinBr Corpus
NestedClinBr is a new corpus containing nested and discontinuous entities in Brazilian Portuguese clinical narratives.
The main goal of NestedClinBr is to provide a human-annotated corpus that can be used for learning and evaluating different machine learning models to extract valuable medical information in the Portuguese language, in special nested and discontinuous entities, an important but less explored task.
In the context of clinical NLP, the recognition… See the full description on the dataset page: https://huggingface.co/datasets/pucpr-br/nestedclinbr.Pregenerated_Synth_Dataset
Pregenerated Synthetic & Real Crop Dataset for Scroll FFN
Training and evaluation corpus for the Vesuvius Challenge 3D Flood-Fill Network (FFN).
See the GitHub repository: https://github.com/nestorvfx/FFN-Scroll-Tracing
bimanual-cup-nestingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 50,
"total_frames": 122650,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cortx-labs/bimanual-cup-nesting.so100-cup-one-nesting-varied-v1_20260731_183655This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/yiju2/so100-cup-one-nesting-varied-v1_20260731_183655.beir-minus-nanobeir-queries-random-nested-subsetscapdyn_match_rzero_global_nested5
Global CapDyn-Match(R-Zero Iter1–Iter3,严格 nested 5-fold)
本目录是唯一允许写入的实验工作区。三个 iteration 的 evaluation JSON 与 checkpoint 只读。
方法
从已保存的 Iter1/Iter2/Iter3 回答中监督选择,不重新生成、不重新 self-evolution。
QA 表示:冻结 Iter2 encoder,对 (question, a_k) 取 layer 23、最后一个非 padding token
f_QA:StandardScaler → TruncatedSVD(≤128) → balanced LogisticRegression
f_cap:Question TF-IDF/SVD ⊕ 精确 7 维 agreement
融合:行内 z-score,搜索 alpha ∈ {0.25,0.50,0.75}、margin ∈ {0,0.05,0.10}
切分单位:split_group_id = benchmark… See the full description on the dataset page: https://huggingface.co/datasets/hpachdpii/capdyn_match_rzero_global_nested5.GM1-E100R2-PF-4m-full-nest
GM1 E100R2-PF-4m: full 4 m nest, every saved frame
10 single-time netCDF files: the whole 4 m moving nest (2 km x 2 km x 17.6 km) of a simulated supercell at every saved history time, 2280 s to 3360 s every 120 s.
What GM1 is
GM1 is a GPU (CUDA) port of CM1 release 21.1, the Cloud Model 1 of George Bryan (NSF NCAR). It keeps CM1's equations and numerics and adds GPU execution and a moving nest. GM1 is research code: its public release is gated on a stock-supercell… See the full description on the dataset page: https://huggingface.co/datasets/deepguess/GM1-E100R2-PF-4m-full-nest.
