datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sana-assetssynthesized_datasetefficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.efficient-cot
To be written.
efficient-agent-experiment-logs
Efficient Agent Experiment Campaigns
This snapshot preserves experiment records and artifacts from the
experiment-campaigns directory of the efficient-agent-research project.
It contains completed, failed, and interrupted runs. A directory or artifact
existing in this snapshot does not imply a valid official evaluation score.
Refer to each run's result and grader report for its status.
Contents
records/*.tar.gz: credential-redacted logs, Agent trajectories, code… See the full description on the dataset page: https://huggingface.co/datasets/VonEquinox/efficient-agent-experiment-logs.efficient-eval-h200-resultsSANA-WM-Bench
SANA-WM-Bench
Minimal 80-scene, 60s SANA-WM benchmark release. It contains the 80 conditioning images, fixed public scene set, simple/hard 60s Sana-WM camera trajectories. Scene IDs are anonymized within each category as 001-020. It intentionally excludes 24s settings, non-Sana-WM export formats, baseline outputs, and generated videos.
Image Provenance
All conditioning images in images/ are AI-generated using Google's Nano Banana
(Gemini native image generation).… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-WM-Bench.efficient-vla-runsVILA-inference-demosSANA-WM-example-training-dataset
SANA-WM Example Training Dataset
This repository contains the Sekai-Game training subset and LTX2 VAE latent cache used by the SANA-WM stage-1 example training configs.
License and Terms
This dataset is redistributed for non-commercial research use only under the same terms as the official Sekai project. See LICENSE and NOTICE.md.
The raw videos, camera trajectories, annotations, and derived VAE latent cache are all subject to the Sekai non-commercial… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-WM-example-training-dataset.vila-q-data-trainefficient-vla-extracts
efficient-vla-extracts
Parser-derived artifacts for the efficient-vla-wiki project.
This dataset is the data layer companion to the main repository:
GitHub: guanweifan/efficient-vla-wiki
Dataset repo: efficient-vla-extracts
What is inside
The staged upload keeps the local directory structure:
extracts/
├── meta/
│ ├── extract_build_index.jsonl
│ └── extract_build_status.json
└── parses/
└── <paper_id>/
Each paper directory may contain:
pdftotext.txt… See the full description on the dataset page: https://huggingface.co/datasets/guanweifan/efficient-vla-extracts.ai2d-no-maskZ1-Code-Reasoning-107K
Z1: Efficient Test-time Scaling with Code
Train Large Language Model to Reason with Shifted Thinking
[📜 Paper] •
[🤗 HF Models] •
[🐱 GitHub]
Details
Please refer to https://github.com/efficientscaling/Z1.
Usage
from datasets import load_dataset
ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"]
ds[0]
Citation
@misc{yu2025efficientscaling,
title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.neuralatlas-attributions-efficientnet_b4
Neural Atlas attributions — efficientnet_b4 on imagenet-pico
Precomputed attribution maps and faithfulness metrics for the torchvision
efficientnet_b4 model (default pretrained weights, no fine-tuning) on imagenet-pico,
a 3000-image subset of ImageNet-1k with three images for each of the 1000
classes.
This repository is part of Neural Atlas, a web tool for comparing
attribution methods across vision architectures on the same image, developed
as an undergraduate thesis at the… See the full description on the dataset page: https://huggingface.co/datasets/Matgc04/neuralatlas-attributions-efficientnet_b4.worldmodelbenchkey_pointdbpedia-entities-efficient-splade-100K
DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding
This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/
Model id used to make these vectors:
model_id = "naver/efficient-splade-VI-BT-large-doc"
For processing the query, use this:
model_id = "naver/efficient-splade-VI-BT-large-query"
If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.E2AM_EfficientNetV2_S
E2AM Ablation Results: EfficientNetV2-S
Energy-aware training ablation study for EfficientNetV2-S across three image-classification datasets: CIFAR-10, CIFAR-100, and Tiny-ImageNet.
Each dataset has 15 training variants (8 individual-method M0..M7, 7 cumulative ablation C0..C6) at 50 epochs, plus a 5-variant deployment pipeline (FP32 baseline, structured pruning, pruning+finetune, INT8 quantization, pruned+INT8).
Status: 45 completed variants, 0 partial.
Quick links… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/E2AM_EfficientNetV2_S.EfficientDatasetCondensation
Data summary
This repository contains small synthetic data for Image datasets; MNIST, SVHN, and CIFAR-10.
Each torch file contains the images and corresponding labels of sizes ranging from 1,10,50 images per class (IPC).
For more details, please refer to our GitHub page and paper below.
Reference
https://github.com/snu-mllab/Efficient-Dataset-Condensation
Citation
@inproceedings{kimICML22,
title = {Dataset Condensation via Efficient Synthetic-Data… See the full description on the dataset page: https://huggingface.co/datasets/ICML2022/EfficientDatasetCondensation.neuralatlas-attributions-efficientnet_b0
Neural Atlas attributions — efficientnet_b0 on imagenet-pico
Precomputed attribution maps and faithfulness metrics for the torchvision
efficientnet_b0 model (default pretrained weights, no fine-tuning) on imagenet-pico,
a 3000-image subset of ImageNet-1k with three images for each of the 1000
classes.
This repository is part of Neural Atlas, a web tool for comparing
attribution methods across vision architectures on the same image, developed
as an undergraduate thesis at the… See the full description on the dataset page: https://huggingface.co/datasets/Matgc04/neuralatlas-attributions-efficientnet_b0.dclm-train-1.64m-tsp-efficientSANA-Streaming-example-training-dataset
SANA-Streaming Example Training Dataset
This repository contains 1,000 aligned reverse video-editing pairs for the
public SANA-Streaming bidirectional V2V training recipe.
License and Terms
This dataset is made available for non-commercial research use only under the
terms in LICENSE. See NOTICE.md for the redistributed content covered by
those terms.
The videos, prompts, annotations, and metadata are all subject to the
non-commercial research-only terms. Do not… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/SANA-Streaming-example-training-dataset.LongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
toy_datavideo_toy_datasana_data_publicbosnian-corpus-v1
Bosnian Corpus v1.0 (cleaned)
This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian,
designed for quantitative linguistic analysis, information-theoretic research,
entropy estimation, corpus linguistics, language modeling, and modern NLP tasks.
The canonical release of the corpus is archived on Zenodo:
Dataset DOI:https://doi.org/10.5281/zenodo.17757098
Corpus composition
The corpus is constructed from three publicly available… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.dpsk-r1-labeling
