datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_circleci_workflow_runsinfini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.cherrl-runsulvr_subset
ULVR stage-0 subsets (latent + source)
Curated, nested subsets of the Unified Visual Latent Reasoning (ULVR) stage-0
training data. Each subset folder is self-contained and ships both:
latent/ — pre-computed teacher latents, identical schema to
RuoliuYang/step0-all
source/ — the matching source samples (images + question/answer +
messages), identical schema to
RuoliuYang/ULVR_v2_clean
Latents and source rows are joinable by sample_id (within a category).
Folder… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ulvr_subset.FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.ditec-wdn--
Dataset Card for DiTEC-WDN
Dataset Summary
DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics.
Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series.
A node can represent a reservoir, junction, or tank, while… See the full description on the dataset page: https://huggingface.co/datasets/rugds/ditec-wdn.tokamark-p2-runssai-osworld-v2-benchmark-runs
Sai on OSWorld-V2 — benchmark runs of record
Two complete 108-task runs of the Sai computer-use agent on OSWorld-V2, with full
per-task evidence: scores, trajectories, agent runtime logs, evaluator logs, API
protocol logs, and run manifests.
Run
Date
Tasks scored
Mean score
Perfect (1.0)
Zeros
run1/
2026-08-12
108/108
0.7276
28
7
run2/
2026-08-20
108/108
0.7329
33
5
Model anthropic/claude-opus-5, thinking max, --max_steps 500, screenshot-only
observation… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v2-benchmark-runs.Public-YAM-runs
Public-YAM-runs
Physical bimanual YAM episodes recorded by the BluPe operator station.
Each run adds an episode to this repository. Failed, interrupted, stopped and
timed-out runs are retained and labeled; these are not all successful demonstrations.
A model saying done is not independently verified task success.
Loading
from datasets import load_dataset
runs = load_dataset("andlyu/Public-YAM-runs", split="train")
usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.benchmarking_sbi_runs
Benchmarking SBI Runs
This dataset contains the raw, per-run results underlying the manuscript
"Benchmarking Simulation-Based Inference"
(Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021).
It is a direct migration of the Git LFS data from
mackelab/benchmarking_sbi_runs on GitHub.
For compiled, ready-to-use dataframes built from these raw results (and the code that produced
them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.infini-news-index
INFINI-NEWS FM-Index
🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference).
Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built
with infini-gram-mini,
Liu et al. 2025) over the
ruggsea/infini-news-corpus
parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.rule-ling-conceptsflame-runsrunwayrelational-generator-runs
relational-generator-runs
Synthetic relational-database corpora, Relational Transformer checkpoints pretrained on them, and evaluation tables for the generator study. Everything here is generated data or derived from public RelBench; no private records.
social-sim-bench-gensgavel-runsrulerpx-run-storerl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.Indonesian-running-photos
Dataset Card for Fotoyu Album Archive
This dataset stores photo and video collections archived from Fotoyu albums using the potoyu-tree-downloader application. It is designed to act as a high-speed Cloudflare-backed CDN for serving static media assets, as well as providing a dataset for image/video classification and machine learning model training.
Dataset Details
Dataset Description
The dataset aggregates scraped photo galleries and video albums… See the full description on the dataset page: https://huggingface.co/datasets/TierKun/Indonesian-running-photos.ULVR_v2_clean
ULVR_v2_clean
Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation splits.
Every sample: input image + question -> assistant produces <abs_vis_token> + intermediate visual step(s) + \boxed{answer}.
subset
train
validation
text_cot
333,911
3,533
bbox_highlight
229,237
2,558
bbox_crop
229,237
2,558
depth
40,000
25
edge
40,000
14
segmentation
40,000
326
helper_interleaved
340,210
3,544
scene_graph
40… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ULVR_v2_clean.GheoLei_BeamNG.drive_Modsoeb-scored-runs
Open-Endedness Bench: scored runs
Every agent run scored in the paper Open-Endedness Bench: Measuring Epistemic
Process from Agent Records, with the output of each scoring stage and the full log
of judge requests and answers. The code is at
github.com/ARA-Labs/oeb.
Layout
out/posttrainbench/<task panel>/<unit>/ PostTrainBench runs (post-training gemma-3-4b on six
held-out tasks); a model's second run is <model>-r2… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/oeb-scored-runs.SPIN-UV
SPIN-UV
SPIN-UV is a multimodal dataset for unstructured scene understanding in dense urban villages. It was collected from a motor-driven, human-steered single-track vehicle and pairs front-facing visual observations with frame-anchored riding-state signals. The dataset is intended to support semantic segmentation, RGB-D perception, state-conditioned traversability, temporal consistency, and motion-aware scene understanding in narrow, weakly structured urban-village corridors.… See the full description on the dataset page: https://huggingface.co/datasets/ruikle123/SPIN-UV.russian-road-signs
Датасет размеченных знаков
Датасет размеченных дорожных знаков для задач компьютерного зрения и детекции объектов.
Загрузка
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="Dognellaf/russian-road-signs",
repo_type="dataset",
local_dir="./russian-road-signs"
)
Описание
Датасет содержит размеченные вручную кадры из видеозаписей с российскими дорожными знаками. Разметка в формате YOLO.
Изображений: 43 851 (JPEG)… See the full description on the dataset page: https://huggingface.co/datasets/Dognellaf/russian-road-signs.coat
Dataset Card for CoAT🧥
Dataset Description
CoAT🧥 (Corpus of Artificial Texts) is a large-scale corpus for Russian, which consists of 246k human-written texts from publicly available resources and artificial texts generated by 13 neural models, varying in the number of parameters, architecture choices, pre-training objectives, and downstream applications.
Each model is fine-tuned for one or more of six natural language generation tasks, ranging from paraphrase generation… See the full description on the dataset page: https://huggingface.co/datasets/RussianNLP/coat.rukopys
RUKOPYS: Ukrainian Handwritten Text Recognition Dataset
RUKOPYS (Ukrainian: рукопис — manuscript) is the first large-scale open dataset for Ukrainian handwritten text recognition (HTR). It spans over a century of Ukrainian handwriting — from 1920s archival documents to present-day school homework — and is designed for end-to-end document understanding: region detection, type classification, and text transcription.
Ukrainian is among the largest Slavic languages (45M+ native… See the full description on the dataset page: https://huggingface.co/datasets/UkrainianCatholicUniversity/rukopys.rlpinn-ablation-runs
RLPINN ablation runs
Логи и результаты запусков абляции DQN-стека RL-агента (PINNacle).
Буферы для этих запусков лежат в
danil-e/rlpinn-ablation-buffers.
Структура
runs/<pde>/<ablation>/<run_tag>/
params.json # гиперпараметры запуска
metrics.jsonl # по строке на лог метрик (шаг + ~30 метрик агента)
others.json
log.txt # полный stdout/stderr запуска
rl_model_snapshots/ # веса агента по шагам… See the full description on the dataset page: https://huggingface.co/datasets/danil-e/rlpinn-ablation-runs.sai-osworld-v21-reference-runs
Sai on OSWorld v2.1 — reference-VM runs
Every clean run of Sai (agent model anthropic/claude-opus-5, thinking effort max) on the 108 tasks of the
osworld-v2.1 release, on a cocoon microVM built from the official v2.1 reference image (Ubuntu 22.04.3).
134 runs across 108 tasks: some tasks ran more than once, and every clean run is here.
Results
108/108 tasks have at least one clean run.
Headline (mean over tasks of each task's mean clean score): 0.7731.
First… See the full description on the dataset page: https://huggingface.co/datasets/simular-ai/sai-osworld-v21-reference-runs.
