datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bartholomew-dataset-v3
BART Dataset v3
The final version of the BART pretraining corpus, focused on removing anything that betrays a
post-1930 origin. This is our best vintage dataset yet.
Documents
146,031 (97.52% of v2)
Characters
102,798,688,961 (96.73% of v2)
Tokens
~23B (estimated)
Shards
473 (one per v2 shard, same basename)
Source
BART Dataset v2
Cutoff
1930
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v3.variouscryptodata
variouscryptodata
Crypto market datasets collected as a by-product of our own research and
published so they are not lost. One sub-folder per dataset; each appended
nightly where collection is still running.
folder
what
coverage
cadence
polymarket_updown_orderbook/
Polymarket Up/Down (5m/15m) order books, 10 levels, BTC/ETH/SOL/XRP/DOGE/HYPE/BNB, with Binance spot reference
2026-05-24 → present
appended nightly (previous UTC day)
hyperliquid_trades/
Hyperliquid perp… See the full description on the dataset page: https://huggingface.co/datasets/Barthel/variouscryptodata.bart-midtrain
BART Midtrain
The midtraining corpus for BART —
pre-1930 mathematics, science, technology, and medicine — plus the full pipeline that built it and
every training mixture it was blended into.
Built by Unbounded Labs.
Corpus documents
11,409
Corpus characters
2,543,809,124
Corpus tokens
~604M
Removed by cleaning
24% of documents (15,075 → 11,409)
Subject focus
math, science, technology, medicine
Cutoff
1930
Schema
single string column text… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/bart-midtrain.rayst3rbartholomew-dataset-v2
BART Dataset v2
The second version of the BART pretraining corpus, focused on stripping low-quality text —
boilerplate and OCR corruption — out of
v1.
Documents
149,745 (93.44% of v1)
Characters
106,274,384,672 (89.50% of v1)
Tokens
~24B (estimated)
Shards
473 (one per v1 shard, same basename)
Source
BART Dataset v1
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:
Institutional Books 1.0 →
v1 →
v2 →
v3… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v2.EventHubDatasetvcc2026-broad-development-20260908
VCC 2026 broad development artifacts
The B7 RunPod science-output recovery is complete: 256 files, 5,106,104,214 bytes.
Contents
Files
Checkpoints
64
Fit receipts
64
Prediction arrays
64
Prediction receipts
64
Use archive-index.json for current locations, SHA-256 values, sizes, and immutable per-file revisions. STATUS.md records completed work and remaining scientific comparisons. The recovery receipt proves all 109 files missing from the previous archive… See the full description on the dataset page: https://huggingface.co/datasets/barthazian/vcc2026-broad-development-20260908.bartholomew-dataset-v1
BART Dataset v1
The first version of the BART pretraining corpus: pre-1930 English books drawn from
Institutional Books 1.0
and filtered hard on OCR quality, language, date, and tokenizability.
Documents
160,263
Characters
118,745,375,871
Tokens
~27B (estimated)
Shards
473 (472 train + 1 val)
Source
Institutional Books 1.0 (242B tokens, ~983K documents)
Schema
single string column text
Lineage — three cumulative filtering stages over the same corpus:… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/bartholomew-dataset-v1.speech_commands
Dataset Card for "speech_commands"
More Information needed
bartenderkaminoglass
Bangumi Image Base of Bartender: Kami No Glass
This is the image base of bangumi Bartender: Kami no Glass, we detected 26 characters, 3350 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/bartenderkaminoglass.SynGallery-1024
SynGallery-1024: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
The 1024×1024 high-resolution edition of
patryk-bartkowiak/SynGallery.
A synthetic dataset for instance-level artwork recognition: 4,898 real
paintings (MET Open Access) hung in a procedurally randomized 3D art-gallery
scene, each rendered from 5 camera viewpoints at 1024×1024 — 24,490
synthetic RGB images paired with their source photos and museum metadata
(title, artist, date, medium… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-1024.tape_to_bin2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 4896,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/tape_to_bin2.dice4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 6198,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/dice4.tape_to_binThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 739,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/tape_to_bin.dice2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 3660,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bartm3/dice2.wikipedia-semantic-searchCompanion dataset to https://bart.degoe.de/building-a-semantic-search-engine-in-250-lines-of-python and https://github.com/bartdegoede/python-searchengine. These chunks contain Wikipedia articles and their embeddings (all-MiniLM-L6-v2). The .npy files are NumPy memory mapped files.
SynGallery-abl4-tex-light-glass-frame
SynGallery-abl4-tex-light-glass-frame: + frame variety
Rung 4 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting, glass and frame molding variant + color/roughness/metallic, while freezing camera pose (the only frozen factor). Same schema, source images and index↔painting… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl4-tex-light-glass-frame.SynGallery
SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
A synthetic dataset for instance-level artwork recognition: 4,898 real
paintings (MET Open Access) hung in a procedurally randomized 3D art-gallery
scene, each rendered from 5 camera viewpoints at 512×512 — 24,490
synthetic RGB images paired with their source photos and museum metadata
(title, artist, date, medium, …). The environment is randomized per scene — wall/floor/roof textures… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery.object-placing-on-Sutton-Barto-30-ep-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ssabats/object-placing-on-Sutton-Barto-30-ep-merged.SynGallery-abl1-tex
SynGallery-abl1-tex: + texture variety
Rung 1 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies wall/floor/roof textures + floor material (on top of the baseline), while freezing lighting, glass, frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl1-tex.SynGallery-abl0
SynGallery-abl0: fixed-environment baseline
Rung 0 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies nothing — the gallery is one fixed configuration for all 24,490 images, while freezing wall/floor/roof textures + floor material, lighting, glass, frame variant/color, camera pose. Same schema… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl0.object-placing-on-Sutton-Barto_20260916_201400This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ssabats/object-placing-on-Sutton-Barto_20260916_201400.GEM__bart_base_schema_guided_dialog__1645547915Leo__bart-large__1645784880summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1SynGallery-abl3-tex-light-glass
SynGallery-abl3-tex-light-glass: + glass
Rung 3 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting and a glass sheet present with probability 0.25, while freezing frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every other rung — they… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl3-tex-light-glass.object-placing-on-Sutton-Barto-60-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ssabats/object-placing-on-Sutton-Barto-60-merged.bartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.SynGallery-abl2-tex-light
SynGallery-abl2-tex-light: + lighting variety
Rung 2 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures and lighting (area-light shape/spread + instanced ceiling lights), while freezing glass, frame variant/color, camera pose. Same schema, source images and index↔painting mapping as… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl2-tex-light.MAPLE-Lua-Corpus
