datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Total_Editing_Synthetic_Video_Albedo_Fulltotal-300-lambda00-s_signal_type6-jh-epoch4
total-300-lambda00-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3875
Action score: 0.43125
Valid samples: 320/320
total-300-lambda10-s_signal_type6-jh-epoch4
total-300-lambda10-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.41875
Valid samples: 320/320
total-300noapp-lambda02-s_signal_type6-jh-epoch4
total-300noapp-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36640625
Action score: 0.409375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-retry-epoch4
total-300-lambda02-s_signal_type6-jh-retry-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.36953125
Action score: 0.3984375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
total-300-lambda02-s_signal_type6-jh-epoch4-reeval1
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38828125
Action score: 0.4234375
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
total-300-lambda02-s_signal_type6-jh-epoch4-reeval2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4125
Action score: 0.4265625
Valid samples: 320/320
total-300-lambda05-s_signal_type6-jh-epoch4
total-300-lambda05-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.35703125
Action score: 0.4375
Valid samples: 320/320
total-300-lambda08-s_signal_type6-jh-epoch4
total-300-lambda08-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.38046875
Action score: 0.4078125
Valid samples: 320/320
total-300-lambda02-s_signal_type6-jh-epoch4
total-300-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.4046875
Action score: 0.4140625
Valid samples: 320/320
total-300app-lambda02-s_signal_type6-jh-epoch4
total-300app-lambda02-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3625
Action score: 0.4015625
Valid samples: 320/320
total-300-random-jh-epoch4
total-300-random-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3890625
Action score: 0.440625
Valid samples: 320/320
total-131-lambda02-residual-s_signal_type6-jh-epoch4
total-131-lambda02-residual-s_signal_type6-jh-epoch4
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3765625
Action score: 0.4171875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
totalsegmentator-organs
TotalSegmentator Organs Dataset
Dataset Description
The TotalSegmentator Organs dataset for multi-organ segmentation (TotalSegmentator Organs subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: adrenal glands, colon, duodenum, esophagus, gallbladder, kidneys, liver, lungs, pancreas, small bowel, spleen, stomach, trachea, bladder
Format: NIfTI (.nii.gz)
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-organs.Total-Text-DatasetTotal Text Dataset.
It consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved, one of a kind.
Original github repo; https://github.com/cs-chan/Total-Text-Dataset
Forked repo; https://github.com/yunusserhat/Total-Text-Dataset
toto_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 902,
"total_frames": 294139,
"total_tasks": 1,
"total_videos": 902,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:902"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/toto_lerobot.Total-Text-Dataset
Dataset Card for Total-Text-Dataset
The Total-Text consists of 1555 images with more than 3 different text orientations: Horizontal, Multi-Oriented, and Curved
This is a FiftyOne dataset with 1555 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Total-Text-Dataset.to-train-my-model
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/40u1d10t/to-train-my-model.TotalSegmentator-CT-Lite
About
This is a derivative of the TotalSegmentator dataset.
1228 CT images and corresponding segmentation mask of 117 structures
We combined multiple segmentation masks into a single nii.gz file under the folder Masks,
and moved all CT images to the folder Images.
All images and masks are renamed according to case IDs.
This dataset is released under the CC-BY-4.0 license.
News 🔥
[10 Oct, 2025] This dataset is integrated into 🔥MedVision🔥
Official… See the full description on the dataset page: https://huggingface.co/datasets/YongchengYAO/TotalSegmentator-CT-Lite.totoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 1003,
"total_frames": 325699,
"total_tasks": 1,
"total_videos": 1003,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1003"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/toto.Indic-total-New-TTS-Merge
Indic Total TTS Merge
Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration.
Languages
assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu
Columns
audio: Audio data
text: Transcript text
duration: Duration in seconds (all >= 3.0s)
language: Language name
lobster-data
LOBSTER Sample Data
LOBSTER (Limit Order Book System) sample L3 order book data for June 21, 2012.
Dataset Description
LOBSTER provides limit order book data from NASDAQ TotalView-ITCH messages.
AAPL: 1, 5, 10, 30, 50 levels
AMZN: 1, 5, 10 levels
GOOG: 1, 5, 10 levels
INTC: 1, 5, 10 levels
MSFT: 1, 5, 10, 30, 50 levels
SPY: 30, 50 levels
Configurations
Each ticker/level combination has its own configuration with two splits:
message: Order book events… See the full description on the dataset page: https://huggingface.co/datasets/totalorganfailure/lobster-data.Totalcap-blazeposeTotalsegmentor_Pelvis_Bone_Recon_Datasettotalsegmentator-vertebrae
TotalSegmentator Vertebrae Dataset
Dataset Description
The TotalSegmentator Vertebrae dataset for vertebrae segmentation (TotalSegmentator Vertebrae subset). This dataset contains CT scans with dense segmentation annotations.
Dataset Details
Modality: CT
Target: cervical, thoracic, and lumbar vertebrae
Format: NIfTI (.nii.gz)
Dataset Structure
Each sample in the JSONL file contains:
{
"image": "path/to/image.nii.gz",
"mask":… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/totalsegmentator-vertebrae.vuntum-robots
Vuntum robotics database
A sourced database of personal robots, humanoids and the companies that build them,
from Vuntum, an independent French and English media outlet on robots
with embedded AI. Every value is tied to a source, a level of evidence and a
verification date. Nothing is estimated silently: a missing value means "not
disclosed", never "no".
Licence: CC BY 4.0. Attribution: "Source: Vuntum (https://vuntum.com)" with a link to the sheet quoted.
Languages: French (fr… See the full description on the dataset page: https://huggingface.co/datasets/Totototoito/vuntum-robots.to-tool-call-datasets
🛠️ To-Tool-Call Datasets
A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training
To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention.
Quick Start ·
At a Glance ·
Format ·
Sources ·
Training Notes
[!IMPORTANT]
This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.
