datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoboReward
RoboReward
Links: Paper · RoboRewardBench Leaderboard
RoboReward is a dataset for training and evaluating general-purpose vision-language reward models for robotics. Each example pairs a task instruction with a real-robot rollout video and a discrete end-of-episode progress reward score in {1,…,5}.
RoboReward is built from large-scale real-robot corpora including Open X-Embodiment (OXE) and RoboArena. Because OXE is success-heavy, we generate additional negatives and near-misses… See the full description on the dataset page: https://huggingface.co/datasets/teetone/RoboReward.raw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character.
language:
- en
size_categories:
- n<1K
Teeth3DSCommonVoice_teensmedical-imaging-combined
Combined Medical Imaging Dataset
Dataset Description
Combined medical imaging dataset with 6793 samples in Alpaca instruction format.
Dataset Statistics
Total Samples: 6793
Training Samples: 5434
Validation Samples: 1359
Modality Distribution
X-ray: 2691 samples
CT: 2257 samples
Unknown: 1329 samples
MRI: 369 samples
Ultrasound: 147 samples
Source Distribution
ROCO: 5000 samples
VQA-RAD: 1793 samples
Sources
ROCO… See the full description on the dataset page: https://huggingface.co/datasets/teeyouteeyou/medical-imaging-combined.ChartQADataset is converted from https://github.com/vis-nlp/ChartQA
vin là tập đã dịch các qa3000 là tập các chart đã dịch
Disclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use.
college-roi-data
College ROI Data — what U.S. colleges and majors actually pay back
Clean, citable tables on the lifetime financial return of U.S. colleges and majors —
30-year net present value by school and state, ROI by major category, the out-of-state
premium, and how exposed each major's career paths are to today's AI. Maintained by
LE TEEN, a college-ROI data project. Every number traces to a
public source; nothing is scraped, modeled behind closed doors, or vibes.
The headline the… See the full description on the dataset page: https://huggingface.co/datasets/le-teen/college-roi-data.all_possible_scans_for_4_5_6_teethcommon_voice_17-pl-speakers-v1.0Available sources(voices):
tomasz
konrad
jan
wojciech
krystian
emilia
kacper
karol
henryk
marek
kinga
szymon
robert
milena
jagoda
marcel
edmund
weronika
maciej
hubert
aniela
artur
stefan
ireneusz
grzegorz
roman
leon
maksymilian
zygmunt
jerzy
piotr
edward
bernard
aleksander
arkadiusz
zuzanna
melania
ignacy
lucjan
iga
patryk
bogdan
adam
dominika
adrian
antoni
bartosz
marian
cezary
ludwik
zenon
ryszard
feliks
filip
franciszek
mateusz
witold
igor
marcin
dominik
julian
kazimierz
sebastian
mariusz… See the full description on the dataset page: https://huggingface.co/datasets/TeeZee/common_voice_17-pl-speakers-v1.0.Pokemon-Captioning-ClassificationDisclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use.
Ornix-data
Ornix-data — Vietnamese TTS (multi-speaker)
Vietnamese TTS corpus, 55,635 clips / 82.02 h, 24 000 Hz mono PCM16
WAV. Filter speaker_id to select a speaker.
speaker_id
gender
source repo
revision
clips
hours
Ngọc Huyền + Hà Trang
female
Teedyyy-rm/Ornix-data (rev 2c7ea5c7) + quocs/hatrang-voice-4h
mixed
10,363
19.07
Hoài Vi
female
TuNguyen1101/luna_vi
af4917d1
30,407
44.68
Hoài Thu
female
vuihocrnd/vi_voice_female
72c02b81
11,012
12.84
Thu Minh
female… See the full description on the dataset page: https://huggingface.co/datasets/Teedyyy-rm/Ornix-data.teenytinystories
TeenyTinyStories (TTS)
A second-order TinyStories-style corpus: short, simple English stories generated by a small open model rather than by a frontier model.
Provenance
Generator: roneneldan/TinyStories-Instruct-33M (GPT-Neo architecture, 4 layers x 768 wide), run from random-free inference with the GPT-Neo tokenizer (EleutherAI/gpt-neo-125M).
Seeding: each story is conditioned on a structured header with three words (verb, noun, adjective, in the order used by… See the full description on the dataset page: https://huggingface.co/datasets/ableian/teenytinystories.camoscio_cleaned
Dataset Card for "camoscio_cleaned"
More Information needed
dolly-15k-pirate-speechDataset for writing style transfer experimentation based on article:
https://ai-r.com/blog/pirate-linguistics-and-tone-of-voice-fine-tuning-llms-to-talk-like-swashbucklers
Only responses are in 'pirate speech'
arrr python library was used to simply change original responses to 'pirate speech' responses
https://pypi.org/project/arrr/
common_voice_17-pl-speakers-v2.1teememo-eq-bench-ja
EQ-Bench3 日本語化パッチ (teememo-eq-bench-ja)
EQ-Bench3 を日本語で評価するための翻訳データと差分パッチのセットです。
リポジトリ構成
teememo-eq-bench-ja/
├── apply_patch.sh # ← パッチ適用スクリプト(これを実行するだけ)
└── data/
├── benchmark.patch # core/benchmark.py 差分 (Fix-B1)
├── conversation.patch # core/conversation.py 差分 (Fix-C1)
│
├── scenario_prompts_ja.txt # シナリオプロンプト (日本語)
├── scenario_prompts_en.txt # シナリオプロンプト (英語・参照用)
├──… See the full description on the dataset page: https://huggingface.co/datasets/YUGOROU/teememo-eq-bench-ja.common-voice-22-en-teens
Common Voice 22 - English Teens Subset
This dataset is a filtered subset of Mozilla's Common Voice Corpus 22.0, containing only recordings from teenage speakers in English.
Dataset Structure
Data Instances
Each instance contains:
audio: Audio file with embedded bytes, sampling rate, and path
text: Transcription of the spoken text
Data Splits
This dataset provides a single unified split containing all 77,500 samples. Users can create their own… See the full description on the dataset page: https://huggingface.co/datasets/cryptolock/common-voice-22-en-teens.medical-research-clean-embeddingsnEMO-speakersfleurs-pl-v1.0teememo-pref-data
TeenEmo Preference (DPO/SimPO) Dataset
Synthetic Japanese counseling dataset for training TeenEmo —
an on-device AI counseling app for teenagers.
Item
Value
Records
2,179
Language
Japanese
Kind
Preference (DPO/SimPO)
Target model
LiquidAI/LFM2.5-1.2B-Base
Synthesized by
LiquidAI/LFM2-24B-A2B
Created
2026-03-17 04:51 UTC
Format
from datasets importload_dataset
ds = load_dataset("YUGOROU/teememo-pref-data", split="train")
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/YUGOROU/teememo-pref-data.Code_Opt_Triton
Overview
This dataset, TEEN-D/Code_Opt_Triton, is an extended version of the publicly available GPUMODE/Inductor_Created_Data_Permissive dataset. It contains pairs of original (PyTorch or Triton) programs and their equivalent Triton code (generated by torch inductor), intended for training models in PyTorch-to-Triton code translation and optimization.
The primary modification in this extended version is that each optimized Triton code snippet is paired with both its original source… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton.wan_brushing_teethThis dataset contains videos generated using Wan 2.1 T2V 14B.
grpo-oumi-synthetic-document-claims
Dataset Card for GRPO Oumi ANLI Subset
Dataset
This dataset is a reformatted version of the oumi-ai/oumi-synthetic-document-claims dataset, specifically structured for use with the GRPO trainer.
You can find more detailed information about the original dataset at the provided link.
Link: https://huggingface.co/datasets/oumi-ai/oumi-synthetic-document-claims
Dataset Structure
The dataset consists of a list of dictionaries, where each dictionary represents a… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/grpo-oumi-synthetic-document-claims.nutrition5k-food-name-geminiBaKaTeQeeurope-ilo-ees-tees-sex-ind-jbc-nb-employees-by-ilo-sector-and-sex-and-type-of-job-co
Employees by ILO sector and sex and type of job contract (thousands) | Europe (ILOSTAT)
🇪🇺 44,599 observations · 15 Europe countries · 2001–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 44,599 observations of Employees data across 15 Europe countries, spanning 2001–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-ind-jbc-nb-employees-by-ilo-sector-and-sex-and-type-of-job-co.camoscio
Camoscio instruction-tuning dataset
This repository contains the dataset used to train Camoscio.
This dataset is an Italian translation with ChatGPT of the Stanford Alpaca dataset.
Please refer to the Camoscio repo for more info.
nEMO-speakers-v2.0-emotion-detection-v2.0common_voice_17-pl-speakers-v2.0
