datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chroma_promptsA collection of prompts captioned using Gemma 2b captioning model. These prompts are intended to be used with FLUX Chroma model.
Download .parquet files to your Google Drive and
run them using the .ipynb notebook in this repo
tsla-historic-priceswine-qualityharbor-benchus-zip-codes-demographics
Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics
Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24.
Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.warn-act-notice-type-codes-crosswalk
WARN Act notice-type codes — the crosswalk
Every US state publishes WARN Act layoff notices with a free-text column saying
what kind of event it is. The statute recognises two: a plant closing and a
mass layoff. Across 48 states that column contains
552 distinct exact strings (531
once you fold case).
This dataset is the crosswalk: one row per raw string, how many notices carry
it, which states emit it, and what it normalizes to.
The finding that matters
520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.code-sante-publique
Code de la santé publique, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sante-publique.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.pick_place_cup_20260918_221239This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/pick_place_cup_20260918_221239.python_codestyles-single-500
Dataset Card for "python_codestyles-single-500"
This dataset contains negative and positive examples with python code of compliance with a code style. A positive
example represents compliance with the code style (label is 1). Each example is composed of two components, the first
component consists of a code that either conforms to the code style or violates it and the second component
corresponding to an example code that already conforms to a code style. In total, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-single-500.jenny-mimi-codes
Jenny TTS — Mimi Codes
Pre-extracted Kyutai Mimi tokens for the
Jenny TTS Dataset —
a single female speaker, ~30h, clean studio-quality recordings. Apache-2.0 license.
Schema
Column
Type
Notes
id
string
e.g. jenny_0
text
string
mixed-case with punctuation
codes
int16[k=8][n_frames]
Mimi codebook indices @ 12.5 fps
n_frames
int32
k_codebooks
int32
8
No speaker_id column — single speaker dataset.
Extraction details
Source:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/jenny-mimi-codes.offsec_redteam_codes
OffSec RedTeam Codes
Token count: ~30B tokens.
OffSec RedTeam Codes is a curated corpus of code (and some auxiliary text) extracted from popular GitHub repositories related to offensive security / red teaming (pentesting, OSINT, C2, privilege escalation, exploitation, forensics, etc.). It is also the largest open-source dataset of red-team and offensive-security code ever compiled.
⚠️ Ethical use only. This dataset is for research, education, and defensive security testing in… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/offsec_redteam_codes.rollout_smolvla_pick_place_cup_20260919_130251This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/rollout_smolvla_pick_place_cup_20260919_130251.CodeScout_Training_Rolloutspython_codestyles-random-1k
Dataset Card for "python_codestyles-random-1k"
This dataset contains negative and positive examples with python code of compliance with a code style. A positive
example represents compliance with the code style (label is 1). Each example is composed of two components, the first
component consists of a code that either conforms to the code style or violates it and the second component
corresponding to an example code that already conforms to a code style. In total, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-random-1k.mls-mimi-codes
Multilingual LibriSpeech (MLS) — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for Multilingual LibriSpeech —
LibriVox audiobooks in 7 non-English languages.
English is intentionally excluded. For English Mimi codes, use:
shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits)
shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native)
shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents)
shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.rollout_smolvla_pick_place_cup_20260919_225813This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/rollout_smolvla_pick_place_cup_20260919_225813.python_codestyles-single-1k
Dataset Card for "python_codestyles-single-1k"
This dataset contains negative and positive examples with python code of compliance with a code style. A positive
example represents compliance with the code style (label is 1). Each example is composed of two components, the first
component consists of a code that either conforms to the code style or violates it and the second component
corresponding to an example code that already conforms to a code style. In total, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-single-1k.neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.python_codestyles-mixed1-1k
Dataset Card for "python_codestyles-mixed1-1k"
This dataset contains negative and positive examples with python code of compliance with a code style. A positive
example represents compliance with the code style (label is 1). Each example is composed of two components, the first
component consists of a code that either conforms to the code style or violates it and the second component
corresponding to an example code that already conforms to a code style.
The dataset combines both… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-mixed1-1k.python_codestyles-random-500
Dataset Card for "python_codestyles-random-500"
This dataset contains negative and positive examples with python code of compliance with a code style. A positive
example represents compliance with the code style (label is 1). Each example is composed of two components, the first
component consists of a code that either conforms to the code style or violates it and the second component
corresponding to an example code that already conforms to a code style. In total, the dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinityofspace/python_codestyles-random-500.rvlcdip-id-codes
Identification Codes in RVL-CDIP and Tobacco3482
Most pages in RVL-CDIP and Tobacco3482 contain an identification code: a Bates number or similar
stamp applied when the documents were produced in tobacco litigation. The codes were added for
record-keeping and are not part of a document's content, yet their position, format, and number
are strongly associated with the category label, so a classifier can use them as a shortcut. This dataset
gives the location and transcription of… See the full description on the dataset page: https://huggingface.co/datasets/stefan-hf/rvlcdip-id-codes.codeswitch-pairs-lase
Codeswitch Pairs LASE — training corpus
1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair).
Schema (manifest.jsonl)
{
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"lang": "en | hi | te | ta",
"text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.CodeScopeus-vehicle-trouble-codes
US vehicle diagnostic trouble codes, linked to manufacturer bulletins and owner reports
1,142 OBD-II diagnostic trouble codes as they actually appear in US manufacturer service bulletins and NHTSA
owner-complaint filings — with the model years, vehicle components and source documents each code shows up in.
Derived from public NHTSA filings. Regenerated nightly.
DOI: 10.5281/zenodo.22891283 — archived on Zenodo. That is the
concept DOI: it always resolves to the newest version… See the full description on the dataset page: https://huggingface.co/datasets/karsonmadden/us-vehicle-trouble-codes.pick_dino_toy_bowlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/pick_dino_toy_bowl.libritts-r-mimi-codes
LibriTTS-R — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for LibriTTS-R — a speech-restored version of
LibriTTS built specifically for TTS research.
Why LibriTTS-R instead of LibriSpeech?
LibriSpeech
LibriTTS-R
Purpose
ASR
TTS
Sample rate
16 kHz
24 kHz (Mimi-native, no resampling)
Segmentation
Arbitrary chunks
Sentence-level
Punctuation
Stripped (ALL CAPS)
Preserved
Audio quality
Raw amateur
Speech restoration appliedNo resampling is needed — 24 kHz… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/libritts-r-mimi-codes.rollout_move_blue_cube_red_area_20260823_100451This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/rollout_move_blue_cube_red_area_20260823_100451.rollout_move_blue_cube_red_area_20260823_112309This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/rollout_move_blue_cube_red_area_20260823_112309.rollout_move_blue_cube_red_area_20260823_094858This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Ishaan-codes11/rollout_move_blue_cube_red_area_20260823_094858.
