datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encoder-decoder-floresp-scorestemp-decoder-train-tokenizedencoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.SylReg-Decodernav2tex-decoder-latex-pretrainptm-relevance-aware-decoder-qwen3-14b
Relevance-aware horizon decoder: Qwen3-14B activations with irrelevant durations
Residual-stream activations of Qwen3-14B on 5,520 short decision prompts, captured for the relevance-aware
horizon decoder experiment of the SPAR (fall 2026) project Planning Temporal Manifolds (PTM). Half the prompts
contain one extra sentence with a duration that is irrelevant to the decision. The data lets you ask whether a linear
readout of the model's time horizon can be made to ignore such… See the full description on the dataset page: https://huggingface.co/datasets/anicola/ptm-relevance-aware-decoder-qwen3-14b.dream-decoder-dataset
Dream Decoder Synthetic Dataset
Size: 1,200 examplesModality: Text (dream_text, interpretation)Fields: id, dream_text, interpretation, symbols, emotions, setting, actions, tags, source
How it was created
Base data generated with templated combinations (symbols, emotions, settings, actions).
~300 dreams were paraphrased with google/flan-t5-base to satisfy the "use a HF model" requirement.
Intended use
For demo/building a dream similarity & recommendation app.… See the full description on the dataset page: https://huggingface.co/datasets/samvlad/dream-decoder-dataset.calcium-decoder-mosaic
Sub-cellular calcium imaging movies with organelle stains and mosaic CaMKII and NFAT reporters
3,800 simulated time-lapse fluorescence microscopy sessions of fields of 15 to 40 coupled human cells: a calcium-indicator movie with sub-cellular calcium gradients, nuclear / mitochondria / ER stain images, and mosaic NFAT-GFP and CaMKII FRET reporter images. Each session has the exact per-cell CaMKII activity and NFAT nuclear fraction at six checkpoints, one seed point per cell, the… See the full description on the dataset page: https://huggingface.co/datasets/amirmmahdavikia/calcium-decoder-mosaic.fintime-decoder-dataset
Dataset Card for FinTime Dataset
Dataset Summary
FinTime Dataset is a comprehensive, large-scale financial time series dataset designed for training and evaluating decoder-only models on financial forecasting tasks across diverse asset classes and market conditions.
Composed of real-world market data spanning equities, cryptocurrencies, forex, commodities, and indices, the dataset captures the complexity, volatility, and multi-scale dynamics typical of financial markets.… See the full description on the dataset page: https://huggingface.co/datasets/thesven/fintime-decoder-dataset.riemannian-generative-decoder
Riemannian generative decoder dataset
This repository contains the data related to the paper "Riemannian generative decoder".
Project Page: https://yhsure.github.io/riemannian-generative-decoder
Code Repository: https://github.com/yhsure/riemannian-generative-decoder
Abstract
Riemannian representation learning typically relies on an encoder to estimate densities on chosen manifolds. This involves optimizing numerically brittle objectives, potentially harming model… See the full description on the dataset page: https://huggingface.co/datasets/yhsure/riemannian-generative-decoder.pet-decoder-audio-fixtures
Pet Decoder Audio Test Fixtures (v1.0) 🧪
This repository contains audio test fixtures used for validating the ingestion pipeline of the Pet Decoder AI application.
These are verified, public domain samples used to test our audio visualization and classification algorithms against known baselines (e.g., Low Frequency vs. High Frequency vocalizations).
Dataset Contents
The dataset consists of 5 reference audio files representing distinct spectral patterns:… See the full description on the dataset page: https://huggingface.co/datasets/petdecoder/pet-decoder-audio-fixtures.MixAtis_for_DecoderOnly
Dataset Card for "MixAtis_for_DecoderOnly"
More Information needed
low-decoder
low-decoding
Author: Surpem
This dataset contains 1200 unique, clean synthetic audio signals representing decoded text commands.
The audio signals represent synthesized Morse Code message blocks.
Dataset Structure
id: A unique UUID string.
audio: The audio wav bytes (16kHz Mono).
text: The decoded string transcription.
Dataset Level: LOW
Low: Slow WPM (~12 WPM), high signal-to-noise ratio (clean), short command strings.
Medium: Fast WPM (~24 WPM), background… See the full description on the dataset page: https://huggingface.co/datasets/Surpem/low-decoder.MixSnips_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split-HALF"
More Information needed
pantheon-ui-decoder-conversations
Pantheon UI Decoder Conversations
Training dataset for the decoder half of the Pantheon UI round-trip translator. The encoder turns natural language into emoji; the decoder takes emoji back to natural language.
Inspired by Anthropic's Natural Language Autoencoders — emoji as a discrete, human-legible intermediate between two model passes.
How it was built
Each row is derived from shreyask/pantheon-ui-conversations by inverting the encoder pairs:
Encoder pair:… See the full description on the dataset page: https://huggingface.co/datasets/shreyask/pantheon-ui-decoder-conversations.decoderstack-gsm8k
decoderstack-gsm8k
GSM8K, pre-tokenized for stacks/decoder-rtx/train_gsm8k.py (the RL sanity-check
pipeline of the DecoderStack backward-pass speedrun): ClimbMix 32k ids with the
nanochat chat template already applied. The trainer downloads these files and never
tokenizes.
file
rows
what
prompts.parquet
7473 train + 1319 test
[bos, user_start, *question, user_end, assistant_start], gold answer, is_val (256 seeded test problems = the in-loop validation tracker)… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/decoderstack-gsm8k.eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 10,
"total_frames": 3498,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 1,
"video_files_size_in_mb": 1,
"fps": 15,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/continuallearning/eval_dit_posttrainv2_union_norm_bs128_seqlora_dit_all_decoder_real_1_on_real_1_seed1000.MixAtis_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split-HALF"
More Information needed
mineru-decoder-finetune-dataset
MinerU Decoder Fine-Tune Dataset (real data)
Note: the dataset viewer is disabled — the data is distributed as zip bundles, not a
HuggingFace-loadable table. Download + unpack with the reproduce script (below), don't use load_dataset.
68,798 real image→tagged-text samples for teaching a VLM to emit inline formatting tags
(bold, italic, sup, sub, underline, strike). Sources: 22 arXiv categories (bold/italic/sup/sub) +
Washington & Texas legislative bills (underline/strike).… See the full description on the dataset page: https://huggingface.co/datasets/ahamad-ai/mineru-decoder-finetune-dataset.decoder-lang-200k#decoder-lang-200k
200,000 instruction of language and programming
Format: JSONL
result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta
Dataset Card for "result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta"
More Information needed
MixSnips_for_DecoderOnly
Dataset Card for "MixSnips_for_DecoderOnly"
More Information needed
prescription-decoder-dataset-cleanprescription-decoder-datasetL7_8_sampled_news_headlines_decoder_classificationdecoder_patch_attributionDecoderLoraMixAtis_for_DecoderOnly_90-10_split
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split"
More Information needed
MixSnips_for_DecoderOnly_90-10_split
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split"
More Information needed
finetune_decoder
