datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.SylReg-Decodernav2tex-decoder-latex-pretrainptm-relevance-aware-decoder-qwen3-14b
Relevance-aware horizon decoder: Qwen3-14B activations with irrelevant durations
Residual-stream activations of Qwen3-14B on 5,520 short decision prompts, captured for the relevance-aware
horizon decoder experiment of the SPAR (fall 2026) project Planning Temporal Manifolds (PTM). Half the prompts
contain one extra sentence with a duration that is irrelevant to the decision. The data lets you ask whether a linear
readout of the model's time horizon can be made to ignore such… See the full description on the dataset page: https://huggingface.co/datasets/anicola/ptm-relevance-aware-decoder-qwen3-14b.dream-decoder-dataset
Dream Decoder Synthetic Dataset
Size: 1,200 examplesModality: Text (dream_text, interpretation)Fields: id, dream_text, interpretation, symbols, emotions, setting, actions, tags, source
How it was created
Base data generated with templated combinations (symbols, emotions, settings, actions).
~300 dreams were paraphrased with google/flan-t5-base to satisfy the "use a HF model" requirement.
Intended use
For demo/building a dream similarity & recommendation app.… See the full description on the dataset page: https://huggingface.co/datasets/samvlad/dream-decoder-dataset.fintime-decoder-dataset
Dataset Card for FinTime Dataset
Dataset Summary
FinTime Dataset is a comprehensive, large-scale financial time series dataset designed for training and evaluating decoder-only models on financial forecasting tasks across diverse asset classes and market conditions.
Composed of real-world market data spanning equities, cryptocurrencies, forex, commodities, and indices, the dataset captures the complexity, volatility, and multi-scale dynamics typical of financial markets.… See the full description on the dataset page: https://huggingface.co/datasets/thesven/fintime-decoder-dataset.riemannian-generative-decoder
Riemannian generative decoder dataset
This repository contains the data related to the paper "Riemannian generative decoder".
Project Page: https://yhsure.github.io/riemannian-generative-decoder
Code Repository: https://github.com/yhsure/riemannian-generative-decoder
Abstract
Riemannian representation learning typically relies on an encoder to estimate densities on chosen manifolds. This involves optimizing numerically brittle objectives, potentially harming model… See the full description on the dataset page: https://huggingface.co/datasets/yhsure/riemannian-generative-decoder.MixAtis_for_DecoderOnly
Dataset Card for "MixAtis_for_DecoderOnly"
More Information needed
pantheon-ui-decoder-conversations
Pantheon UI Decoder Conversations
Training dataset for the decoder half of the Pantheon UI round-trip translator. The encoder turns natural language into emoji; the decoder takes emoji back to natural language.
Inspired by Anthropic's Natural Language Autoencoders — emoji as a discrete, human-legible intermediate between two model passes.
How it was built
Each row is derived from shreyask/pantheon-ui-conversations by inverting the encoder pairs:
Encoder pair:… See the full description on the dataset page: https://huggingface.co/datasets/shreyask/pantheon-ui-decoder-conversations.low-decoder
low-decoding
Author: Surpem
This dataset contains 1200 unique, clean synthetic audio signals representing decoded text commands.
The audio signals represent synthesized Morse Code message blocks.
Dataset Structure
id: A unique UUID string.
audio: The audio wav bytes (16kHz Mono).
text: The decoded string transcription.
Dataset Level: LOW
Low: Slow WPM (~12 WPM), high signal-to-noise ratio (clean), short command strings.
Medium: Fast WPM (~24 WPM), background… See the full description on the dataset page: https://huggingface.co/datasets/Surpem/low-decoder.MixSnips_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split-HALF"
More Information needed
decoderstack-gsm8k
decoderstack-gsm8k
GSM8K, pre-tokenized for stacks/decoder-rtx/train_gsm8k.py (the RL sanity-check
pipeline of the DecoderStack backward-pass speedrun): ClimbMix 32k ids with the
nanochat chat template already applied. The trainer downloads these files and never
tokenizes.
file
rows
what
prompts.parquet
7473 train + 1319 test
[bos, user_start, *question, user_end, assistant_start], gold answer, is_val (256 seeded test problems = the in-loop validation tracker)… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/decoderstack-gsm8k.MixAtis_for_DecoderOnly_90-10_split-HALF
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split-HALF"
More Information needed
result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta
Dataset Card for "result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta"
More Information needed
decoder-lang-200k#decoder-lang-200k
200,000 instruction of language and programming
Format: JSONL
prescription-decoder-dataset-cleanprescription-decoder-datasetL7_8_sampled_news_headlines_decoder_classificationnios_nllb_decoder_biasedMixSnips_for_DecoderOnly
Dataset Card for "MixSnips_for_DecoderOnly"
More Information needed
DecoderLoraMixAtis_for_DecoderOnly_90-10_split
Dataset Card for "MixAtis_for_DecoderOnly_90-10_split"
More Information needed
MixSnips_for_DecoderOnly_90-10_split
Dataset Card for "MixSnips_for_DecoderOnly_90-10_split"
More Information needed
Encoder-DecoderDoctor_minifinetune_decoderDecoder_onlyDecoder-OnlyE31_decoder_gen_fsqnomic-decoder-data
