datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
dl_alchemy_seq9p6m_context1024less-is-moe-s1-calibration-128-seq8192
Less-is-MoE S1K calibration data — 128 samples, seq_length 8192
This is the fixed calibration artifact used to prune GPT-OSS-120B,
Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant:
yentinglin/s1K-1.1-trl-format revision
58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by
Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows.
For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.nca-paper-share20-seq_len_2048-657M
nca-paper-share20-seq_len_2048-657M
Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002.
Configuration
param
value
grid
12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share20-seq_len_2048-657M.dyck-k128-seq_len_2048-1B
dyck-k128-seq_len_2048-1B
Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256.
Grammar parameters
param
value
k (bracket types)
128
max_depth
16
p_open
0.5
seq_length
2048
file
split
tokens
train.bin
train
999,999,488
val.bin
val
10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.CultriX__SeQwence-14B-details
Dataset Card for Evaluation run of CultriX/SeQwence-14B
Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-details.flowzap-sequence-workflows
sequence-workflows
A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case.
Organization Model
Top-level folders are primary Use Cases from the FlowZap Templates dropdown.
Second-level folders preserve the original source domain from the FlowZap app index.
Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes.
Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.grpo-qwen3-1.7b-taco-easy-3200-bs32-n8-seqs16-32k-146102-rollouts
grpo_Qwen3-1.7B_TACO-easy-3200_bs32_n8_seqs16_32k_1epoch rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
SeqGeo-VL
🛰️ SeqGeo-VL: A Multimodal Cross-view Sequence Geo-localization Dataset
🌐 Project Page |
📄 Paper |
💻 GitHub |
🤗 Pretrained Weights
SeqGeo-VL is a sequential cross-view geo-localization dataset pairing satellite imagery, street-view videos, GPS trajectories, and natural-language route descriptions.
The dataset is introduced in TrajLoc: Trajectory-aware Cross-view Geo-localization with Sequential Observations, accepted at ECCV 2026.
⚠️ This repository contains annotations… See the full description on the dataset page: https://huggingface.co/datasets/MVRL/SeqGeo-VL.sequelbox__Llama3.1-8B-PlumCode-details
Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumCode
Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumCode
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumCode-details.prerequisite-sequencing-regulated-processes
What must pass before the next step: mandated step order in regulated processes
Canonical, always-current version: https://referencesource.org/prerequisite-sequencing-regulated-processes/
Machine-readable: https://referencesource.org/prerequisite-sequencing-regulated-processes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-11
Stale after: 2027-08-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/prerequisite-sequencing-regulated-processes.deepmath_103k_chat_train_seq8192SeqAtom-Coder-Dataset
SeqAtom Coder Dataset
SeqAtom Coder Dataset is the training dataset for:
tomjnet/SeqAtom-Coder-1.5B-Instruct
The dataset is designed to teach the SeqAtom model the Seq programming language, including syntax, execution flow, AI instructions, compiler behavior, and code translation.
Base Model
The initial SeqAtom model is based on:
Qwen/Qwen2.5-Coder-1.5B-Instruct
Dataset Format
The dataset uses conversational instruction examples.
Each record… See the full description on the dataset page: https://huggingface.co/datasets/tomjnet/SeqAtom-Coder-Dataset.CultriX__SeQwence-14B-v5-details
Dataset Card for Evaluation run of CultriX/SeQwence-14B-v5
Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B-v5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-v5-details.openvid-frame-sequences-1M
OpenVid Frame Sequences — 1M adjacent frame pairs
Short, single-shot frame sequences cut from OpenVid-1M,
built to train and evaluate models on what changes between two frames half a second apart.
One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs.
[f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09]
^ the thing you describe / predict
Sequences
116,596
Frames per sequence
10 (0.5 s apart, t = 0.0 … 4.5 s)
Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.CultriX__SeQwence-14Bv1-details
Dataset Card for Evaluation run of CultriX/SeQwence-14Bv1
Dataset automatically created during the evaluation run of model CultriX/SeQwence-14Bv1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14Bv1-details.nca-paper-seq_len_1024-164M
nca-paper-seq_len_1024-164M
Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002.
Configuration
param
value
grid
12×12
colors… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-seq_len_1024-164M.clinical-intervention-sequencing-and-state-control-v0.2
Clinical Multi-Evidence State Integration Benchmark
CMESI v0.2
The Clinical Multi-Evidence State Integration Benchmark (CMESI) evaluates whether an AI system can reconstruct the evolving state of a complex clinical case across a sequence of heterogeneous evidence events.
CMESI does not test whether a model can identify a diagnosis from a static vignette alone. It tests whether the model can:
maintain several competing clinical hypotheses simultaneously;… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-intervention-sequencing-and-state-control-v0.2.qwen35-08b-seqlen-ablation-0919i448-poc-datasetthermo_seq_instaffinity_seq_instomnimind-genomic-sequences-6sp
Genomic Sequences — 6 Species (20kb windows, Etapa C replication)
Sequências genômicas reais (NCBI/Ensembl) usadas no pipeline OmniMind Etapa C/D:
20.000 bp ACGT filtrados por espécie, mesmas regiões e janela da Etapa C
(nucleotide-transformer-v2-50m-multi-species).
Species
Chromosome
Source
Homo sapiens
22
NCBI
Saccharomyces cerevisiae
chrI
NCBI
Caenorhabditis elegans
chrI
NCBI
Drosophila melanogaster
2L
NCBI
Arabidopsis thaliana
1
NCBI
Escherichia coli… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-genomic-sequences-6sp.tuluv2_100k_seq
Dataset Card for Tuluv2-100k-Seq
It is a converted 100k subsample of Tulu-v2-sft-mixture (ODC-BY) using Seq-Instruct method from the paper SIT: Fine-tuning Large Language Models with Sequential Instructions
It includes massive sequential instructions which contains subtask more than one in each instructions.
Sequence-of-action-prediction-mind2websequelbox__gemma-2-9B-MOTH-details
Dataset Card for Evaluation run of sequelbox/gemma-2-9B-MOTH
Dataset automatically created during the evaluation run of model sequelbox/gemma-2-9B-MOTH
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__gemma-2-9B-MOTH-details.DeepSeek-7B-Sequential-Agent-Results
🤖 Advanced Sequential Processing Agent (DeepSeek 7B)
🖋️ Lead AI Engineer & Developer
Abdullah Ali Bahaaldeen
📌 Project Overview
This repository contains the automated execution pipeline and generated datasets from a highly optimized Sequential AI Agent. Built upon the deepseek-ai/deepseek-llm-7b-chat architecture, this robust system is engineered to operate seamlessly in severely constrained hardware environments (e.g., dual NVIDIA T4 GPUs or single-node… See the full description on the dataset page: https://huggingface.co/datasets/aab20abdullah/DeepSeek-7B-Sequential-Agent-Results.sequelbox__Llama3.1-8B-MOTH-details
Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-MOTH
Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-MOTH
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-MOTH-details.nca-paper-share10-seq_len_1024-164M
nca-paper-share10-seq_len_1024-164M
Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002.
Configuration
param
value
grid
12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share10-seq_len_1024-164M.
