Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hi-todayis-jh /l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes461 downloads9d agoHugging Face02hi-todayis-jh /f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes353 downloads8d agoHugging Face03kothasuhas /dl_alchemy_seq9p6m_context1024tabularn<1K0 likes312 downloads1mo agoHugging Face04jayzou3773 /less-is-moe-s1-calibration-128-seq8192 Less-is-MoE S1K calibration data — 128 samples, seq_length 8192 This is the fixed calibration artifact used to prune GPT-OSS-120B, Qwen3.5-122B-A10B, and the Gemma-4-26B-A4B causal language tower. It uses the same 128 source rows as the full-length variant: yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251, train, followed by Dataset.shuffle(seed=1234) and the first 128 nonempty messages rows. For pruning, concatenate messages[].content with one… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-s1-calibration-128-seq8192.tabulartext-generationn<1K0 likes149 downloads17d agoHugging Face05alexkstern /nca-paper-share20-seq_len_2048-657M nca-paper-share20-seq_len_2048-657M Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002. Configuration param value grid 12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share20-seq_len_2048-657M.tabularn<1K0 likes146 downloads3mo agoHugging Face06alexkstern /dyck-k128-seq_len_2048-1B dyck-k128-seq_len_2048-1B Procedurally generated k-shuffle Dyck bracket sequences (Hu et al. 2025, arXiv:2502.19249), as flat uint16 token-id .bin files. Token ids are 0-based: opening bracket type i is id i and its matching close is i + k, so ids span [0, 2k) and the vocabulary is 2k = 256. Grammar parameters param value k (bracket types) 128 max_depth 16 p_open 0.5 seq_length 2048 file split tokens train.bin train 999,999,488 val.bin val 10,000… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/dyck-k128-seq_len_2048-1B.tabularn<1K0 likes99 downloads5mo agoHugging Face07open-llm-leaderboard /CultriX__SeQwence-14B-detailsgated Dataset Card for Evaluation run of CultriX/SeQwence-14B Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-details.tabular10K<n<100K0 likes81 downloads2y agoHugging Face08Jules-OC /flowzap-sequence-workflows sequence-workflows A synchronized FlowZap template corpus with 242 canonical templates sourced from https://flowzap.xyz/sitemap-templates.xml and organized by primary Use Case. Organization Model Top-level folders are primary Use Cases from the FlowZap Templates dropdown. Second-level folders preserve the original source domain from the FlowZap app index. Each template keeps all matched Use Cases in metadata.json and the generated JSON/CSV indexes. Templates that do not… See the full description on the dataset page: https://huggingface.co/datasets/Jules-OC/flowzap-sequence-workflows.tabularn<1K0 likes73 downloads7mo agoHugging Face09hi-todayis-jh /grpo-qwen3-1.7b-taco-easy-3200-bs32-n8-seqs16-32k-146102-rollouts grpo_Qwen3-1.7B_TACO-easy-3200_bs32_n8_seqs16_32k_1epoch rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular10K<n<100K0 likes73 downloads12d agoHugging Face10open-llm-leaderboard /sequelbox__Llama3.1-8B-PlumCode-detailsgated Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumCode Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumCode The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumCode-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face11Unggi /modernbert_encoder_sp_seq_512_csedm_fold1tabularn<1K0 likes49 downloads2y agoHugging Face12open-llm-leaderboard /CultriX__SeQwence-14B-v5-detailsgated Dataset Card for Evaluation run of CultriX/SeQwence-14B-v5 Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B-v5 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-v5-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Face13junha1125 /openvid-frame-sequences-1M OpenVid Frame Sequences — 1M adjacent frame pairs Short, single-shot frame sequences cut from OpenVid-1M, built to train and evaluate models on what changes between two frames half a second apart. One sample = 10 consecutive frames, 0.5 s apart (a 4.5 s span) → 9 adjacent frame pairs. [f00] --0.5s--> [f01] --0.5s--> [f02] ... [f09] ^ the thing you describe / predict Sequences 116,596 Frames per sequence 10 (0.5 s apart, t = 0.0 … 4.5 s) Adjacent frame… See the full description on the dataset page: https://huggingface.co/datasets/junha1125/openvid-frame-sequences-1M.tabularimage-to-text100K<n<1M0 likes46 downloads2mo agoHugging Face14open-llm-leaderboard /CultriX__SeQwence-14Bv1-detailsgated Dataset Card for Evaluation run of CultriX/SeQwence-14Bv1 Dataset automatically created during the evaluation run of model CultriX/SeQwence-14Bv1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14Bv1-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face15alexkstern /nca-paper-seq_len_1024-164M nca-paper-seq_len_1024-164M Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002. Configuration param value grid 12×12 colors… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-seq_len_1024-164M.tabularn<1K0 likes41 downloads5mo agoHugging Face16Unggi /modernbert_encoder_sp_seq_512_dbe22kt_fold1tabularn<1K0 likes38 downloads2y agoHugging Face17CodeMasterCody3D /qwen35-08b-seqlen-ablation-0919tabularn<1K0 likes37 downloads21d agoHugging Face18open-llm-leaderboard /sequelbox__gemma-2-9B-MOTH-detailsgated Dataset Card for Evaluation run of sequelbox/gemma-2-9B-MOTH Dataset automatically created during the evaluation run of model sequelbox/gemma-2-9B-MOTH The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__gemma-2-9B-MOTH-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face19open-llm-leaderboard /sequelbox__Llama3.1-8B-MOTH-detailsgated Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-MOTH Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-MOTH The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-MOTH-details.tabular10K<n<100K0 likes24 downloads2y agoHugging Face20alexkstern /nca-paper-share10-seq_len_1024-164M nca-paper-share10-seq_len_1024-164M Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002. Configuration param value grid 12×12… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share10-seq_len_1024-164M.tabularn<1K0 likes23 downloads4mo agoHugging Face21alexkstern /nca-paper-share10-seq_len_1024-164M-seed2 nca-paper-share10-seq_len_1024-164M-seed2 Procedurally generated Neural Cellular Automata trajectories (Lee et al. 2026), as flat uint16 token-id .bin files. Random NCA rules are rolled out on a 12×12 grid of 10 cell states and tokenized by 2×2 patches (base-10); only high-complexity rules survive a gzip-ratio filter (kept iff in (0.5, 1.0)). Token ids: 10,000 patch ids plus two grid delimiters (start=10000, end=10001); vocab = 10,002. Configuration param value grid… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/nca-paper-share10-seq_len_1024-164M-seed2.tabularn<1K0 likes23 downloads3mo agoHugging Face22LLMTeamAkiyama /cleand_sequelbox_Celestia3-DeepSeek-R1-0528元データ: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528 データ件数: 88,443 平均トークン数: 2143 最大トークン数: 31,680 合計トークン数: 189,577,005 ファイル形式: JSONL ファイルサイズ: 812.4 MB tabularquestion-answering10K<n<100K0 likes22 downloads1y agoHugging Face23joduor /adaption-nist-biosafety-seq-bench This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-nist_biosafety_seq_bench This dataset comprises structured biological sequence records designed as a strategic benchmark for training AI systems in biosafety, biosecurity, and synthetic biology governance. Each sample includes genomic data, organism identifiers, risk classification labels, and review status metadata to support pathogen detection and function prediction tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-nist-biosafety-seq-bench.tabular10K<n<100K0 likes21 downloads4mo agoHugging Face24open-llm-leaderboard /sequelbox__Llama3.1-8B-PlumChat-detailsgated Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumChat Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumChat The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumChat-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face25open-llm-leaderboard /CultriX__SeQwence-14B-EvolMerge-detailsgated Dataset Card for Evaluation run of CultriX/SeQwence-14B-EvolMerge Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B-EvolMerge The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-EvolMerge-details.tabular10K<n<100K0 likes16 downloads2y agoHugging Face26alexkstern /nca-paper-share200-seq_len_2048-6.5Btabularn<1K0 likes16 downloads3mo agoHugging Face27open-llm-leaderboard /sequelbox__Llama3.1-8B-PlumMath-detailsgated Dataset Card for Evaluation run of sequelbox/Llama3.1-8B-PlumMath Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-8B-PlumMath The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-8B-PlumMath-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face28joduor /adaption-ebolavirus-protein-sequences This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ebolavirus_protein_sequences This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-ebolavirus-protein-sequences.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face29joduor /ebolavirus_protein_sequences_INITIAL This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-ebolavirus_protein_sequences This dataset contains amino acid sequences for seven key proteins from various Ebola and Marburg virus genomes, including strains like Zaire, Sudan, and Tai Forest. Each entry provides the protein identifier, name, strain information, and the full sequence intended for generating embeddings using models like ESM-2 or ProtT5. The collection includes major… See the full description on the dataset page: https://huggingface.co/datasets/joduor/ebolavirus_protein_sequences_INITIAL.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face30open-llm-leaderboard /sequelbox__Llama3.1-70B-PlumChat-detailsgated Dataset Card for Evaluation run of sequelbox/Llama3.1-70B-PlumChat Dataset automatically created during the evaluation run of model sequelbox/Llama3.1-70B-PlumChat The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sequelbox__Llama3.1-70B-PlumChat-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.