Team Ai
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset 0 likes269 downloads6mo agoHugging Face02gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk16-data GSA-PT-Qwen2-7B-Instruct-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset tabular10K<n<100K0 likes238 downloads6mo agoHugging Face03gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk32-data GSA-PT-Qwen2-7B-Instruct-chunk32-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset 0 likes227 downloads6mo agoHugging Face04gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk8-data GSA-PT-Llama-3.2-1B-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset 0 likes217 downloads6mo agoHugging Face05gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk32-data GSA-FT-Qwen2-7B-Instruct-chunk32-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset tabular10K<n<100K0 likes163 downloads6mo agoHugging Face06GwendalTsang /repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle Reproduction bundle — Incremental Learning of Sparse Attention Patterns in Transformers ICML 2026 paper #12503 · OpenReview vSRh1qU5sH · arXiv 2602.19143 Everything needed to re-run the reproduction whose results are recorded in the Trackio logbook GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers. Layout upstream/ the authors' official code, vendored unchanged… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle.0 likes153 downloads3mo agoHugging Face07gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk16-data GSA-PT-Llama-3.2-1B-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset 0 likes141 downloads6mo agoHugging Face08gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset tabular10K<n<100K0 likes134 downloads6mo agoHugging Face09gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-data GSA-PT-Qwen2-7B-Instruct-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset tabular10K<n<100K0 likes130 downloads6mo agoHugging Face10gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk16-data GSA-FT-Qwen2-7B-Instruct-chunk16-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset 0 likes111 downloads6mo agoHugging Face11gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset tabular10K<n<100K0 likes93 downloads6mo agoHugging Face12gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset 0 likes88 downloads6mo agoHugging Face13gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk4-chunk4-data GSA-PT-Llama-3.2-1B-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk4-chunk4 — model trained on this dataset 0 likes51 downloads6mo agoHugging Face14gist-sparse-attention /GSA-FT-Llama-3.2-1B-data GSA-FT-Llama-3.2-1B-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Llama-3.2-1B. The answers to the questiosn are generated by Deepseek-V3.2 Each sample is formatted with GSA gist tokens for instruction fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-FT-Llama-3.2-1B-chunk8 yuzhenm/GSA-FT-Llama-3.2-1B-chunk16… See the full description on the dataset page: https://huggingface.co/datasets/gist-sparse-attention/GSA-FT-Llama-3.2-1B-data.0 likes40 downloads6mo agoHugging Face15gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk8-data GSA-FT-Qwen2-7B-Instruct-chunk8-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset 0 likes23 downloads6mo agoHugging Face16shiqihe /WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3Bgated Sparse Attention for Web Agents — Offline Replay Records Step-level records from an offline replay study comparing three sparse-attention methods against full attention on browser-agent trajectories. Model: Qwen3-VL-30B-A3B-Instruct (48 layers, 128 experts / 8 active, GQA 32:4, head_dim 128) Hardware: NVIDIA GB10 (sm_121) Tasks: 50 complete trajectories sampled from WebVoyager + GAIA — 332 steps, of which 190 emit an element index. Every task was completed successfully by the… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B.textother1K<n<10K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.