datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma-4-e2b-atlas
gemma-4-e2b-it-atlasgemma-4-E2B-it-pi-mono-agent-evalgemma4-bestckpt-traces-topk128-v2-e2b-mediumgemma-4-E2B-it-pi-mono-agent-loss-fpgemma-4-e2b-it-SAE
gemma-4-e2b-it — 35-layer SAE atlas
Sparse autoencoders on every decoder layer of gemma-4-e2b-it. Trained from scratch in one rolling pipeline with an event-aware controller. 35 layers, 49,152 features per layer, no per-layer hand-tuning.
The base model is a stubborn one. 15 sliding-window layers, then BAM no KV cache, thick and getting thicker the deeper you go. This atlas was built the whole way through it anyway.
What this is
Three months of work. My first… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-it-SAE.financial-english-source-corpus-gemma4-e2b-1280gemma4-e2b-base-topk128-hf-overlay-v128-seed42
Gemma 4 E2B base top-k-128 HF training overlay
This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into
Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from
JWei05/gemma4-e2b-base-topk128-traces,
but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training
engine.
This repository is a reproducibility artifact for the corresponding distillation run. It is not a
new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.gemma4-e2b-base-topk128-tracesgemma4-bestckpt-traces-topk128-v2-e2b-hardgemma4-bestckpt-traces-topk128-v2-e2b-easygemma-4-E2B-distill-qlora-kaggle
gemma4-qlora-kaggle-notebook
Kaggle-ready notebook: QLoRA SFT of google/gemma-4-E2B-it
on the sft config of r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation.
Built for Kaggle T4 x2 (runs on a single T4 by design), 12 h session limit, internet ON.
Import: Kaggle -> File -> Import Notebook -> upload gemma4_qlora_kaggle.ipynb (or add this dataset as input and open the file).
Sections: 0 config (N_ROWS=4000, MAX_LEN=2048, teacher/task filters, LoRA, smoke/resume/push knobs) - 1… See the full description on the dataset page: https://huggingface.co/datasets/adffedccasfe/gemma-4-E2B-distill-qlora-kaggle.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.open-perfectblend_gemma-4-E2B-itgemma-4-e2b-acoustic-verifier-dataSynthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-E2B-it.omiv-unsloth-gemma4-e2b-it-q8-evidence
OMIV Case Study 01 — Public Evidence Summary
The export succeeded, and both current artifact identities matched retained historical C1 size/SHA-256 observations.
This Dataset repository contains compact, machine-readable public summaries and
SHA-256 commitments for an independent OMIV validation of an Unsloth Q8_0 GGUF
export. It contains no model weights, GGUF bytes, tokenizer files, source
snapshots, private evidence, execution envelopes, transcripts, or infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/200lz/omiv-unsloth-gemma4-e2b-it-q8-evidence.Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-original-gemma-4-E2B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-E2B-it.ConvAI2-Qwen-original-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-E2B-it.Gemma4-E2B-SFT-WebCode
Gemma4-E2B-SFT-WebCode
Synthetic frontend web development dataset. Natural language component description → production-ready code.
Frameworks: React, TypeScript, Tailwind CSS, Vanilla HTML/CSS/JS.
Components: Navigation, forms, modals, data tables, charts, infinite scroll, etc.
Format: ShareGPT/ChatML. Includes accessibility attributes and comments.
Use: Fine-tune models for frontend copilot tasks.
Generator: DuoNeural/TurboGemma4E2B, temperature 0.65.
gemma-4-e2b-SAE-sqlite
Gemma 4 E2B SAE SQLite Atlas
An exact, queryable SQLite representation of all 35 residual-stream
sparse autoencoders from
juiceb0xc0de/gemma-4-e2b-it-SAE.
The database contains 1,720,320 feature rows. Encoder and decoder
vectors preserve the source checkpoints' float32 values exactly.
Files
gemma-4-e2b-sae.sqlite3 — SQLite database (20.63 GiB)
manifest.json — source revision, dimensions, SHA-256, and integrity result
Database SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e2b-SAE-sqlite.ConvAI2-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-E2B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-E2B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-E2B-it.PersonaChat-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E2B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-E2B-it.ConvAI2-Qwen-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-E2B-it.ConvAI2-FLUX-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E2B-it.ConvAI2-ERNIE-enhanced-gemma-4-E2B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E2B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E2B-it.
