fp8
Datasets
All datasets matching “fp8”qwen3-1.7b-polaris-fp8-rollouts-20260912
Qwen3-1.7B Polaris FP8 rollouts
Snapshot of three FP8 rollout datasets taken on 2026-09-12. The model is Qwen3-1.7B-Base. Original Parquet files and attempt, file, and checkpoint-lineage metadata are preserved without rewriting.
Configuration
Training batches
Training responses
Validation responses
Total bytes
maxrl_strict
101
827392
144320
6654288117
maxrl_permissive
107
876544
144320
8168064932
dppo
126
1032192
173184
6733682534
Provenance and… See the full description on the dataset page: https://huggingface.co/datasets/steviel/qwen3-1.7b-polaris-fp8-rollouts-20260912.magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedSIGNAL-Dataset-Hiddens-Qwen-Qwen3-4B-Instruct-FP8This dataset contains hidden states of Qwen3-4B-Instruct model generated using SIGNAL Dataset.
Sentence tokenization
from transformers import AutoTokenizer
from datasets import load_dataset
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507")
# TBD
dbrx-instruct-fp8Eikos-27B-FP8-AgentRewardBench
Eikos-27B-FP8 as an AgentRewardBench judge
Eikos-27B-FP8 is an open-weight typed-decision model (MIT) that
answers typed questions in one forward pass, with a probability for every option. This repository holds its judgments
on AgentRewardBench (Lù et al., 2025), in the
official judgment format, and the official scorer's output.
Results (test split, 1,106 trajectories, official scripts/score_judgments.py)
Overall
AssistantBench
VisualWebArena
WebArena… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/Eikos-27B-FP8-AgentRewardBench.commitmoe-qwen35-fp8-expert-routing-traces
