datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
checkpoint
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/YYF111/checkpoint.officeqa-checkpoint-eval-data
Checkpoint evaluation plot data
Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures.
No model execution, grading, publication, or source-result changes were performed to make this export.
Contents
checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds.
pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.finetuning-checkpointsKrea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.weact-native-science-checkpoint
Native WeAct science checkpoint
The fixed eight-question, seven-arm pilot has reached its human-review checkpoint. All 56 planned task records are in pilot_v4/. The frozen 3,000-question test has not been run. No accuracy score or retraining conclusion is claimed.
The independent Serper, Jina and E2B checks passed, and each backend passed native webpage extraction with the live auxiliary model. The corrected runtime uses original questions, native hard/soft routing, the… See the full description on the dataset page: https://huggingface.co/datasets/Corning/weact-native-science-checkpoint.rhan-checkpointsphysam4d-dynamics-checkpointsdetails_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private
Dataset Card for Evaluation run of hosted_vllm//fsx/anton/deepseek-r1-checkpoint
Dataset automatically created during the evaluation run of model hosted_vllm//fsx/anton/deepseek-r1-checkpoint.
The dataset is composed of 15 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private.GHN_checkpoints_CE_vs_KDwebvid_10m_part_2-bootstapir_checkpoint_v2AlphaPanda_checkpointsbridge_data_v2-bootstapir_checkpoint_v2fractal-bootstapir_checkpoint_v2bc_z-bootstapir_checkpoint_v2gradlab-breakout-checkpoint-evals
Checkpoint monitoring trajectories
Each complete episode row links to immutable transition tables and lossless PNG
chunks. The episode index retains Run, training seed, Checkpoint, evaluation,
episode conditions, actions, rewards, boundaries and source/runtime provenance.
Frames are full unmasked native RGB at the contracted action cadence. Chunk joins
preserve the initial and true terminal image without duplicated transitions.
Staged chunks without a complete episode index are… See the full description on the dataset page: https://huggingface.co/datasets/tsilva/gradlab-breakout-checkpoint-evals.reinforcement-learning-checkpoint-downloadstranslation-checkpointsmess3-sfp-checkpointsvisual-question-answering-checkpoint-downloadsgenerations-olmo-3-1025-7b-simnpo-gentle-checkpoint-192rec_rl_checkpoints
rec_rl_checkpoints
Checkpoints for Semantic-ID reasoning generative recommendation, trained with
the SIDReasoner three-stage recipe:
Stage-1 SFT — direct Semantic-ID (SID) prediction, no reasoning.
Stage-2 reasoning activation — cold-start "think" format activation.
Stage-3 GRPO RL — reinforcement learning over SID reasoning traces.
Base model: Qwen3-1.7B. Domains: Office_Products, Video_Games,
Industrial_and_Scientific (Amazon Reviews).
Layout
reproduced/… See the full description on the dataset page: https://huggingface.co/datasets/yufan/rec_rl_checkpoints.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.cameo_checkpointsgenerations-21-DEBUG-qwen3-8b-simnpo-gentle-igm-10b-target-100-localtrain-checkpoint-1unconditional-image-generation-checkpoint-downloadsgenerations-17-DEBUG-qwen3-8b-simnpo-gentle-baseline-target-100-localtrain-checkpoint-1generations-18-DEBUG-llama-3_1-8b-simnpo-gentle-bm25-10b-target-100-localtrain-checkpoint-1maniskill-bootstapir_checkpoint_v2agentboard-babyai-v1-v071-four-run-checkpoints
AgentBoard BabyAI v1 Prime-RL Four-Run Checkpoints
This dataset preserves four runs from the v071 experiment series using
Qwen/Qwen3.5-9B, Prime-RL 0.7.0 at commit
cb74ae2bebe9710b11551970a9661d10116c7179, and the packaged
agentboard-babyai-v1-context 0.2.2 Verifiers environment.
Runs
Directory
Algorithm
ECHO user alpha
Exact resume step
Final trainer step
Retained adapters
rlonly/
GRPO
0
95
100
21
echo005/
ECHO
0.05
100
100
23
echo050/
ECHO
0.5
95… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-four-run-checkpoints.feature-extraction-checkpoint-downloads
