OpenMOSS-Team/MOSS-VL-Instruct-0708-FP8
per-sample media budgets in processor; per-sample MRoPE position ids for batched vision inputs
per-sample media budgets in processor; per-sample MRoPE position ids for batched vision inputs
sync offline system prompts with training data: no_thinking -> no system message, deep_thinking -> <think>/<answer> tags
Align tokenizer with the standard (add silence/response tokens, embed chat_template)
fix: apply query RoPE in cross-attention when reusing cached vision KV during decode
Align processor with MOSS-VL-Instruct-0708 (Fast image/video processor)
Upload README.md with huggingface_hub
Fix language switch links
Separate English and Chinese benchmark model cards
Add bilingual quantization benchmark comparison
Restore model card assets and logo
Add files using upload-large-folder tool
initial commit
