Team Ai
Modelpublic

LS-ML/Ornith-1.0-35B-ModelOpt-NVFP4-Expert

sourceHugging Facemitupdated 4mo agoView on Hugging Face
2likes94downloads
Model Card

Ornith 1.0 35B ModelOpt NVFP4 Expert

This repository contains a community ModelOpt NVFP4 experts-only quantization of `deepreinforce-ai/Ornith-1.0-35B`.

This is not an official DeepReinforce release. The source BF16 checkpoint is unchanged and remains available from the upstream repository.

Model Details

FieldValue
Base modeldeepreinforce-ai/Ornith-1.0-35B
Base revision5df2ed3f675c7beaa490328cc70bb573b65fb660
Release repoLS-ML/Ornith-1.0-35B-ModelOpt-NVFP4-Expert
ArchitectureQwen3_5MoeForConditionalGeneration
Model typeqwen3_5_moe
ModalityText and vision-language
Max context tested262144 tokens
QuantizationModelOpt NVFP4, experts-only target
Group size16
Exported checkpoint size23G on disk
Source BF16 checkpoint size66G on disk
LicenseMIT

The repository name uses Expert, but the quantization target is experts-only: the MoE expert linear weights are quantized while embeddings, lm head, visual modules, attention, linear-attention modules, and shared experts are excluded. See hf_quant_config.json and config.json for the exact ModelOpt quantization metadata.

Benchmark Summary

The checkpoint was benchmarked against the upstream BF16 model on the same DGX Spark / GB10, same vLLM nightly container, same 262K serving profile, same FP8 KV cache allocation, and same benchmark harness.

Performance summary:

MetricBF16ModelOpt NVFP4Result
Model directory size66G23G2.9x smaller
vLLM model loading memory65.53 GiB22.41 GiB2.9x lower
Weight loading time454.55s153.13s3.0x faster
Text prefill speedup rangebaseline1.08x to 1.41xfaster in every row
Text decode speedup rangebaseline1.16x to 1.45xfaster in every row
4K image prefill1204.6 tok/s1357.6 tok/s1.13x faster
4K image decode30.0 tok/s36.1 tok/s1.20x faster

Accuracy summary:

BenchmarkBF16 acc / scoreNVFP4 acc / scoreDelta NVFP4-BF16
MMLU-Pro55.8%55.0%-0.8 pp
GPQA Diamond mirror22.7%13.6%-9.1 pp
MATH-50034.6%36.6%+2.0 pp
HumanEval+39.0%21.3%-17.7 pp
MBPP+78.0%76.5%-1.6 pp
MMMU validation57.4%56.1%-1.3 pp
OCRBench69.9%70.2%+0.3 pp
BFCL v3 10-each subset1.0%0.0%-1.0 pp
ToolEvalBench8288+6

Detailed performance and accuracy benchmark tables are included below.

Benchmark Graphics

Performance benchmark:

[image]

Accuracy and quality benchmark:

[image]

Quantization

The checkpoint was produced with NVIDIA ModelOpt from a fused-expert staging copy of the upstream BF16 model. The upstream checkpoint stores split expert tensors; the staging copy fused expert gate/up/down tensors into the layout expected by current Transformers and ModelOpt. The upstream BF16 checkpoint itself was not modified.

Quantization metadata:

FieldValue
ModelOpt source version0.46.0.dev106+g6cc522658
ModelOpt commit6cc5226588f0668679df03ba4646b7dfec32f99c
Quant algorithmNVFP4
KV cache quantization during exportnone
Calibration datasetnvidia/Nemotron-SFT-Agentic-v2, split search
Calibration settingscalib_size=16, calib_seq=512, batch_size=1
Export modelow-memory ModelOpt path

The quantization notes used during export are included in QUANTIZATION_NOTES.md.

Serving With vLLM

Validated serving stack:

  • —vllm/vllm-openai:nightly
  • —ModelOpt quantization loader: --quantization modelopt
  • —FlashInfer attention backend
  • —Blackwell / GB10 tested with CUTE_DSL_ARCH=sm_121a
  • —OpenAI-compatible chat completions
  • —Text, image, and Qwen3 XML tool-call smoke tests

Example full-context DGX Spark / GB10 launch:

bash
docker run --rm \
  --name ornith35-nvfp4-vllm \
  --gpus all --ipc=host --ulimit memlock=-1 --ulimit stack=67108864 --shm-size=32g \
  -p 8000:8000 \
  -v /path/to/models:/models \
  -v ~/.cache:/root/.cache \
  -e FLASHINFER_DISABLE_VERSION_CHECK=1 \
  -e CUTE_DSL_ARCH=sm_121a \
  vllm/vllm-openai:nightly \
  --model /models/Ornith-1.0-35B-ModelOpt-NVFP4-Expert \
  --host 0.0.0.0 \
  --port 8000 \
  --served-model-name ornith35-nvfp4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --quantization modelopt \
  --kv-cache-dtype fp8 \
  --kv-cache-memory-bytes 28G \
  --attention-backend flashinfer \
  --max-model-len 262144 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 8192 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --enable-chunked-prefill

Lower-memory smoke profile:

bash
docker run --rm \
  --name ornith35-nvfp4-vllm-smoke \
  --gpus all --ipc=host --ulimit memlock=-1 --ulimit stack=67108864 --shm-size=32g \
  -p 8000:8000 \
  -v /path/to/models:/models \
  -v ~/.cache:/root/.cache \
  -e FLASHINFER_DISABLE_VERSION_CHECK=1 \
  -e CUTE_DSL_ARCH=sm_121a \
  vllm/vllm-openai:nightly \
  --model /models/Ornith-1.0-35B-ModelOpt-NVFP4-Expert \
  --host 0.0.0.0 \
  --port 8000 \
  --served-model-name ornith35-nvfp4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --quantization modelopt \
  --kv-cache-dtype fp8 \
  --kv-cache-memory-bytes 4G \
  --attention-backend flashinfer \
  --max-model-len 8192 \
  --max-num-seqs 1 \
  --max-num-batched-tokens 8192 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --enable-chunked-prefill

For non-thinking deterministic chat/eval requests, use:

json
{
  "temperature": 0,
  "chat_template_kwargs": {
    "enable_thinking": false
  }
}

Validation

The quantized checkpoint was validated on a local DGX Spark / GB10 against the BF16 upstream checkpoint using the same vLLM runtime profile, same maximum context, same FP8 KV cache allocation, and the same benchmark harness.

Performance Benchmark Results

All text rows below used unique prompts, prefix caching disabled, max_tokens=128, and submitted concurrency 1, 2, or 4.

ContextConcBF16 prefillNVFP4 prefillPrefill speedupBF16 decodeNVFP4 decodeDecode speedupBF16 TTFTNVFP4 TTFT
32k13288.84626.21.41x29.635.51.20x10.0s7.1s
32k24349.75412.51.24x55.470.21.27x15.1s12.1s
32k44425.15512.51.25x82.1118.91.45x29.6s23.8s
64k13656.14348.01.19x28.233.51.19x17.9s15.1s
64k23625.84315.61.19x48.563.61.31x36.1s30.4s
64k43594.44289.91.19x79.7103.31.30x72.9s61.1s
128k12653.13003.21.13x26.130.71.18x49.4s43.6s
128k22632.22988.21.14x45.055.51.23x99.6s87.7s
128k42607.52969.31.14x64.084.41.32x201.1s176.6s
full11713.51853.71.08x22.225.71.16x152.8s141.3s
full21704.91838.81.08x37.543.81.17x307.2s284.8s
full41702.51840.81.08x56.771.01.25x615.3s569.0s

Throughput units are tokens per second. TTFT is max time to first token for the concurrent batch.

The serial 4096 x 4096 image test:

CaseBF16 prefillNVFP4 prefillPrefill speedupBF16 decodeNVFP4 decodeDecode speedupBF16 TTFTNVFP4 TTFT
4K image1204.61357.61.13x30.036.11.20x13.6s12.1s

Startup and memory observations:

MetricBF16ModelOpt NVFP4
Remote model directory size66G23G
vLLM model loading memory65.53 GiB22.41 GiB
Weight loading time454.55s153.13s
Engine init observed to healthabout 610sabout 290s
MemAvailable after healthabout 17 GiBabout 59 GiB

Accuracy Benchmark Results

Quality runs used deterministic decoding with temperature=0, chat_template_kwargs={"enable_thinking": false}, and paired per-item scoring. The key quantization regression count is "BF16 correct / NVFP4 wrong".

BenchmarkItemsBF16 accNVFP4 accDelta NVFP4-BF16BF16 correct/NVFP4 wrongNVFP4 correct/BF16 wrong
MMLU-Pro1203255.8%55.0%-0.8 pp527429
GPQA Diamond mirror19822.7%13.6%-9.1 pp235
MATH-50050034.6%36.6%+2.0 pp1727
HumanEval+16439.0%21.3%-17.7 pp334
MBPP+37878.0%76.5%-1.6 pp137
MMMU validation90057.4%56.1%-1.3 pp5442
OCRBench100069.9%70.2%+0.3 pp1720
BFCL v3 10-each subset1001.0%0.0%-1.0 pp10

ToolEvalBench version 2.0.7 was run sequentially over 69 standard scenarios:

ModelFinal scorePointsDeployabilityResponsivenessSafety warnings
BF1682113 / 13874543
NVFP488121 / 13881641

Interpretation: this NVFP4 export is close to BF16 on broad multiple-choice, multimodal, OCR, and MBPP-style code tasks, but it showed clear regressions on GPQA Diamond and HumanEval+ in this run.

Caveats

  • —This is a community quantization, not an official upstream release.
  • —The checkpoint has been validated with vLLM ModelOpt loading. Other loaders may not support this ModelOpt NVFP4 format.
  • —vLLM marks ModelOpt NVFP4 support as experimental, so revalidate after major vLLM, ModelOpt, CUDA, or FlashInfer changes.
  • —The export does not include FP8 KV q/prob scaling factors. When serving with --kv-cache-dtype fp8, vLLM reports that it uses scale 1.0; treat this as a quality caveat for accuracy-sensitive workloads.
  • —GPQA used the ungated fingertap/GPQA-Diamond mirror because the official dataset was gated at the time of evaluation.
  • —HumanEval+ and MBPP+ used EvalPlus prompts and expanded inputs, but scoring used a local subprocess checker because the official EvalPlus sandbox failed on the local macOS host with resource-limit errors.
  • —BFCL v3 10-each should be treated as raw paired signal only; ToolEvalBench is the stronger tool-use benchmark in this report.

Responsible Use

Use this model consistently with the upstream model license and any applicable laws or platform policies. Because this is a quantized derivative, evaluate it for your own target domain before relying on it in production or accuracy-sensitive workflows.

License And Attribution

The upstream model is MIT licensed. This quantized release preserves the MIT license and attribution to deepreinforce-ai/Ornith-1.0-35B.

Upstream model: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B

Quantized by LS-ML.