Team Ai
Modelpublic

Shoolife/Qwen3-1.7B-TensorRT-LLM-Checkpoint-FP8

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes21downloads
Model Card

Qwen3-1.7B TensorRT-LLM Checkpoint (FP8)

This repository contains a community-converted TensorRT-LLM checkpoint for `Qwen/Qwen3-1.7B`.

It is a TensorRT-LLM checkpoint-format repository, not a prebuilt engine. The intent is to let you download the checkpoint from Hugging Face and build an engine locally for your own GPU and TensorRT-LLM version.

Who This Repo Is For

This repository is for users who already work with TensorRT-LLM and want a ready-made TensorRT-LLM checkpoint that they can turn into a local engine for their own GPU.

It is not:

  • —a prebuilt TensorRT engine
  • —a plain Transformers checkpoint
  • —an Ollama model
  • —a one-click chat model that can be run directly after download

How to Use

  1. 1.Download this repository from Hugging Face.
  2. 2.Build a local engine with trtllm-build for your own GPU and TensorRT-LLM version.
  3. 3.Run inference with the engine you built.

The Build Example section below shows the validated local command used for the benchmark snapshot in this README.

Model Characteristics

  • —Base model: Qwen/Qwen3-1.7B
  • —License: apache-2.0
  • —Architecture: Qwen3ForCausalLM
  • —Upstream maximum context length (max_position_embeddings): 40960
  • —Hidden size: 2048
  • —Intermediate size: 6144
  • —Layers: 28
  • —Attention heads: 16
  • —KV heads: 8
  • —Vocabulary size: 151936

These values come from the upstream model/checkpoint configuration. They describe the model family itself, not a specific locally built TensorRT engine.

Checkpoint Details

  • —TensorRT-LLM version used for conversion: 1.2.0rc6
  • —Checkpoint dtype: bfloat16
  • —Quantization: FP8 (weights and activations)
  • —KV cache quantization: FP8
  • —Calibration dataset: cnn_dailymail (64 samples, max seq length 256)
  • —Tensor parallel size: 1
  • —Checkpoint files:
  • —config.json
  • —rank0.safetensors
  • —tokenizer and generation files copied from the upstream Hugging Face model

Files

  • —config.json: TensorRT-LLM checkpoint config
  • —rank0.safetensors: TensorRT-LLM checkpoint weights
  • —generation_config.json: upstream generation config
  • —tokenizer.json: upstream tokenizer
  • —tokenizer_config.json: upstream tokenizer config
  • —merges.txt: upstream merges file
  • —vocab.json: upstream vocabulary

Build Example

The following command is the validated local engine build used for the benchmarks in this README. These values are build-time/runtime settings for one local engine, not limits of the checkpoint itself.

Build an engine locally with TensorRT-LLM:

bash
huggingface-cli download Shoolife/Qwen3-1.7B-TensorRT-LLM-Checkpoint-FP8 --local-dir ./checkpoint

trtllm-build \
  --checkpoint_dir ./checkpoint \
  --output_dir ./engine \
  --gemm_plugin auto \
  --gpt_attention_plugin auto \
  --max_batch_size 1 \
  --max_input_len 512 \
  --max_seq_len 1024 \
  --max_num_tokens 256 \
  --workers 1 \
  --monitor_memory

If you rebuild the engine with different limits, memory usage and supported request shapes will change accordingly.

Quantization

This checkpoint was produced using TensorRT-LLM quantization tooling:

bash
python quantize.py \
  --model_dir ./Qwen3-1.7B \
  --output_dir ./checkpoint_fp8 \
  --dtype bfloat16 \
  --qformat fp8 \
  --kv_cache_dtype fp8 \
  --calib_dataset cnn_dailymail \
  --calib_size 64 \
  --batch_size 1 \
  --calib_max_seq_length 256 \
  --tokenizer_max_seq_length 2048

Validation

The checkpoint was validated by building a local engine and running inference on:

  • —GPU: NVIDIA GeForce RTX 5070 Laptop GPU
  • —Runtime: TensorRT-LLM 1.2.0rc6

Validated Local Engine Characteristics

Local build and runtime characteristics from the validated engine used for the benchmark snapshot below:

PropertyValue
Checkpoint size2.5 GB
Built engine size2.6 GB
Tested GPUNVIDIA GeForce RTX 5070 Laptop GPU
GPU memory reported by benchmark host7.53 GiB
Engine build max_batch_size1
Engine build max_input_len512
Engine build max_seq_len1024
Engine build max_num_tokens256

Important: the 1024 / 256 limits above belong only to this particular local engine build. They are not the intrinsic maximum context or generation limits of Qwen3-1.7B itself.

These values are specific to the local engine build used for validation and will change if you rebuild with different TensorRT-LLM settings and memory budgets.

Benchmark Snapshot

Local single-GPU measurements from the validated local engine on RTX 5070 Laptop GPU, using TensorRT-LLM synthetic fixed-length requests, 20 requests per profile, 2 warmup requests, and concurrency=1.

ProfileInputOutputTTFTTPOTOutput tok/sAvg latency
tiny_16_3216329.51 ms7.26 ms136.5234.4 ms
short_chat_42_64426410.79 ms7.55 ms131.6486.3 ms
balanced_128_12812812814.53 ms7.51 ms132.2968.0 ms
long_prompt_192_641926418.97 ms7.56 ms129.3495.1 ms
long_generation_42_1924219210.16 ms7.20 ms138.71384.7 ms

These numbers are local measurements from one machine and should be treated as reference values, not portability guarantees.

Quick Parity Check

A quick parity check was run on ARC-Challenge (20 examples) and OpenBookQA (20 examples) to verify that the TensorRT-LLM FP8 engine produces comparable answers to the upstream Hugging Face model.

BenchmarkHF AccuracyTRT AccuracyAgreement
arc_challenge0.750.600.80
openbookqa0.600.750.70
Overall`0.675``0.675``0.75`

The FP8 quantized engine achieves the same overall accuracy as the HF baseline, with some per-example variance.

Local Comparison

The table below compares locally validated TensorRT-LLM variants built for the same GPU family and the same local engine limits (max_batch_size=1, max_seq_len=1024, max_num_tokens=256).

VariantCheckpointEngine`short_chat_42_64``balanced_128_128``long_generation_42_192`Quick-check overallQuick-check change vs BF16Practical reading
BF163.9 GB3.9 GB98.5 tok/s97.7 tok/s98.4 tok/s0.675baselineNative precision, best numerical stability
FP163.9 GB3.9 GB98.5 tok/s97.7 tok/s98.4 tok/s0.675sameEquivalent precision, identical results
FP82.5 GB2.6 GB131.6 tok/s132.2 tok/s138.7 tok/s0.675same~35% faster, same accuracy on this subset
NVFP41.9 GB1.4 GB176.5 tok/s176.8 tok/s177.2 tok/s0.25-42.5 pts on this quick-checkFastest and smallest, but severe quality drop

This comparison is intentionally local and narrow. It should not be treated as a universal benchmark across all prompts, datasets, GPUs, or TensorRT-LLM versions.

Notes

  • —This is not an official Qwen or NVIDIA release.
  • —This repository does not include a prebuilt TensorRT engine.
  • —Engine compatibility and performance depend on your GPU, driver, CUDA, TensorRT, and TensorRT-LLM versions.