RedHatAI/phi-4-quantized.w8a8
54.3k
1---2language:3- en4base_model:5- microsoft/phi-46pipeline_tag: text-generation7tags:8- phi9- phi310- nlp11- math12- code13- chat14- conversational15- neuralmagic16- redhat17- llmcompressor18- quantized19- W8A820- INT821- compressed-tensors22license: mit23license_name: mit24name: RedHatAI/phi-4-quantized.w8a825description: This model was obtained by quantizing activations and weights of phi-4 to INT8 data type.26readme: https://huggingface.co/RedHatAI/phi-4-quantized.w8a8/main/README.md27tasks:28- text-to-text29provider: Microsoft30license_link: https://choosealicense.com/licenses/mit/31validated_on:32 - RHOAI 2.2033 - RHAIIS 3.034 - RHELAI 1.535 - vLLM 0.8.436---37 38<h1 style="display: flex; align-items: center; gap: 10px; margin: 0;">39 phi-4-quantized.w8a840 <img src="https://www.redhat.com/rhdc/managed-files/Catalog-Validated_model_0.png" alt="Model Icon" width="40" style="margin: 0; padding: 0;" />41</h1>42 43<a href="https://www.redhat.com/en/products/ai/validated-models" target="_blank" style="margin: 0; padding: 0;">44<img src="https://www.redhat.com/rhdc/managed-files/Validated_badge-Dark.png" alt="Validated Badge" width="250" style="margin: 0; padding: 0;" />45</a>46 47## Model Overview48- **Model Architecture:** Phi3ForCausalLM49 - **Input:** Text50 - **Output:** Text51- **Model Optimizations:**52 - **Activation quantization:** INT853 - **Weight quantization:** INT854- **Intended Use Cases:** This model is designed to accelerate research on language models, for use as a building block for generative AI powered features. It provides uses for general purpose AI systems and applications (primarily in English) which require:55 1. Memory/compute constrained environments.56 2. Latency bound scenarios.57 3. Reasoning and logic.58- **Out-of-scope:** This model is not specifically designed or evaluated for all downstream purposes, thus:59 1. Developers should consider common limitations of language models as they select use cases, and evaluate and mitigate for accuracy, safety, and fairness before using within a specific downstream use case, particularly for high-risk scenarios.60 2. Developers should be aware of and adhere to applicable laws or regulations (including privacy, trade compliance laws, etc.) that are relevant to their use case, including the model’s focus on English.61 3. Nothing contained in this Model Card should be interpreted as or deemed a restriction or modification to the license the model is released under.62- **Release Date:** 03/03/202563- **Version:** 1.064- **Validated on:** RHOAI 2.20, RHAIIS 3.0, RHELAI 1.565- **Model Developers:** Red Hat (Neural Magic)66 67 68### Model Optimizations69 70This model was obtained by quantizing activations and weights of [phi-4](https://huggingface.co/microsoft/phi-4) to INT8 data type.71This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x).72Weight quantization also reduces disk size requirements by approximately 50%.73 74Only weights and activations of the linear operators within transformers blocks are quantized.75Weights are quantized with a symmetric static per-channel scheme, whereas activations are quantized with a symmetric dynamic per-token scheme.76A combination of the [SmoothQuant](https://arxiv.org/abs/2211.10438) and [GPTQ](https://arxiv.org/abs/2210.17323) algorithms is applied for quantization, as implemented in the [llm-compressor](https://github.com/vllm-project/llm-compressor) library.77 78 79## Deployment80 81This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend, as shown in the example below.82 83```python84from vllm import LLM, SamplingParams85from transformers import AutoTokenizer86 87model_id = "neuralmagic-ent/phi-4-quantized.w8a8"88number_gpus = 189 90sampling_params = SamplingParams(temperature=0.7, top_p=0.8, max_tokens=256)91 92tokenizer = AutoTokenizer.from_pretrained(model_id)93 94messages = [95 {"role": "user", "content": "Give me a short introduction to large language model."},96]97 98prompts = tokenizer.apply_chat_template(messages, tokenize=False)99 100llm = LLM(model=model_id, tensor_parallel_size=number_gpus)101 102outputs = llm.generate(prompts, sampling_params)103 104generated_text = outputs[0].outputs[0].text105print(generated_text)106```107 108vLLM aslo supports OpenAI-compatible serving. See the [documentation](https://docs.vllm.ai/en/latest/) for more details.109 110<details>111 <summary>Deploy on <strong>Red Hat AI Inference Server</strong></summary>112 113```bash114podman run --rm -it --device nvidia.com/gpu=all -p 8000:8000 \115 --ipc=host \116--env "HUGGING_FACE_HUB_TOKEN=$HF_TOKEN" \117--env "HF_HUB_OFFLINE=0" -v ~/.cache/vllm:/home/vllm/.cache \118--name=vllm \119registry.access.redhat.com/rhaiis/rh-vllm-cuda \120vllm serve \121--tensor-parallel-size 8 \122--max-model-len 32768 \123--enforce-eager --model RedHatAI/phi-4-quantized.w8a8124```125See [Red Hat AI Inference Server documentation](https://docs.redhat.com/en/documentation/red_hat_ai_inference_server/) for more details.126</details>127 128<details>129 <summary>Deploy on <strong>Red Hat Enterprise Linux AI</strong></summary>130 131```bash132# Download model from Red Hat Registry via docker133# Note: This downloads the model to ~/.cache/instructlab/models unless --model-dir is specified.134ilab model download --repository docker://registry.redhat.io/rhelai1/phi-4-quantized-w8a8:1.5135```136 137```bash138# Serve model via ilab139ilab model serve --model-path ~/.cache/instructlab/models/phi-4-quantized-w8a8140 141# Chat with model142ilab model chat --model ~/.cache/instructlab/models/phi-4-quantized-w8a8143```144See [Red Hat Enterprise Linux AI documentation](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux_ai/1.4) for more details.145</details>146 147<details>148 <summary>Deploy on <strong>Red Hat Openshift AI</strong></summary>149 150```python151# Setting up vllm server with ServingRuntime152# Save as: vllm-servingruntime.yaml153apiVersion: serving.kserve.io/v1alpha1154kind: ServingRuntime155metadata:156 name: vllm-cuda-runtime # OPTIONAL CHANGE: set a unique name157 annotations:158 openshift.io/display-name: vLLM NVIDIA GPU ServingRuntime for KServe159 opendatahub.io/recommended-accelerators: '["nvidia.com/gpu"]'160 labels:161 opendatahub.io/dashboard: 'true'162spec:163 annotations:164 prometheus.io/port: '8080'165 prometheus.io/path: '/metrics'166 multiModel: false167 supportedModelFormats:168 - autoSelect: true169 name: vLLM170 containers:171 - name: kserve-container172 image: quay.io/modh/vllm:rhoai-2.20-cuda # CHANGE if needed. If AMD: quay.io/modh/vllm:rhoai-2.20-rocm173 command:174 - python175 - -m176 - vllm.entrypoints.openai.api_server177 args:178 - "--port=8080"179 - "--model=/mnt/models"180 - "--served-model-name={{.Name}}"181 env:182 - name: HF_HOME183 value: /tmp/hf_home184 ports:185 - containerPort: 8080186 protocol: TCP187```188 189```python190# Attach model to vllm server. This is an NVIDIA template191# Save as: inferenceservice.yaml192apiVersion: serving.kserve.io/v1beta1193kind: InferenceService194metadata:195 annotations:196 openshift.io/display-name: phi-4-quantized.w8a8 # OPTIONAL CHANGE197 serving.kserve.io/deploymentMode: RawDeployment198 name: phi-4-quantized.w8a8 # specify model name. This value will be used to invoke the model in the payload199 labels:200 opendatahub.io/dashboard: 'true'201spec:202 predictor:203 maxReplicas: 1204 minReplicas: 1205 model:206 modelFormat:207 name: vLLM208 name: ''209 resources:210 limits:211 cpu: '2' # this is model specific212 memory: 8Gi # this is model specific213 nvidia.com/gpu: '1' # this is accelerator specific214 requests: # same comment for this block215 cpu: '1'216 memory: 4Gi217 nvidia.com/gpu: '1'218 runtime: vllm-cuda-runtime # must match the ServingRuntime name above219 storageUri: oci://registry.redhat.io/rhelai1/modelcar-phi-4-quantized-w8a8:1.5220 tolerations:221 - effect: NoSchedule222 key: nvidia.com/gpu223 operator: Exists224```225 226```bash227# make sure first to be in the project where you want to deploy the model228# oc project <project-name>229# apply both resources to run model230# Apply the ServingRuntime231oc apply -f vllm-servingruntime.yaml232# Apply the InferenceService233oc apply -f qwen-inferenceservice.yaml234```235 236```python237# Replace <inference-service-name> and <cluster-ingress-domain> below:238# - Run `oc get inferenceservice` to find your URL if unsure.239# Call the server using curl:240curl https://<inference-service-name>-predictor-default.<domain>/v1/chat/completions241 -H "Content-Type: application/json" \242 -d '{243 "model": "phi-4-quantized.w8a8",244 "stream": true,245 "stream_options": {246 "include_usage": true247 },248 "max_tokens": 1,249 "messages": [250 {251 "role": "user",252 "content": "How can a bee fly when its wings are so small?"253 }254 ]255}'256```257 258See [Red Hat Openshift AI documentation](https://docs.redhat.com/en/documentation/red_hat_openshift_ai/2025) for more details.259</details>260 261 262## Creation263 264<details>265 <summary>Creation details</summary>266 This model was created with [llm-compressor](https://github.com/vllm-project/llm-compressor) by running the code snippet below. 267 268 269 ```python270 from transformers import AutoModelForCausalLM, AutoTokenizer271 from llmcompressor.modifiers.quantization import GPTQModifier272 from llmcompressor.modifiers.smoothquant import SmoothQuantModifier273 from llmcompressor.transformers import oneshot274 from datasets import load_dataset275 276 # Load model277 model_stub = "microsoft/phi-4"278 model_name = model_stub.split("/")[-1]279 280 num_samples = 1024281 max_seq_len = 8192282 283 tokenizer = AutoTokenizer.from_pretrained(model_stub)284 285 model = AutoModelForCausalLM.from_pretrained(286 model_stub,287 device_map="auto",288 torch_dtype="auto",289 )290 291 def preprocess_fn(example):292 return {"text": tokenizer.apply_chat_template(example["messages"], add_generation_prompt=False, tokenize=False)}293 294 ds = load_dataset("neuralmagic/LLM_compression_calibration", split="train")295 ds = ds.map(preprocess_fn)296 297 # Configure the quantization algorithm and scheme298 recipe = [299 SmoothQuantModifier(300 smoothing_strength=0.7,301 mappings=[302 [["re:.*qkv_proj"], "re:.*input_layernorm"],303 [["re:.*gate_up_proj"], "re:.*post_attention_layernorm"],304 ],305 ),306 GPTQModifier(307 ignore=["lm_head"],308 sequential_targets=["Phi3DecoderLayer"],309 dampening_frac=0.01,310 targets="Linear",311 scheme="W8A8",312 ),313 ]314 315 # Apply quantization316 oneshot(317 model=model,318 dataset=ds, 319 recipe=recipe,320 max_seq_length=max_seq_len,321 num_calibration_samples=num_samples,322 )323 324 # Save to disk in compressed-tensors format325 save_path = model_name + "-quantized.w8a8"326 model.save_pretrained(save_path)327 tokenizer.save_pretrained(save_path)328 print(f"Model and tokenizer saved to: {save_path}")329 ```330</details>331 332 333 334## Evaluation335 336The model was evaluated on the OpenLLM leaderboard tasks (version 1) with the [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) and the [vLLM](https://docs.vllm.ai/en/stable/) engine, using the following command:337```338lm_eval \339 --model vllm \340 --model_args pretrained="neuralmagic-ent/phi-4-quantized.w8a8",dtype=auto,gpu_memory_utilization=0.6,max_model_len=4096,enable_chunk_prefill=True,tensor_parallel_size=1 \341 --tasks openllm \342 --batch_size auto343```344 345### Accuracy346 347#### Open LLM Leaderboard evaluation scores348<table>349 <tr>350 <td><strong>Benchmark</strong>351 </td>352 <td><strong>phi-4</strong>353 </td>354 <td><strong>phi-4-quantized.w8a8<br>(this model)</strong>355 </td>356 <td><strong>Recovery</strong>357 </td>358 </tr>359 <tr>360 <td>MMLU (5-shot)361 </td>362 <td>80.30363 </td>364 <td>80.39365 </td>366 <td>100.1%367 </td>368 </tr>369 <tr>370 <td>ARC Challenge (25-shot)371 </td>372 <td>64.42373 </td>374 <td>64.33375 </td>376 <td>99.9%377 </td>378 </tr>379 <tr>380 <td>GSM-8K (5-shot, strict-match)381 </td>382 <td>90.07383 </td>384 <td>90.30385 </td>386 <td>100.3%387 </td>388 </tr>389 <tr>390 <td>Hellaswag (10-shot)391 </td>392 <td>84.37393 </td>394 <td>84.30395 </td>396 <td>99.9%397 </td>398 </tr>399 <tr>400 <td>Winogrande (5-shot)401 </td>402 <td>80.58403 </td>404 <td>79.95405 </td>406 <td>99.2%407 </td>408 </tr>409 <tr>410 <td>TruthfulQA (0-shot, mc2)411 </td>412 <td>59.37413 </td>414 <td>58.82415 </td>416 <td>99.1%417 </td>418 </tr>419 <tr>420 <td><strong>Average</strong>421 </td>422 <td><strong>76.52</strong>423 </td>424 <td><strong>76.35</strong>425 </td>426 <td><strong>99.8%</strong>427 </td>428 </tr>429</table>430 431 