Team Ai
Modelpublic

RedHatAI/phi-4-quantized.w8a8

sourceHugging Facemitupdated 4mo agoView on Hugging Face
5likes4.3kdownloads
README.md431 linesDownload Raw Back to root
1---2language:3- en4base_model:5- microsoft/phi-46pipeline_tag: text-generation7tags:8- phi9- phi310- nlp11- math12- code13- chat14- conversational15- neuralmagic16- redhat17- llmcompressor18- quantized19- W8A820- INT821- compressed-tensors22license: mit23license_name: mit24name: RedHatAI/phi-4-quantized.w8a825description: This model was obtained by quantizing activations and weights of phi-4 to INT8 data type.26readme: https://huggingface.co/RedHatAI/phi-4-quantized.w8a8/main/README.md27tasks:28- text-to-text29provider: Microsoft30license_link: https://choosealicense.com/licenses/mit/31validated_on:32  - RHOAI 2.2033  - RHAIIS 3.034  - RHELAI 1.535  - vLLM 0.8.436---37 38<h1 style="display: flex; align-items: center; gap: 10px; margin: 0;">39  phi-4-quantized.w8a840  <img src="https://www.redhat.com/rhdc/managed-files/Catalog-Validated_model_0.png" alt="Model Icon" width="40" style="margin: 0; padding: 0;" />41</h1>42  43<a href="https://www.redhat.com/en/products/ai/validated-models" target="_blank" style="margin: 0; padding: 0;">44<img src="https://www.redhat.com/rhdc/managed-files/Validated_badge-Dark.png" alt="Validated Badge" width="250" style="margin: 0; padding: 0;" />45</a>46 47## Model Overview48- **Model Architecture:** Phi3ForCausalLM49  - **Input:** Text50  - **Output:** Text51- **Model Optimizations:**52  - **Activation quantization:** INT853  - **Weight quantization:** INT854- **Intended Use Cases:** This model is designed to accelerate research on language models, for use as a building block for generative AI powered features. It provides uses for general purpose AI systems and applications (primarily in English) which require:55  1. Memory/compute constrained environments.56  2. Latency bound scenarios.57  3. Reasoning and logic.58- **Out-of-scope:** This model is not specifically designed or evaluated for all downstream purposes, thus:59  1. Developers should consider common limitations of language models as they select use cases, and evaluate and mitigate for accuracy, safety, and fairness before using within a specific downstream use case, particularly for high-risk scenarios.60  2. Developers should be aware of and adhere to applicable laws or regulations (including privacy, trade compliance laws, etc.) that are relevant to their use case, including the model’s focus on English.61  3. Nothing contained in this Model Card should be interpreted as or deemed a restriction or modification to the license the model is released under.62- **Release Date:** 03/03/202563- **Version:** 1.064- **Validated on:** RHOAI 2.20, RHAIIS 3.0, RHELAI 1.565- **Model Developers:** Red Hat (Neural Magic)66 67 68### Model Optimizations69 70This model was obtained by quantizing activations and weights of [phi-4](https://huggingface.co/microsoft/phi-4) to INT8 data type.71This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x).72Weight quantization also reduces disk size requirements by approximately 50%.73 74Only weights and activations of the linear operators within transformers blocks are quantized.75Weights are quantized with a symmetric static per-channel scheme, whereas activations are quantized with a symmetric dynamic per-token scheme.76A combination of the [SmoothQuant](https://arxiv.org/abs/2211.10438) and [GPTQ](https://arxiv.org/abs/2210.17323) algorithms is applied for quantization, as implemented in the [llm-compressor](https://github.com/vllm-project/llm-compressor) library.77 78 79## Deployment80 81This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend, as shown in the example below.82 83```python84from vllm import LLM, SamplingParams85from transformers import AutoTokenizer86 87model_id = "neuralmagic-ent/phi-4-quantized.w8a8"88number_gpus = 189 90sampling_params = SamplingParams(temperature=0.7, top_p=0.8, max_tokens=256)91 92tokenizer = AutoTokenizer.from_pretrained(model_id)93 94messages = [95    {"role": "user", "content": "Give me a short introduction to large language model."},96]97 98prompts = tokenizer.apply_chat_template(messages, tokenize=False)99 100llm = LLM(model=model_id, tensor_parallel_size=number_gpus)101 102outputs = llm.generate(prompts, sampling_params)103 104generated_text = outputs[0].outputs[0].text105print(generated_text)106```107 108vLLM aslo supports OpenAI-compatible serving. See the [documentation](https://docs.vllm.ai/en/latest/) for more details.109 110<details>111  <summary>Deploy on <strong>Red Hat AI Inference Server</strong></summary>112  113```bash114podman run --rm -it --device nvidia.com/gpu=all -p 8000:8000 \115 --ipc=host \116--env "HUGGING_FACE_HUB_TOKEN=$HF_TOKEN" \117--env "HF_HUB_OFFLINE=0" -v ~/.cache/vllm:/home/vllm/.cache \118--name=vllm \119registry.access.redhat.com/rhaiis/rh-vllm-cuda \120vllm serve \121--tensor-parallel-size 8 \122--max-model-len 32768  \123--enforce-eager --model RedHatAI/phi-4-quantized.w8a8124```125​​See [Red Hat AI Inference Server documentation](https://docs.redhat.com/en/documentation/red_hat_ai_inference_server/) for more details.126</details>127 128<details>129  <summary>Deploy on <strong>Red Hat Enterprise Linux AI</strong></summary>130  131```bash132# Download model from Red Hat Registry via docker133# Note: This downloads the model to ~/.cache/instructlab/models unless --model-dir is specified.134ilab model download --repository docker://registry.redhat.io/rhelai1/phi-4-quantized-w8a8:1.5135```136 137```bash138# Serve model via ilab139ilab model serve --model-path ~/.cache/instructlab/models/phi-4-quantized-w8a8140  141# Chat with model142ilab model chat --model ~/.cache/instructlab/models/phi-4-quantized-w8a8143```144See [Red Hat Enterprise Linux AI documentation](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux_ai/1.4) for more details.145</details>146 147<details>148  <summary>Deploy on <strong>Red Hat Openshift AI</strong></summary>149  150```python151# Setting up vllm server with ServingRuntime152# Save as: vllm-servingruntime.yaml153apiVersion: serving.kserve.io/v1alpha1154kind: ServingRuntime155metadata:156 name: vllm-cuda-runtime # OPTIONAL CHANGE: set a unique name157 annotations:158   openshift.io/display-name: vLLM NVIDIA GPU ServingRuntime for KServe159   opendatahub.io/recommended-accelerators: '["nvidia.com/gpu"]'160 labels:161   opendatahub.io/dashboard: 'true'162spec:163 annotations:164   prometheus.io/port: '8080'165   prometheus.io/path: '/metrics'166 multiModel: false167 supportedModelFormats:168   - autoSelect: true169     name: vLLM170 containers:171   - name: kserve-container172     image: quay.io/modh/vllm:rhoai-2.20-cuda # CHANGE if needed. If AMD: quay.io/modh/vllm:rhoai-2.20-rocm173     command:174       - python175       - -m176       - vllm.entrypoints.openai.api_server177     args:178       - "--port=8080"179       - "--model=/mnt/models"180       - "--served-model-name={{.Name}}"181     env:182       - name: HF_HOME183         value: /tmp/hf_home184     ports:185       - containerPort: 8080186         protocol: TCP187```188 189```python190# Attach model to vllm server. This is an NVIDIA template191# Save as: inferenceservice.yaml192apiVersion: serving.kserve.io/v1beta1193kind: InferenceService194metadata:195  annotations:196    openshift.io/display-name: phi-4-quantized.w8a8 # OPTIONAL CHANGE197    serving.kserve.io/deploymentMode: RawDeployment198  name: phi-4-quantized.w8a8        # specify model name. This value will be used to invoke the model in the payload199  labels:200    opendatahub.io/dashboard: 'true'201spec:202  predictor:203    maxReplicas: 1204    minReplicas: 1205    model:206      modelFormat:207        name: vLLM208      name: ''209      resources:210        limits:211          cpu: '2'			# this is model specific212          memory: 8Gi		# this is model specific213          nvidia.com/gpu: '1'	# this is accelerator specific214        requests:			# same comment for this block215          cpu: '1'216          memory: 4Gi217          nvidia.com/gpu: '1'218      runtime: vllm-cuda-runtime	# must match the ServingRuntime name above219      storageUri: oci://registry.redhat.io/rhelai1/modelcar-phi-4-quantized-w8a8:1.5220    tolerations:221    - effect: NoSchedule222      key: nvidia.com/gpu223      operator: Exists224```225 226```bash227# make sure first to be in the project where you want to deploy the model228# oc project <project-name>229# apply both resources to run model230# Apply the ServingRuntime231oc apply -f vllm-servingruntime.yaml232# Apply the InferenceService233oc apply -f qwen-inferenceservice.yaml234```235 236```python237# Replace <inference-service-name> and <cluster-ingress-domain> below:238# - Run `oc get inferenceservice` to find your URL if unsure.239# Call the server using curl:240curl https://<inference-service-name>-predictor-default.<domain>/v1/chat/completions241        -H "Content-Type: application/json" \242        -d '{243    "model": "phi-4-quantized.w8a8",244    "stream": true,245    "stream_options": {246        "include_usage": true247    },248    "max_tokens": 1,249    "messages": [250        {251            "role": "user",252            "content": "How can a bee fly when its wings are so small?"253        }254    ]255}'256```257 258See [Red Hat Openshift AI documentation](https://docs.redhat.com/en/documentation/red_hat_openshift_ai/2025) for more details.259</details>260 261 262## Creation263 264<details>265  <summary>Creation details</summary>266  This model was created with [llm-compressor](https://github.com/vllm-project/llm-compressor) by running the code snippet below. 267 268 269  ```python270  from transformers import AutoModelForCausalLM, AutoTokenizer271  from llmcompressor.modifiers.quantization import GPTQModifier272  from llmcompressor.modifiers.smoothquant import SmoothQuantModifier273  from llmcompressor.transformers import oneshot274  from datasets import load_dataset275 276  # Load model277  model_stub = "microsoft/phi-4"278  model_name = model_stub.split("/")[-1]279  280  num_samples = 1024281  max_seq_len = 8192282  283  tokenizer = AutoTokenizer.from_pretrained(model_stub)284  285  model = AutoModelForCausalLM.from_pretrained(286      model_stub,287      device_map="auto",288      torch_dtype="auto",289  )290  291  def preprocess_fn(example):292    return {"text": tokenizer.apply_chat_template(example["messages"], add_generation_prompt=False, tokenize=False)}293  294  ds = load_dataset("neuralmagic/LLM_compression_calibration", split="train")295  ds = ds.map(preprocess_fn)296  297  # Configure the quantization algorithm and scheme298  recipe = [299      SmoothQuantModifier(300          smoothing_strength=0.7,301          mappings=[302              [["re:.*qkv_proj"], "re:.*input_layernorm"],303              [["re:.*gate_up_proj"], "re:.*post_attention_layernorm"],304          ],305      ),306      GPTQModifier(307          ignore=["lm_head"],308          sequential_targets=["Phi3DecoderLayer"],309          dampening_frac=0.01,310          targets="Linear",311          scheme="W8A8",312      ),313  ]314  315  # Apply quantization316  oneshot(317      model=model,318      dataset=ds, 319      recipe=recipe,320      max_seq_length=max_seq_len,321      num_calibration_samples=num_samples,322  )323  324  # Save to disk in compressed-tensors format325  save_path = model_name + "-quantized.w8a8"326  model.save_pretrained(save_path)327  tokenizer.save_pretrained(save_path)328  print(f"Model and tokenizer saved to: {save_path}")329  ```330</details>331 332 333 334## Evaluation335 336The model was evaluated on the OpenLLM leaderboard tasks (version 1) with the [lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) and the [vLLM](https://docs.vllm.ai/en/stable/) engine, using the following command:337```338lm_eval \339  --model vllm \340  --model_args pretrained="neuralmagic-ent/phi-4-quantized.w8a8",dtype=auto,gpu_memory_utilization=0.6,max_model_len=4096,enable_chunk_prefill=True,tensor_parallel_size=1 \341  --tasks openllm \342  --batch_size auto343```344 345### Accuracy346 347#### Open LLM Leaderboard evaluation scores348<table>349  <tr>350   <td><strong>Benchmark</strong>351   </td>352   <td><strong>phi-4</strong>353   </td>354   <td><strong>phi-4-quantized.w8a8<br>(this model)</strong>355   </td>356   <td><strong>Recovery</strong>357   </td>358  </tr>359  <tr>360   <td>MMLU (5-shot)361   </td>362   <td>80.30363   </td>364   <td>80.39365   </td>366   <td>100.1%367   </td>368  </tr>369  <tr>370   <td>ARC Challenge (25-shot)371   </td>372   <td>64.42373   </td>374   <td>64.33375   </td>376   <td>99.9%377   </td>378  </tr>379  <tr>380   <td>GSM-8K (5-shot, strict-match)381   </td>382   <td>90.07383   </td>384   <td>90.30385   </td>386   <td>100.3%387   </td>388  </tr>389  <tr>390   <td>Hellaswag (10-shot)391   </td>392   <td>84.37393   </td>394   <td>84.30395   </td>396   <td>99.9%397   </td>398  </tr>399  <tr>400   <td>Winogrande (5-shot)401   </td>402   <td>80.58403   </td>404   <td>79.95405   </td>406   <td>99.2%407   </td>408  </tr>409  <tr>410   <td>TruthfulQA (0-shot, mc2)411   </td>412   <td>59.37413   </td>414   <td>58.82415   </td>416   <td>99.1%417   </td>418  </tr>419  <tr>420   <td><strong>Average</strong>421   </td>422   <td><strong>76.52</strong>423   </td>424   <td><strong>76.35</strong>425   </td>426   <td><strong>99.8%</strong>427   </td>428  </tr>429</table>430 431