Team Ai
Modelpublic

MiniMaxAI/MiniMax-M2.5

sourceHugging Faceotherupdated 7mo agoView on Hugging Face
1.5klikes579kdownloads
vllm_deploy_guide.md118 linesDownload Raw Back to docs
1# MiniMax M2.5 Model vLLM Deployment Guide2 3[English Version](./vllm_deploy_guide.md) | [Chinese Version](./vllm_deploy_guide_cn.md)4 5We recommend using [vLLM](https://docs.vllm.ai/en/stable/) to deploy the [MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5) model. vLLM is a high-performance inference engine with excellent serving throughput, efficient and intelligent memory management, powerful batch request processing capabilities, and deeply optimized underlying performance. We recommend reviewing vLLM's official documentation to check hardware compatibility before deployment.6 7## Applicable Models8 9This document applies to the following models. You only need to change the model name during deployment.10 11- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)12- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)13- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)14 15The deployment process is illustrated below using MiniMax-M2.5 as an example.16 17## System Requirements18 19- OS: Linux20 21- Python: 3.9 - 3.1222 23- GPU:24 25  - compute capability 7.0 or higher26 27  - Memory requirements: 220 GB for weights, 240 GB per 1M context tokens28 29The following are recommended configurations; actual requirements should be adjusted based on your use case:30 31- **96G x4** GPU: Supports a total KV Cache capacity of 400K tokens.32 33- **144G x8** GPU: Supports a total KV Cache capacity of up to 3M tokens.34 35> **Note**: The values above represent the total aggregate hardware KV Cache capacity. The maximum context length per individual sequence remains **196K** tokens.36 37## Deployment with Python38 39It is recommended to use a virtual environment (such as **venv**, **conda**, or **uv**) to avoid dependency conflicts. 40 41We recommend installing vLLM in a fresh Python environment:42 43```bash44uv venv45source .venv/bin/activate46uv pip install vllm --torch-backend=auto47```48 49Run the following command to start the vLLM server. vLLM will automatically download and cache the MiniMax-M2.5 model from Hugging Face.50 514-GPU deployment command:52 53```bash54SAFETENSORS_FAST_GPU=1 vllm serve \55    MiniMaxAI/MiniMax-M2.5 --trust-remote-code \56    --tensor-parallel-size 4 \57    --enable-auto-tool-choice --tool-call-parser minimax_m2 \58    --reasoning-parser minimax_m2_append_think59```60 618-GPU deployment command:62 63```bash64SAFETENSORS_FAST_GPU=1 vllm serve \65    MiniMaxAI/MiniMax-M2.5 --trust-remote-code \66    --enable_expert_parallel --tensor-parallel-size 8 \67    --enable-auto-tool-choice --tool-call-parser minimax_m2 \68    --reasoning-parser minimax_m2_append_think 69```70 71## Testing Deployment72 73After startup, you can test the vLLM OpenAI-compatible API with the following command:74 75```bash76curl http://localhost:8000/v1/chat/completions \77    -H "Content-Type: application/json" \78    -d '{79        "model": "MiniMaxAI/MiniMax-M2.5",80        "messages": [81            {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},82            {"role": "user", "content": [{"type": "text", "text": "Who won the world series in 2020?"}]}83        ]84    }'85```86 87## Common Issues88 89### MiniMax-M2 model is not currently supported90 91This vLLM version is outdated. Please upgrade to the latest version.92 93### torch.AcceleratorError: CUDA error: an illegal memory access was encountered94Add `--compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"` to the startup parameters to resolve this issue. For example:95 96```bash97SAFETENSORS_FAST_GPU=1 vllm serve \98    MiniMaxAI/MiniMax-M2.5 --trust-remote-code \99    --enable_expert_parallel --tensor-parallel-size 8 \100    --enable-auto-tool-choice --tool-call-parser minimax_m2 \101    --reasoning-parser minimax_m2_append_think \102    --compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"103```104 105### Output is garbled106 107If you encounter corrupted output when using vLLM to serve these models, you can upgrade to the nightly version (ensure it is a version after commit [cf3eacfe58fa9e745c2854782ada884a9f992cf7](https://github.com/vllm-project/vllm/commit/cf3eacfe58fa9e745c2854782ada884a9f992cf7))108 109## Getting Support110 111If you encounter any issues while deploying the MiniMax model:112 113- Contact our technical support team through official channels such as email at [model@minimax.io](mailto:model@minimax.io)114 115- Submit an issue on our [GitHub](https://github.com/MiniMax-AI) repository116 117We continuously optimize the deployment experience for our models. Feedback is welcome!118