Deploy vLLM with Flat Model Architecture
Set up vLLM's new Flat Model abstraction for day-0 model support, faster cold starts, and the Model Runner V2 execution engine — the Q3 2026 production upgrade path.
vLLM's Q3 2026 roadmap ships two major architectural upgrades: the Flat Model abstraction and Model Runner V2 (MRV2). Flat Model simplifies the model integration layer so new models are supported day-0 without custom adapters. MRV2 replaces the legacy execution engine with a cleaner, more composable runner. Together, they reduce cold start time, improve production stability, and make day-0 model support the default rather than the exception.
This recipe deploys vLLM with Flat Model and MRV2 enabled, using a recent model as the test case.
- vLLM running with Flat Model architecture and MRV2 execution engine
- Day-0 support for new model architectures without custom code
- Reduced cold start time vs. legacy V0 architecture
- OpenAI-compatible API endpoint
- Production-ready serving with improved failure ergonomics
- Docker and NVIDIA Container Toolkit installed
- GPU with at least 24GB VRAM (for a 13B-class model; more for larger)
- vLLM v0.10+ (Flat Model and MRV2 are rolling out across Q3 2026 — check the release notes for your version)
| Check | Command |
|---|---|
| Container running | docker ps |
| API up | curl http://localhost:8000/v1/models |
| Flat Model active | docker logs vllm-flat 2>&1 | grep -i flat |
| MRV2 active | docker logs vllm-flat 2>&1 | grep -i mrv2 |
| GPU active | docker exec vllm-flat nvidia-smi |
| Speculative decoding | Check logs for draft token acceptance rate |
| Symptom | Fix |
|---|---|
| Flat Model not supported for your model | Not all architectures are migrated yet. Check the vLLM model support list. Q3 2026 target: top 20 architectures. |
| MRV2 missing features | MRV2 is newer than V0; some niche features may not be ported yet. Fall back to V0 with VLLM_USE_MRV2=0 if needed. |
| Cold start still slow | Flat Model reduces but does not eliminate cold start. The Q3 roadmap also targets KV Cache Manager redesign and scheduler refactoring for further improvements. |
| DFlash drafter not found | Ensure the model ships with a drafter, or provide --spec-model with a compatible draft model. |
| Out of memory | Muse Glimmer 30B at full precision needs 55GB+. Use a quantized variant or a smaller model. |
- You want day-0 support for new models without waiting for vLLM to add custom adapters
- You need reduced cold start time for autoscaling deployments
- You want the production stability improvements in the Q3 2026 roadmap
- You are deploying Muse Glimmer 30B or other recently released models with speculative decoding
- Add LiteLLM Proxy in front for multi-model routing
- Connect Open WebUI for a chat interface
- See the vLLM Q3 2026 Roadmap for the full list of architectural changes
Steps
docker pull vllm/vllm-openai:latest
Check the release notes to confirm Flat Model and MRV2 are enabled in your version. If you need a nightly build:
docker pull vllm/vllm-openai:nightly
docker run -d \
--name vllm-flat \
--gpus all \
-p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e VLLM_USE_FLAT_MODEL=1 \
-e VLLM_USE_MRV2=1 \
--restart unless-stopped \
vllm/vllm-openai:latest \
--model meta-models/Muse-Glimmer-30B \
--spec-type draft-dflash \
--spec-draft-n-max 15 \
--max-model-len 131072
Key flags:
VLLM_USE_FLAT_MODEL=1: Enables the Flat Model abstractionVLLM_USE_MRV2=1: Enables Model Runner V2 execution engine--spec-type draft-dflash: Enables DFlash speculative decoding (Muse Glimmer ships with this drafter)--max-model-len 131072: Sets the context window to the model's full 131K
curl http://localhost:8000/v1/models
You should see the model listed. Test a completion:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-models/Muse-Glimmer-30B",
"messages": [{"role": "user", "content": "Write a Python function to check if a number is prime."}],
"max_tokens": 200
}'
Check the server logs for Flat Model initialization:
docker logs vllm-flat 2>&1 | grep -i "flat model\|mrv2"
You should see confirmation that the Flat Model path and MRV2 runner are active.
Point any OpenAI-compatible client at http://localhost:8000/v1:
- Hermes Agent: Set
OPENAI_API_BASE=http://localhost:8000/v1in your.env - Open WebUI: Set the model endpoint to
http://localhost:8000/v1 - LiteLLM Gateway: Add as a custom provider
Recipe verified 2026-08-19. Commands are tested but your environment may differ.
Browse related services