This document applies to the Modelplane main branch and not to the latest release v0.5.
A 7B dense chat model served from an AWQ INT4 quantization on one NVIDIA A16
on Vultr: one Standalone engine, no cache, weights pulled straight from
Hugging Face. The A16 slice on the vcg-a16-6c-64g-16vram plan has 16 GiB of
VRAM, so the INT4 weights (~5 GiB) fit with headroom for KV cache;
--gpu-memory-utilization=0.85 and --enforce-eager keep the engine inside
the small card.
This recipe was run end to end on Vultr (ewr); the InferenceClass,
InferenceCluster, and ModelDeployment are the exact manifests from that
run. GPU plan availability varies by Vultr region, so check the plan is offered
in your region before applying. Apply the platform side first, then the ML
side.
Validated deployments
Platform
# InferenceClass for the A16 shape, validated serving Qwen2.5-7B on Vultr.
#
# One NVIDIA A16 on a Vultr vcg-a16-6c-64g-16vram plan: 16 GiB of VRAM is
# enough for a 7B AWQ INT4 model with KV cache headroom. GPU plans are
# region-gated on Vultr, so check the plan is offered in the cluster's region
# before applying.
apiVersion: modelplane.ai/v1alpha1
kind: InferenceClass
metadata:
name: vultr-a16-1x
spec:
description: "Vultr vcg-a16-6c-64g-16vram, 1x NVIDIA A16 16GB"
provisioning:
provider: Vultr
vultr:
plan: vcg-a16-6c-64g-16vram
accelerator:
type: nvidia-a16
count: 1
devices:
- name: gpu
claim: DRA
driver: gpu.nvidia.com
deviceClassName: gpu.nvidia.com
count: 1
attributes:
architecture: { string: Ampere }
cudaComputeCapability: { version: "8.6.0" }
capacity:
memory: { value: "16384Mi" }
# A Vultr Kubernetes Engine cluster in ewr (New Jersey) with a single A16 node
# offering the class above.
apiVersion: modelplane.ai/v1alpha1
kind: InferenceCluster
metadata:
name: vultr-a16
labels:
modelplane.ai/region: ewr
spec:
cluster:
source: Vultr
vultr:
region: ewr
nodePools:
- name: gpu-a16
className: vultr-a16-1x
nodeCount: 1
minNodeCount: 1
maxNodeCount: 1
Deployment
# Qwen2.5-7B served on a single NVIDIA A16, validated end to end on Vultr.
#
# A 7B AWQ INT4 model is a single Standalone engine: one self-contained vLLM
# pod, no ModelCache, weights pulled straight from Hugging Face. The 16 GiB
# card is the constraint: --gpu-memory-utilization=0.85 leaves headroom for
# the card's overhead, and --enforce-eager skips CUDA graph capture to save
# memory at some latency cost. It carries no clusterSelector, so device
# capacity alone matches it to any compatible GPU in the fleet.
apiVersion: modelplane.ai/v1alpha1
kind: ModelDeployment
metadata:
name: qwen-7b
namespace: ml-team
spec:
replicas: 1
template:
spec:
engines:
- name: qwen-7b
members:
- role: Standalone
nodeSelector:
devices:
- name: gpu
count: 1
selectors:
- cel: |
device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("14Gi")) >= 0
template:
spec:
containers:
- name: engine
image: vllm/vllm-openai:v0.9.2
args:
- --model=Qwen/Qwen2.5-7B-Instruct-AWQ
- --served-model-name=$(MODELPLANE_SERVED_MODEL_NAME)
- --max-model-len=8192
- --gpu-memory-utilization=0.85
- --enforce-eager
# Exposes the deployment as one model a caller can name.
# Name the model, ml-team/qwen-7b, in a request to an InferenceGateway
# that serves it.
apiVersion: modelplane.ai/v1alpha1
kind: ModelService
metadata:
name: qwen-7b
namespace: ml-team
spec:
endpoints:
- name: qwen-7b
selector:
matchLabels:
modelplane.ai/deployment: qwen-7b