# Qwen2.5-7B

A 7B dense chat model (AWQ INT4) on one NVIDIA A16 on Vultr.

Source: /recipes/qwen2.5-7b/

<!-- vale write-good.Passive = NO -->
A 7B dense chat model served from an AWQ INT4 quantization on one NVIDIA A16
on Vultr: one `Standalone` engine, no cache, weights pulled straight from
Hugging Face. The A16 slice on the `vcg-a16-6c-64g-16vram` plan has 16 GiB of
VRAM, so the INT4 weights (~5 GiB) fit with headroom for KV cache;
`--gpu-memory-utilization=0.85` and `--enforce-eager` keep the engine inside
the small card.

This recipe was run end to end on Vultr (`ewr`); the `InferenceClass`,
`InferenceCluster`, and `ModelDeployment` are the exact manifests from that
run. GPU plan availability varies by Vultr region, so check the plan is offered
in your region before applying. Apply the platform side first, then the ML
side.

## Validated deployments

{{< validated-deployments >}}

## Platform

{{< manifests "recipes/qwen2.5-7b/inference-class.yaml" >}}

{{< manifests "recipes/qwen2.5-7b/inference-cluster.yaml" >}}

## Deployment

{{< manifests "recipes/qwen2.5-7b/model-deployment.yaml" >}}

{{< manifests "recipes/qwen2.5-7b/model-service.yaml" >}}
<!-- vale write-good.Passive = YES -->
