# Qwen2.5-72B

A 72B dense chat model (AWQ INT4) on one 80 GB GPU, on AKS and Nebius.

Source: /recipes/qwen2.5-72b/

<!-- vale write-good.Passive = NO -->
A 72B dense chat model served from an AWQ INT4 quantization on one 80 GB GPU
per replica: one `Standalone` engine fed by a `ModelCache`. The platform side
covers an A100 on AKS and an H100 on Nebius, and the ML side is the same
manifest for both. The deployment has two replicas and no
`clusterSelector`. Each pool has exactly one GPU, so with both platforms
applied one replica runs on each. The service then splits traffic between the
two GPUs by weight, which the last section uses to compare them. To serve on
just one platform, apply one tab and drop `replicas` to 1.

These manifests mirror the repository's AKS and Nebius demos. Apply the
platform side first, then the ML side.

## Validated deployments

{{< validated-deployments >}}

## Platform

{{< tabs >}}
{{< tab "AKS" >}}
{{< manifests "recipes/qwen2.5-72b/inference-class-aks.yaml" >}}

{{< manifests "recipes/qwen2.5-72b/inference-cluster-aks.yaml" >}}
{{< /tab >}}
{{< tab "Nebius" >}}
{{< manifests "recipes/qwen2.5-72b/inference-class-nebius.yaml" >}}

{{< manifests "recipes/qwen2.5-72b/inference-cluster-nebius.yaml" >}}
{{< /tab >}}
{{< /tabs >}}

## Deployment

{{< manifests "recipes/qwen2.5-72b/model-cache.yaml" >}}

{{< manifests "recipes/qwen2.5-72b/model-deployment.yaml" >}}

{{< manifests "recipes/qwen2.5-72b/model-service.yaml" >}}

## Compare the A100 and the H100

Replicas are fleet-wide, not per-cluster: the deployment's `replicas: 2` means
two complete serving instances, and because each pool has exactly one 80 GB GPU
they land one on the A100 and one on the H100. Modelplane labels each
replica's endpoint with the cluster it runs on, so the service can split
traffic between the platforms by weight. This service pairs the deployment
label with each cluster label and gives each GPU half of the live traffic
under the same model name:

{{< manifests "recipes/qwen2.5-72b/model-service-split.yaml" >}}

Both GPUs now serve the same workload, so their engine metrics give a direct
performance comparison: scrape each replica's latency and throughput as in
[Monitor the Fleet]({{< ref "/platform/telemetry.md" >}}) and
read the two side by side. Weights are relative, so you can shift traffic
toward whichever platform performs better (80/20, or as far as 100/0) without
touching the deployment.
<!-- vale write-good.Passive = YES -->
