Placement Deployment Guide | Kamiwaza Docs

Kamiwaza 1.0.1 Deployment Guide

This is documentation for Kamiwaza 1.0.1, which is no longer actively maintained. For the current GA release, see 1.0.1.

Version: 1.0.1

Overview

Placement is automatic — deploying a model with placement is the same deploy flow you already use. This guide shows how to deploy multiple models onto shared GPUs, read where each deployment landed, and troubleshoot placement problems.

For background on how placement decides, read the Model Placement Overview and Fractional GPU Serving.

Before you start

Deploy a model in the UI

  1. Navigate to the Models page and select the model.
  2. Click Deploy. In Novice Mode, Kamiwaza picks a platform-appropriate variant with sensible defaults; in Advanced Mode you can select the engine and parameters yourself. See the GUI Walkthrough.
  3. Kamiwaza estimates the model's footprint, picks a GPU (or memory pool) with enough free budget, and starts the deployment. No placement input is required.
  4. Watch the status move through DEPLOYING and INITIALIZING to DEPLOYED. The statuses are described in Model Deployment.

To run a second model on the same hardware, just deploy it the same way. If the combined budgets fit, both models run side by side; if not, the second deployment fails fast with a NoFit error.

Deploy a model with the SDK

from kamiwaza_sdk import KamiwazaClient

client = KamiwazaClient(base_url="https://<your-host>/api")

deployment_id = client.serving.deploy_model(model_id=model_id)

deployment_id = client.serving.deploy_model(
    repo_id="Qwen/Qwen3-8B",
    m_config_id=config_id,
)

status = client.serving.get_deployment_status(deployment_id)

The deploy API is asynchronous: the server accepts the request and returns the deployment ID immediately. By default the SDK blocks client-side (wait=True), polling until the deployment reaches DEPLOYED; it raises DeploymentFailedError if the deployment reaches a terminal failure state (FAILED, ERROR, or MUST_REDOWNLOAD) and TimeoutError if the deployment is not ready within timeout_seconds. Pass wait=False to get the deployment ID back as soon as the server accepts the request, then observe progress with get_deployment_status or `wait_deployment_ready.

Monitor placement

Open a deployment's details to see where it landed:

Field Meaning
topology managed_cluster or standalone_cluster
node_name The node the model was placed on
gpu_index, gpu_vendor Which GPU on the node, and its vendor
hardware_class hardware_isolated, software_shared, or unified_memory
sharing_class How the device is shared
allocated_capacity_gb The memory budget reserved for this deployment, in GB

Verify

  1. Confirm each deployment shows DEPLOYED in the UI.
  2. Send a short test prompt to each deployment's endpoint (shown in the UI).
  3. If you deployed multiple models to one GPU, confirm both respond — they are serving concurrently from the same card.

Troubleshooting

The deployment fails immediately with a placement (NoFit) error

This is placement telling you the model does not fit anywhere, before any pod starts. The deployment shows FAILED with last_error_code set to NoFitError and the no-fit reason in last_error_message. The quick version:

A SharingNotConfigured notice appears

On a managed cluster where the GPU Operator is installed but no sharing strategy is configured, deployments succeed as whole-GPU and carry a SharingNotConfigured notice. This is informational: density is limited to one model per GPU until your cluster admin enables a sharing strategy.

The deployment reaches ERROR or FAILED after placement

Placement succeeded, but the model failed at runtime. The deployment record carries three fields that tell you what happened:

The deployment sits in INITIALIZING

Normal for a short period: routing is up but the model is still loading. If it persists well beyond the expected load time, check last_error_code for STARTUP_TIMEOUT.

Checking capacity on a standalone cluster

On a standalone cluster you can inspect what placement sees:

# GPU labels detected on a node
kubectl get node <node-name> -o jsonpath='{.metadata.labels}' | tr ',' '\n' | grep gpu
# Per-GPU memory budgets advertised on a node (GB)
kubectl get node <node-name> -o jsonpath='{.status.allocatable}' | tr ',' '\n' | grep vram-gb-gpu