# Documentation for Kamiwaza 0.9.3

This version of Kamiwaza is no longer actively maintained. For the current GA release, see [1.0.1](https://docs.kamiwaza.ai/).

---

## Base System Requirements

### Supported Operating Systems & Architecture
- **Linux**:
  - Ubuntu 24.04 and 22.04 LTS via .deb package installation (x64/amd64 architecture only)
  - Redhat Enterprise Linux (RHEL) 9
- **Windows**: 11 (x64 architecture) via WSL with MSI installer
- **macOS**: 12.0 or later, Apple Silicon (ARM64) only (community edition only)

### CPU Requirements
- **Architecture**:
  - Linux: x64/amd64 (64-bit)
  - macOS: ARM64 (Apple Silicon) only
  - Windows: x64 (64-bit)
- **Minimum Cores**: 8+ cores
- **Recommended Cores**: 16+ cores for CPU-based inference workloads

### Core Software Requirements
- **Python**: Python 3.10 for tarball installations; Python 3.12 for .deb/.msi installations
- **Docker**: Docker Engine with Compose v2
- **Node.js**: 22.x (installed via NVM during setup)
- **Browser**: Chrome Version 141+ (tested and recommended)
- **GPU Support**: NVIDIA GPU with compute capability 7.0+ (Linux only) or NVIDIA RTX/Intel Arc (Windows via WSL)

### Memory Requirements
#### System RAM
- **Minimum**: 16GB RAM
- **Recommended**: 32GB+ RAM for CPU-based inference workloads
- **GPU Workloads**: 16GB+ system RAM (32GB+ recommended)

#### GPU Memory (vRAM)
- **GPU Inference**: 16GB+ vRAM required
- **Recommended**: 32GB+ vRAM for optimal GPU inference performance

### Storage Requirements
#### Storage Performance
- **Required**: SSD (Solid State Drive)
- **Preferred**: NVMe SSD for optimal performance
- **Minimum**: SATA SSD
- **Note**: Models weights can be on a separate HDD but load time will increase

#### Storage Capacity
**Linux/macOS**
- **Minimum**: 100GB free disk space
- **Recommended**: 200GB+ free disk space
**Windows**
- **Minimum**: 100GB free disk space
- **Recommended**: 200GB+ free space on SSD

## Hardware Recommendation Tiers
Kamiwaza is a distributed AI platform built on Ray that supports both CPU-only and GPU-accelerated inference. Hardware requirements vary significantly based on:
- **Model size**: From 0.6B to 70B+ parameters
- **Deployment scale**: Single-node development vs multi-node production
- **Inference engine**: LlamaCpp (CPU/GPU), VLLM (GPU), MLX (Apple Silicon)
- **Workload type**: Interactive chat, batch processing, RAG pipelines

### GPU Memory Requirements by Model Size
The table below provides real-world GPU memory requirement estimates for representative models at different scales.

| Model Example | Parameters | Minimum vRAM | Notes |
| --- | --- | --- | --- |
| **GPT-OSS 20B** | 20B | 24GB | Fits 1x 24GB GPU |
| **GPT-OSS 120B** | 120B | 80GB | 1x H100/H200 or 2x A100 80GB recommended |
| **Qwen 3 235B A22B** | 235B | 150GB | 2x H200 (282GB) or 2x B200 (384GB) ideal |
| **Qwen 3-VL 235B A22B** | 235B | 150GB | Budget +20-30% vRAM for high-res vision inputs |

### Tier 1: Development & Small Models
**Use Case:** Local development, testing, small to medium model deployment (up to 13B parameters)
**Hardware Specifications:**
- **CPU:** 8-16 cores / 16-32 threads
- **RAM:** 32GB (16GB minimum)
- **Storage:** 200GB NVMe SSD (100GB minimum)
- **GPU:** Optional - Single GPU with 16-24GB VRAM

### Tier 2: Production - Medium to Large Models
**Use Case:** Production deployment of medium to large models (13B-70B parameters), high throughput
**Hardware Specifications:**
- **CPU:** 32 cores / 64 threads
- **RAM:** 128-256GB system RAM
- **Storage:** 1-2TB NVMe SSD

### Tier 3: Enterprise Multi-Node Cluster
**Use Case:** Enterprise deployment with multiple models, high availability, horizontal scaling, 99.9%+ SLA
**Cluster Architecture:**
- **Head Node (Control Plane):** 16 cores / 32 threads, 64GB RAM, 500GB NVMe SSD
- **Worker Nodes (3+ nodes for HA):** 32-64 cores / 64-128 threads per node, 256-512GB RAM per node, 2TB NVMe SSD per node

## Cloud Provider Instance Mapping
### AWS EC2 Instance Types
| Tier | Instance Type | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `m6i.2xlarge` | 8 | 32GB | None | 200GB gp3 |
| **Tier 1: With GPU** | `g5.xlarge` | 4 | 16GB | 1x A10G (24GB) | 200GB gp3 |

### Google Cloud Platform (GCP) Instance Types
| Tier | Machine Type | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `n2-standard-8` | 8 | 32GB | None | 200GB SSD |
| **Tier 1: With GPU** | `n1-standard-8` + `1x T4` | 8 | 30GB | 1x T4 (16GB) | 200GB SSD |

### Microsoft Azure Instance Types
| Tier | VM Size | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `Standard_D8s_v5` | 8 | 32GB | None | 200GB Premium SSD |
| **Tier 1: With GPU** | `Standard_NC4as_T4_v3` | 4 | 28GB | 1x T4 (16GB) | 200GB Premium SSD |

## Dependencies & Components
### Required System Packages
See platform-specific installation instructions
### NVIDIA Components (Linux GPU Support)
- NVIDIA Driver (550-server recommended)
- NVIDIA Container Toolkit
- nvidia-docker2
### Docker Configuration Requirements
- Docker Engine with Compose v2
- User must be in docker group

### Network Configuration
#### Network Bandwidth Requirements
**Single Node Deployment**
- **Minimum:** 1 Gbps (for model downloads, API traffic)
- **Recommended:** 10 Gbps (for high-throughput inference)

## Special Considerations
### Apple Silicon (M-Series)
**MLX Engine Support:**
- Unified memory architecture (shared CPU/GPU RAM)
- Performance for models up to 13B parameters; reasonable performance for larger models when context is appropriately restricted and RAM is available.
