# Hardware Requirements

Documentation for Kamiwaza **0.11.0**.

## Hardware Requirements

### CPU
- **Minimum Cores**: 8+ cores
- **Recommended Cores**: 16+ cores for CPU-based inference workloads
- **Architecture**:
  - Linux: x64/amd64 (64-bit)
  - Windows: x64 (64-bit)
  - macOS: ARM64 (Apple Silicon) only

### Memory
#### System RAM
| Mode | Minimum | Recommended | Notes |
| --- | --- | --- | --- |
| **Lite Mode** | 16GB | 32GB | SQLite database; limited capacity for apps/tools |
| **Full Mode** | 32GB | 64GB+ | CockroachDB + DataHub; production workloads |
| **GPU Workloads** | 32GB | 64GB+ | System RAM alongside GPU vRAM |

#### GPU Memory (vRAM)
- **GPU Inference**: 16GB+ vRAM required
- **Recommended**: 32GB+ vRAM for optimal GPU inference performance

### GPU (Optional)
Kamiwaza supports multiple GPU and accelerator platforms:

**Discrete GPUs:**
- NVIDIA GPUs with compute capability 7.0+ (Linux)
- NVIDIA RTX / Intel Arc (Windows via WSL)

**Unified Memory Systems:**
- **NVIDIA DGX Spark** 
- **AMD Ryzen AI Max+ 395** 
- **Apple Silicon M-series**

### Storage
Storage requirements are the same across all platforms.

#### Storage Performance
- **Required**: SSD (Solid State Drive)
- **Preferred**: NVMe SSD for optimal performance
- **Minimum**: SATA SSD
- **Note**: Model weights can be on a separate HDD but load times will increase significantly

#### Storage Capacity
- **Minimum**: 100GB free disk space
- **Recommended**: 200GB+ free disk space
- **Enterprise Edition**: Additional space for /opt/kamiwaza persistence

#### Capacity Planning
| Component | Minimum | Recommended | Notes |
| --- | --- | --- | --- |
| **Operating System** | 20GB | 50GB | Ubuntu/RHEL base + dependencies |
| **Kamiwaza Platform** | 50GB | 50GB | Python environment, Ray, services |
| **Model Storage** | 50GB | 500GB+ | Depends on number and size of models |
| **Database** | 10GB | 50GB | CockroachDB for metadata |
| **Vector Database** | 10GB | 100GB+ | For embeddings (if enabled) |
| **Logs & Metrics** | 10GB | 50GB | Rotated logs, Ray dashboard data |
| **Scratch Space** | 20GB | 100GB | Temporary files, downloads, builds |
| **Total** | **170GB** | **900GB+** | |

#### Storage Performance Requirements
**Local Storage (Single Node):**
- **Minimum:** SATA SSD (500 MB/s sequential read)
- **Recommended:** NVMe SSD (2000+ MB/s sequential read)

### Supported Operating Systems
#### Linux
- **Ubuntu**: 24.04 and 22.04 LTS via .deb package installation (x64/amd64 architecture only)
- **Red Hat Enterprise Linux (RHEL)**: 9

#### Windows
- **Windows 11** (x64 architecture) via WSL with MSI installer
- Requires Windows Subsystem for Linux (WSL) installed and enabled
- Administrator access required for initial setup

#### macOS
- **macOS 15.0 (Sequoia) or later**, Apple Silicon (ARM64) only

## Software Dependencies
### Pre-requisites (User Must Install)
Before running the Kamiwaza installer, ensure the following are installed:
| Component | Requirement | Installation Guide |
| --- | --- | --- |
| **Docker** | Docker Engine 24.0+ with Compose 2.23+ | [Docker Install Guide](https://docs.docker.com/engine/install/) |
| **Browser** | Chrome 141+ (tested and recommended) | [Download Chrome](https://www.google.com/chrome/) |

### Verifying System Requirements
Use these commands to verify your system meets the requirements before installation.

### Docker
```bash
docker --version
```

### Python
```bash
python3 --version
```

## Hardware Recommendation Tiers
### GPU Memory Requirements by Model Size
| Model Example | Parameters | Minimum vRAM | Notes |
| --- | --- | --- | --- |
| **GPT-OSS 20B** | 20B | 24GB | Includes weights + 1-batch max context; fits 1x 24GB GPU |
| **GPT-OSS 120B** | 120B | 80GB | ~40GB weights + 1-batch max context; 1x H100/H200 or 2x A100 80GB recommended |
| **Qwen 3 235B A22B** | 235B | 150GB | ~120GB weights + 1-batch max context; 2x H200 (282GB) or 2x B200 (384GB) ideal for max context |
| **Qwen 3-VL 235B A22B** | 235B | 150GB | Same base minimum (includes 1-batch max context); budget +20-30% vRAM for high-res vision inputs |

### Tier 1: Development & Small Models
**Use Case:** Local development, testing, small to medium model deployment (up to 13B parameters)

**Hardware Specifications:**
- **CPU:** 8-16 cores / 16-32 threads
- **RAM:** 32GB (16GB minimum for lite mode only)
- **Storage:** 200GB NVMe SSD (100GB minimum)
- **GPU:**Optional - Single GPU with 16-24GB VRAM
  - NVIDIA RTX 4090 (24GB)  
  - NVIDIA RTX 4080 (16GB)  
  - NVIDIA T4 (16GB)  
- **Network:** 1-10 Gbps

### Tier 2: Production - Medium to Large Models
**Use Case:** Production deployment of medium to large models (13B-70B parameters), high throughput

**Hardware Specifications:**
- **CPU:** 32 cores / 64 threads
- **RAM:** 128-256GB system RAM
- **Storage:** 1-2TB NVMe SSD
- **GPU:**1-4 GPUs with 40GB+ VRAM each
  - 1-4x NVIDIA B200 (192GB HBM3e)
  - 1-4x NVIDIA H200 (141GB HBM3e)
  - 1-4x NVIDIA RTX 6000 Pro Blackwell (48GB)
  - 1-2x NVIDIA H100 (80GB)
  - 1-4x NVIDIA A100 (40GB or 80GB)
  - 1-2x NVIDIA L40S (48GB)
  - 2-4x NVIDIA A10G (24GB) for tensor parallelism
- **Network:** 25-40 Gbps

### Tier 3: Enterprise Multi-Node Cluster
**Use Case:** Enterprise deployment with multiple models, high availability, horizontal scaling, 99.9%+ SLA

**Cluster Architecture:**
**Head Node (Control Plane):**
- **CPU:** 16 cores / 32 threads
- **RAM:** 64GB
- **Storage:** 500GB NVMe SSD
- **GPU:** Same class as worker nodes (homogeneous cluster recommended)

**Worker Nodes (3+ nodes for HA):**
- **CPU:** 32-64 cores / 64-128 threads per node
- **RAM:** 256-512GB per node
- **Storage:** 2TB NVMe SSD per node (local cache)
- **GPU:** 4-8 GPUs per node (same class as head node)
- **Network:** 40-100 Gbps (InfiniBand for HPC workloads)

## Cloud Provider Instance Mapping
### AWS EC2 Instance Types
| Tier | Instance Type | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `m6i.2xlarge` | 8 | 32GB | None | 200GB gp3 |
| **Tier 1: With GPU** | `g5.xlarge` | 4 | 16GB | 1x A10G (24GB) | 200GB gp3 |
| **Tier 2: Multi-GPU** | `g5.12xlarge` | 48 | 192GB | 4x A10G (96GB) | 2TB gp3 |

### Google Cloud Platform (GCP) Instance Types
| Tier | Machine Type | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `n2-standard-8` | 8 | 32GB | None | 200GB SSD |
| **Tier 2: Multi-GPU** | `a2-highgpu-4g` | 48 | 340GB | 4x A100 (160GB) | 2TB SSD |

### Microsoft Azure Instance Types
| Tier | VM Size | vCPU | RAM | GPU | Storage |
| --- | --- | --- | --- | --- | --- |
| **Tier 1: CPU-only** | `Standard_D8s_v5` | 8 | 32GB | None | 200GB Premium SSD |
| **Tier 2: A100 Multi-GPU** | `Standard_NC96ads_A100_v4` | 96 | 880GB | 4x A100 (320GB) | 2TB Premium SSD |

## Network Configuration
### Network Bandwidth Requirements
#### Single Node Deployment
- **Minimum:** 1 Gbps (for model downloads, API traffic)
- **Recommended:** 10 Gbps (for high-throughput inference)

#### Multi-Node Cluster
- **Minimum:** 10 Gbps Ethernet
- **Recommended:** 25-40 Gbps Ethernet or InfiniBand

### Network Ports
#### Linux/macOS Enterprise Edition
- 443/tcp: HTTPS primary access

#### Windows Edition
- 443/tcp: HTTPS primary access (via WSL)

### Required Kernel Modules (Enterprise Edition Linux Only)
- overlay
- br_netfilter

### System Network Parameters (Enterprise Edition Linux Only)
```bash
# Required sysctl settings for Swarm networking
net.bridge.bridge-nf-call-iptables = 1
net.bridge.bridge-nf-call-ip6tables = 1
net.ipv4.ip_forward = 1
```

## Directory Structure
### Enterprise Edition
```text
/etc/kamiwaza/
├── config/
├── ssl/      # Cluster certificates
└── swarm/    # Swarm tokens
/opt/kamiwaza/
├── containers/  # Docker root (configurable)
├── logs/
└── runtime/    # Runtime files
```
### Community Edition
```text
$KAMIWAZA_ROOT/
├── env.sh
├── runtime/
└── logs/
```

## Special Considerations
### Apple Silicon (M-Series)
**MLX Engine Support:**
- Kamiwaza supports Apple Silicon via the MLX inference engine
- Unified memory architecture (shared CPU/GPU RAM)

### NVIDIA DGX Spark
The NVIDIA DGX Spark is a compact AI workstation powered by the GB10 Grace Blackwell Superchip:
- **CPU:** 20-core ARM

### AMD Ryzen AI Max+ 395 "Strix Halo"
AMD's Strix Halo platform provides powerful AI inference in a compact form factor:
- **CPU:** 16-core Zen 5 (up to 5.1 GHz)

## Shared Storage (Multi-Node Clusters)
**Network Filesystem Requirements:**
- **Protocol:** NFSv4, Lustre, CephFS, or S3-compatible object storage
- **Network Bandwidth:** 10 Gbps minimum, 40+ Gbps for production

### Storage Configuration by Edition
#### Enterprise Edition Requirements
- Primary mountpoint for persistent storage (/opt/kamiwaza)
- Shared storage for multi-node clusters.

### Community Edition
- Local filesystem storage.

## Version Compatibility
- Docker Engine: 24.0 or later with Compose 2.23+
- NVIDIA Driver: 450.80.02 or later
