# Hardware Requirements

## CPU
- **Minimum Cores**: 8+ cores
- **Recommended Cores**: 16+ cores for CPU-based inference workloads
- **Architecture**:
  - Linux: x64/amd64 (64-bit)
  - Windows: x64 (64-bit)
  - macOS: ARM64 (Apple Silicon) only

## Memory

### System RAM

| Mode         | Minimum | Recommended | Notes                                                        |
|--------------|---------|-------------|--------------------------------------------------------------|
| **Lite Mode**| 16GB    | 32GB        | SQLite database; limited capacity for apps/tools            |
| **Full Mode**| 32GB    | 64GB+       | CockroachDB + DataHub; production workloads                  |
| **GPU Workloads**| 32GB| 64GB+       | System RAM alongside GPU vRAM                                 |

### GPU Memory (vRAM)
- **GPU Inference**: 16GB+ vRAM required
- **Recommended**: 32GB+ vRAM for optimal GPU inference performance

### GPU (Optional)
Kamiwaza supports multiple GPU and accelerator platforms:

**Discrete GPUs:**
- NVIDIA GPUs with compute capability 7.0+ (Linux)
- NVIDIA RTX / Intel Arc (Windows via WSL)

**Unified Memory Systems:**
- **NVIDIA DGX Spark** \- GB10 Grace Blackwell, 128GB unified memory
- **AMD Ryzen AI Max+ 395** \- "Strix Halo" platform, up to 128GB unified memory
- **Apple Silicon M-series** \- Unified memory architecture (macOS only)

### Storage
Storage requirements are the same across all platforms.

#### Storage Performance
- **Required**: SSD (Solid State Drive)
- **Preferred**: NVMe SSD for optimal performance
- **Minimum**: SATA SSD
- **Note**: Model weights can be on a separate HDD but load times will increase significantly

#### Storage Capacity
- **Minimum**: 100GB free disk space
- **Recommended**: 200GB+ free disk space
- **Enterprise Edition**: Additional space for /opt/kamiwaza persistence

#### Capacity Planning

| Component               | Minimum | Recommended | Notes                                               |
|-------------------------|---------|-------------|-----------------------------------------------------|
| **Operating System**    | 20GB    | 50GB        | Ubuntu/RHEL base + dependencies                      |
| **Kamiwaza Platform**   | 50GB    | 50GB        | Python environment, Ray, services                     |
| **Model Storage**       | 50GB    | 500GB+      | Depends on number and size of models                 |
| **Database**            | 10GB    | 50GB        | CockroachDB for metadata                              |
| **Vector Database**     | 10GB    | 100GB+      | For embeddings (if enabled)                          |
| **Logs & Metrics**      | 10GB    | 50GB        | Rotated logs, Ray dashboard data                      |
| **Scratch Space**       | 20GB    | 100GB       | Temporary files, downloads, builds                    |
| **Total**               | **170GB**| **900GB+**  |                                                     |

#### Storage Performance Requirements
**Local Storage (Single Node):**
- **Minimum:** SATA SSD (500 MB/s sequential read)
- **Recommended:** NVMe SSD (2000+ MB/s sequential read)
- **Note:** HDD is only recommended for non-dynamic model loads and low KV cache usage - model load times can be very long (15+ minutes); models are in memory after load

**Performance Targets:**
- **Sequential Read:** 2000+ MB/s (model loading)
- **Sequential Write:** 1000+ MB/s (model downloads, checkpoints)
- **4K Random Read IOPS:** 50,000+ (database, concurrent access)
- **4K Random Write IOPS:** 20,000+ (database writes, logs)

### Supported Operating Systems

#### Linux
- **Ubuntu**: 24.04 and 22.04 LTS via .deb package installation (x64/amd64 architecture only)
- **Red Hat Enterprise Linux (RHEL)**: 9

#### Windows
- **Windows 11** (x64 architecture) via WSL with MSI installer
- Requires Windows Subsystem for Linux (WSL) installed and enabled
- Administrator access required for initial setup
- Windows Terminal recommended for optimal WSL experience

#### macOS
- **macOS 15.0 (Sequoia) or later**, Apple Silicon (ARM64) only
- Community edition only
- Single-node deployments only (Enterprise edition not available on macOS)

## Software Dependencies

### Pre-requisites (User Must Install)
Before running the Kamiwaza installer, ensure the following are installed:

| Component | Requirement                           | Installation Guide                           |
|-----------|---------------------------------------|----------------------------------------------|
| **Docker**| Docker Engine 24.0+ with Compose 2.23+| [Docker Install Guide](https://docs.docker.com/engine/install/) |
| **Browser**| Chrome 141+ (tested and recommended)| [Download Chrome](https://www.google.com/chrome/) |

### GPU Drivers (Required for GPU Inference)
Install the appropriate driver for your GPU hardware:

**NVIDIA GPUs:**

| Component                   | Requirement           | Installation Guide                         |
|-----------------------------|-----------------------|--------------------------------------------|
| NVIDIA Driver               | 550-server or later   | [NVIDIA Driver Downloads](https://www.nvidia.com/download/index.aspx) |
| NVIDIA Container Toolkit     | Required for GPU containers | [Container Toolkit Install](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) |

**AMD GPUs (ROCm):**

| Component     | Requirement              | Installation Guide                          |
|---------------|--------------------------|---------------------------------------------|
| ROCm          | 7.1.1+ (see note for gfx1151) | [ROCm Installation](https://rocm.docs.amd.com/en/latest/deploy/linux/index.html) |
| Docker ROCm support | `--device /dev/kfd --device /dev/dri` | [ROCm Docker Guide](https://rocm.docs.amd.com/en/latest/how-to/docker.html) |

### System Resources

```bash
# Check available memory
free -h

# Check CPU cores
nproc

# Check available disk space
df -h /
```

## Hardware Recommendation Tiers
Kamiwaza is a distributed AI platform built on Ray that supports both CPU-only and GPU-accelerated inference. Hardware requirements vary significantly based on:
- **Model size**: From 0.6B to 70B+ parameters
- **Deployment scale**: Single-node development vs multi-node production
- **Inference engine**: LlamaCpp (CPU/GPU), VLLM (GPU), MLX (Apple Silicon)
- **Workload type**: Interactive chat, batch processing, RAG pipelines

### Tier 1: Development & Small Models
**Workload Capacity:**
- Low-volume workloads: 1-10 concurrent requests (supports dozens of interactive users)
- Development, testing, and proof-of-concept deployments

### Tier 2: Production - Medium to Large Models
**Use Case:** Production deployment of medium to large models (13B-70B parameters), high throughput

### Tier 3: Enterprise Multi-Node Cluster
**Use Case:** Enterprise deployment with multiple models, high availability, horizontal scaling, 99.9%+ SLA

## Cloud Provider Instance Mapping
### AWS EC2 Instance Types

| Tier                      | Instance Type       | vCPU | RAM   | GPU                         | Storage  |
|---------------------------|---------------------|------|-------|-----------------------------|----------|
| **Tier 1: CPU-only**      | `m6i.2xlarge`       | 8    | 32GB  | None                        | 200GB gp3 |
| **Tier 2: Multi-GPU**    | `g5.12xlarge`       | 48   | 192GB | 4x A10G (96GB)             | 2TB gp3 |
| **Tier 3: All Nodes**    | `p4d.24xlarge`      | 96   | 1152GB| 8x A100 (320GB)            | 2TB gp3 |

### Microsoft Azure Instance Types

| Tier                      | VM Size              | vCPU | RAM  | GPU                         | Storage  |
|---------------------------|---------------------|------|------|-----------------------------|----------|
| **Tier 1: CPU-only**      | `Standard_D8s_v5`   | 8    | 32GB | None                        | 200GB Premium SSD |
| **Tier 2: H100 (recommended)** | `Standard_NC40ads_H100_v5` | 40 | 320GB | 1x H100 (80GB)          | 2TB Premium SSD |
| **Tier 3: H100 (recommended)** | `Standard_ND96isr_H100_v5` | 96 | 1900GB| 8x H100 (640GB)        | 2TB Premium SSD |

## Important Notes
- **System Impact**: Network and kernel configurations can affect other services
- **Security**: Certificate generation and management for cluster communications
- **GPU Support**: Available on Linux (NVIDIA GPUs) and Windows (NVIDIA RTX, Intel Arc via WSL)
- **Storage**: Enterprise Edition requires specific storage configuration
- **Network**: Enterprise Edition requires specific network ports for cluster communication
- **Docker**: Custom Docker root configuration may affect other containers
- **Windows Edition**: Requires WSL 2 and will create a dedicated Ubuntu 24.04 instance
- **Administrator Access**: Windows installation requires administrator privileges for initial setup
