System Requirements | Kamiwaza Docs
Base System Requirements
Supported Operating Systems & Architecture
- Linux:
- Ubuntu 24.04 and 22.04 LTS via .deb package installation (x64/amd64 architecture only)
- Redhat Enterprise Linux (RHEL) 9
- Windows: 11 (x64 architecture) via WSL with MSI installer
- macOS: 12.0 or later, Apple Silicon (ARM64) only (community edition only)
CPU Requirements
- Architecture:
- Linux: x64/amd64 (64-bit)
- macOS: ARM64 (Apple Silicon) only
- Windows: x64 (64-bit)
- Minimum Cores: 8+ cores
- Recommended Cores: 16+ cores for CPU-based inference workloads
Core Software Requirements
- Python: Python 3.10 for tarball installations; Python 3.12 for .deb/.msi installations
- Docker: Docker Engine with Compose v2
- Node.js: 22.x (installed via NVM during setup)
- Browser: Chrome Version 141+ (tested and recommended)
- GPU Support: NVIDIA GPU with compute capability 7.0+ (Linux only) or NVIDIA RTX/Intel Arc (Windows via WSL)
Memory Requirements
System RAM
- Minimum: 16GB RAM
- Recommended: 32GB+ RAM for CPU-based inference workloads
- GPU Workloads: 16GB+ system RAM (32GB+ recommended)
GPU Memory (vRAM)
- GPU Inference: 16GB+ vRAM required
- Recommended: 32GB+ vRAM for optimal GPU inference performance
Windows (WSL-based) Specific
- Minimum: 16GB RAM
- Recommended: 32GB+ RAM
- Memory Allocation: 50-75% of system RAM dedicated to Kamiwaza during installation
Storage Requirements
Storage Performance
- Required: SSD (Solid State Drive)
- Preferred: NVMe SSD for optimal performance
- Minimum: SATA SSD
Storage Capacity
Linux/macOS
- Minimum: 100GB free disk space
- Recommended: 200GB+ free disk space
Windows
- Minimum: 100GB free disk space
- Recommended: 200GB+ free space on SSD
Hardware Recommendation Tiers
Kamiwaza is a distributed AI platform built on Ray that supports both CPU-only and GPU-accelerated inference. Hardware requirements vary significantly based on:
- Model size: From 0.6B to 70B+ parameters
- Deployment scale: Single-node development vs multi-node production
- Inference engine: LlamaCpp (CPU/GPU), VLLM (GPU), MLX (Apple Silicon)
- Workload type: Interactive chat, batch processing, RAG pipelines
GPU Memory Requirements by Model Size
The table below provides real-world GPU memory requirement estimates for representative models at different scales. These estimates assume FP8 and include overhead for context windows and batch processing.
| Model Example | Parameters | Minimum vRAM | Notes |
|---|---|---|---|
| GPT-OSS 20B | 20B | 24GB | Includes weights + 1-batch max context; fits 1x 24GB GPU (e.g., L4/RTX 4090) |
| GPT-OSS 120B | 120B | 80GB | ~40GB weights + 1-batch max context; 1x H100/H200 or 2x A100 80GB recommended |
| Qwen 3 235B A22B | 235B | 150GB | ~120GB weights + 1-batch max context; 2x H200 (282GB) or 2x B200 (384GB) ideal for max context |
| Qwen 3-VL 235B A22B | 235B | 150GB | Same base minimum (includes 1-batch max context); budget +20-30% vRAM for high-res vision inputs |
Tier 1: Development & Small Models
Use Case: Local development, testing, small to medium model deployment (up to 13B parameters)
Hardware Specifications:
- CPU: 8-16 cores / 16-32 threads
- RAM: 32GB (16GB minimum)
- Storage: 200GB NVMe SSD (100GB minimum)
- GPU: Optional - Single GPU with 16-24GB VRAM
- NVIDIA RTX 4090 (24GB)
- NVIDIA RTX 4080 (16GB)
- NVIDIA T4 (16GB)
Tier 2: Production - Medium to Large Models
Use Case: Production deployment of medium to large models (13B-70B parameters), high throughput
Hardware Specifications:
- CPU: 32 cores / 64 threads
- RAM: 128-256GB system RAM
- Storage: 1-2TB NVMe SSD
- GPU: 1-4 GPUs with 40GB+ VRAM each
Tier 3: Enterprise Multi-Node Cluster
Use Case: Enterprise deployment with multiple models, high availability, horizontal scaling, 99.9%+ SLA
Cluster Architecture: Head Node (Control Plane):
- CPU: 16 cores / 32 threads
- RAM: 64GB
- Storage: 500GB NVMe SSD
Worker Nodes (3+ nodes for HA):
- CPU: 32-64 cores / 64-128 threads per node
- RAM: 256-512GB per node
- Storage: 2TB NVMe SSD per node (local cache)
Shared Storage:
- High-performance NAS or distributed filesystem (Lustre, CephFS)
Cloud Provider Instance Mapping
AWS EC2 Instance Types
| Tier | Instance Type | vCPU | RAM | GPU | Storage |
|---|---|---|---|---|---|
| Tier 1: CPU-only | m6i.2xlarge |
8 | 32GB | None | 200GB gp3 |
| Tier 1: With GPU | g5.xlarge |
4 | 16GB | 1x A10G (24GB) | 200GB gp3 |
| Tier 1: Alternative | g5.2xlarge |
8 | 32GB | 1x A10G (24GB) | 200GB gp3 |
| Tier 2: Multi-GPU | g5.12xlarge |
48 | 192GB | 4x A10G (96GB) | 2TB gp3 |
| Tier 2: Alternative | p4d.24xlarge |
96 | 1152GB | 8x A100 (320GB) | 2TB gp3 |
| Tier 3: All Nodes | p4d.24xlarge |
96 | 1152GB | 8x A100 (320GB) | 2TB gp3 |
Google Cloud Platform (GCP) Instance Types
| Tier | Machine Type | vCPU | RAM | GPU | Storage |
|---|---|---|---|---|---|
| Tier 1: CPU-only | n2-standard-8 |
8 | 32GB | None | 200GB SSD |
| Tier 1: With GPU | n1-standard-8 + 1x T4 |
8 | 30GB | 1x T4 (16GB) | 200GB SSD |
| Tier 2: Multi-GPU | a2-highgpu-4g |
48 | 340GB | 4x A100 (160GB) | 2TB SSD |
Microsoft Azure Instance Types
| Tier | VM Size | vCPU | RAM | GPU | Storage |
|---|---|---|---|---|---|
| Tier 1: CPU-only | Standard_D8s_v5 |
8 | 32GB | None | 200GB Premium SSD |
| Tier 1: With GPU | Standard_NC4as_T4_v3 |
4 | 28GB | 1x T4 (16GB) | 200GB Premium SSD |
| Tier 2: H100 Multi-GPU | Standard_NC80adis_H100_v5 |
80 | 640GB | 2x H100 (160GB) | 2TB Premium SSD |
Windows-Specific Prerequisites
- Windows Subsystem for Linux (WSL) installed and enabled
- Administrator access required for initial setup
- Windows Terminal (recommended for optimal WSL experience)
Dependencies & Components
Required System Packages
See platform-specific installation instructions
NVIDIA Components (Linux GPU Support)
- NVIDIA Driver (550-server recommended)
- NVIDIA Container Toolkit
- nvidia-docker2
Windows Components (Automated via MSI Installer)
- Windows Subsystem for Linux (WSL 2)
- Ubuntu 24.04 LTS (automatically downloaded and configured)
- Docker Engine (configured within WSL)
- GPU drivers and runtime (automatically detected and configured)
- Node.js 22 (via NVM within WSL environment)
Network Configuration
Network Bandwidth Requirements
Single Node Deployment
Network Bandwidth:
- Minimum: 1 Gbps (for model downloads, API traffic)
- Recommended: 10 Gbps (for high-throughput inference)
Multi-Node Cluster
Inter-Node Network:
- Minimum: 10 Gbps Ethernet
- Recommended: 25-40 Gbps Ethernet or InfiniBand
Detailed Storage Requirements
Capacity Planning
| Component | Minimum | Recommended | Notes |
|---|---|---|---|
| Operating System | 20GB | 50GB | Ubuntu/RHEL base + dependencies |
| Kamiwaza Platform | 50GB | 50GB | Python environment, Ray, services |
| Model Storage | 50GB | 500GB+ | Depends on number and size of models |
| Database | 10GB | 50GB | CockroachDB for metadata |
Important Notes
- System Impact: Network and kernel configurations can affect other services
- Security: Certificate generation and management for cluster communications
- Storage: Enterprise Edition requires specific storage configuration
- Network: Enterprise Edition requires specific network ports for cluster communication
- Windows Edition: Requires WSL 2 and will create a dedicated Ubuntu 24.04 instance
Additional Considerations
Network Ports
Linux/macOS Enterprise Edition
- 443/tcp: HTTPS primary access
- 51100-51199/tcp: Deployment ports for model instances