Showing 11 of 11
Intel-Based Self-Hosting
Intel Arc GPU and Xeon CPU setups for OpenVINO, IPEX, and CPU-optimized inference.
Kubernetes Self-Hosting
Run agent workloads, vector databases, and inference APIs on Kubernetes for scale, resilience, and reproducible deployments.
Mac-Based Self-Hosting
Apple Silicon local inference from entry-level MacBook Air to Mac Studio Ultra.
macOS Self-Hosting
Apple Silicon native inference with MLX, llama.cpp Metal backends, and local agent tools.
NixOS Self-Hosting
Reproducible, declarative AI infrastructure with NixOS — ideal for teams that want versioned system configurations and rollback safety.
Single-Board Edge Self-Hosting
Run small models and lightweight agents on Raspberry Pi, Orange Pi, and other edge boards for offline, low-power inference.
Proxmox VE Self-Hosting
Run AI workloads and agent VMs on Proxmox VE. Combine bare-metal performance with VM isolation, GPU passthrough, and easy snapshots.
NVIDIA-Based Self-Hosting
The complete guide to CUDA-accelerated local AI — from DGX Spark and RTX workstations to A100/H100 datacenter clusters.
AMD-Based Self-Hosting
A practical guide to ROCm-accelerated local AI on AMD Radeon and Instinct hardware — from budget gaming GPUs to MI300X datacenter clusters.
Linux Self-Hosting
The definitive OS for local AI. Ubuntu setup, drivers, containers, and the inference engines that power agent workloads.
Windows Self-Hosting
Run local LLMs on Windows with WSL2, native CUDA, and tools like LM Studio and Ollama.