Skip to content
whatismyalternative

LLMKube

A Kubernetes operator for self-hosted LLM inference across multiple runtimes and hardware

Visit LLMKube

LLMKube is an Apache-2.0 open-source Kubernetes operator for deploying and managing self-hosted LLM inference on your own hardware. It addresses the operational complexity of running models for a team—model downloads, GPU scheduling, health checks, autoscaling, and observability—by treating inference as a first-class Kubernetes workload through purpose-built Model and InferenceService CRDs declared in YAML.

The operator supports pluggable runtime backends including vLLM, TGI, llama.cpp, mlx-server, vllm-swift, and a generic custom container option, running on NVIDIA CUDA, Apple Silicon Metal, AMD (Vulkan/ROCm), and Intel (oneAPI/SYCL) hardware. Capabilities include HPA autoscaling driven by inference metrics, Prometheus metrics with Grafana dashboards, an OpenAI-compatible API, a ModelRouter for routing across local and external providers, and the Foreman agentic coding control plane that dispatches coder, verifier, and reviewer agents across a fleet. It also offers OpenShift support, hybrid GPU/CPU offloading, and memory-pressure protection.

LLMKube is aimed at developers and teams running LLM inference on Kubernetes across cloud, on-prem, air-gapped, edge, or Apple Silicon environments. It is packaged as free, open-source software under Apache 2.0, with a Helm chart, CLI, and documentation available alongside its GitHub repository.

12 alternatives to LLMKube

Ranked by how well each tool replaces LLMKube: shared features, audience, price and popularity.

  1. Open-source AI gateway that puts your AI stack behind one OpenAI-compatible key.

    Covers 0 of 15 key features.

    Free planOpen source
    62 out of 100 matchFree
  2. Take control of your ML and AI complexity

    Covers 0 of 15 key features.

    Open source
    61 out of 100 match—
  3. Inference from Kernel to Cloud

    Covers 0 of 15 key features.

    Free planOpen source
    60 out of 100 matchUsage-based
  4. Independent AI evaluations lab

    Covers 0 of 15 key features.

    Free plan
    59 out of 100 matchFree
  5. Make AI run on every machine

    Covers 0 of 15 key features.

    Open source
    58 out of 100 match—
  6. Ultra-fast inference for latency-sensitive agents.

    Covers 0 of 15 key features.

    Free plan
    58 out of 100 matchUsage-based
  7. Cloud platform for running open-source machine learning models via API

    Covers 0 of 15 key features.

    Free planOpen source
    58 out of 100 matchUsage-based
  8. An API for open models, usable locally or in the cloud

    Covers 0 of 15 key features.

    Free planOpen source
    58 out of 100 match$100/mo
  9. Fastest inference for open models

    Covers 0 of 15 key features.

    57 out of 100 match—
  10. LLM evaluation platform for enterprises

    Covers 0 of 15 key features.

    Free plan
    56 out of 100 matchContact sales
  11. Inference infrastructure for AI-native teams

    Covers 0 of 15 key features.

    Free plan
    56 out of 100 match$250/mo
  12. The Open Source AI Platform

    Covers 0 of 15 key features.

    Open source
    55 out of 100 match—