LLMKube

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.

Healthy · 74
74/100
Healthy

Actively maintained with regular releases.

Maintenance 31/40
Releases 25/25
Community 2/20
Momentum 15/15

Commit activity more than 78% of tracked projects; popularity more than 12%. How this is calculated.

Commits, last 11 months

2025-10: 0 commits 2025-11: 0 commits 2025-12: 0 commits 2026-01: 0 commits 2026-02: 40 commits 2026-03: 42 commits 2026-04: 69 commits 2026-05: 120 commits 2026-06: 196 commits 2026-07: 250 commits 2026-08: 192 commits
2025-10 909 total 2026-08
Stars217
Latest releasev0.9.30 · 25 Sept 2026
Repo updated27 Sept 2026
LicenceApache-2.0
Built withGo, Docker, K8S

Is LLMKube actively maintained?

LLMKube is under active development: 909 commits landed over the last 11 months, averaging roughly 213 commits a month in the most recent quarter.

Activity is accelerating: the last quarter carried 176% more commits than the one before it.

The most recent tagged release, v0.9.30, shipped within the last month — a current, installable version exists today.

LLMKube ranks #11 of 14 tracked generative artificial intelligence (genai) projects. Several better-maintained options exist in the same category — they are listed below.

What LLMKube actually does

LLMKube is a Kubernetes operator designed to manage self-hosted large language model inference workloads directly on cluster infrastructure. It provides support for multiple backend runtimes, device accelerators including NVIDIA CUDA and Apple Silicon Metal, and exposes an OpenAI-compatible API.

Best fit: Platform engineers running Kubernetes who need to orchestrate and scale self-hosted language model inference across heterogeneous hardware.

Worth knowing: Early-stage projects with modest community adoption carry inherent risks regarding API stability, long-term maintenance, and rapid upstream changes.

Deployment notes

As a Kubernetes operator built with Go, it typically requires cluster-admin privileges for custom resource definitions and RBAC configurations, alongside persistent storage for model weights if caching locally.


Alternatives to LLMKube

Projects in the same categories, ordered by health score.