LLMKube
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Actively maintained with regular releases.
Commit activity more than 78% of tracked projects; popularity more than 12%. How this is calculated.
Commits, last 11 months
| Stars | 217 |
|---|---|
| Latest release | v0.9.30 · 25 Sept 2026 |
| Repo updated | 27 Sept 2026 |
| Licence | Apache-2.0 |
| Built with | Go, Docker, K8S |
Is LLMKube actively maintained?
LLMKube is under active development: 909 commits landed over the last 11 months, averaging roughly 213 commits a month in the most recent quarter.
Activity is accelerating: the last quarter carried 176% more commits than the one before it.
The most recent tagged release, v0.9.30, shipped within the last month — a current, installable version exists today.
LLMKube ranks #11 of 14 tracked generative artificial intelligence (genai) projects. Several better-maintained options exist in the same category — they are listed below.
What LLMKube actually does
LLMKube is a Kubernetes operator designed to manage self-hosted large language model inference workloads directly on cluster infrastructure. It provides support for multiple backend runtimes, device accelerators including NVIDIA CUDA and Apple Silicon Metal, and exposes an OpenAI-compatible API.
Best fit: Platform engineers running Kubernetes who need to orchestrate and scale self-hosted language model inference across heterogeneous hardware.
Worth knowing: Early-stage projects with modest community adoption carry inherent risks regarding API stability, long-term maintenance, and rapid upstream changes.
Deployment notes
As a Kubernetes operator built with Go, it typically requires cluster-admin privileges for custom resource definitions and RBAC configurations, alongside persistent storage for model weights if caching locally.
Links
Alternatives to LLMKube
Projects in the same categories, ordered by health score.
Modern design AI chat framework supporting multiple AI providers, one click install MCP Marketplace and Artifacts / Thinking.
User-friendly AI Interface, supports Ollama, OpenAI API.
Run your AI models locally and generate images and audio (alternative to OpenAI and Claude).
Enhanced ChatGPT-compatible AI chat interface supporting multiple AI providers, with multi-user auth, message search, and plugin support.
Chat UI that works with any LLM. It comes loaded with advanced features like agents, web search, RAG, MCP, deep research, Connectors to 40+ knowledge sources, and more.
LLMOps platform for prompt management, LLM evaluation, and observability. Build, evaluate, and monitor production-grade LLM applications with collaborative prompt engineering.