LLMKube

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.

Beliebtheit
★ 198
Kategorien
Generative Artificial Intelligence (GenAI)
Plattform
Go, Docker, K8S
Lizenz
Apache-2.0
Letzte Version
v0.9.20 (2026-08-24)
Status
Aktiv

Ähnliche Apps