Skip to content
llama.cpp

llama.cpp

Local and server inference runtime focused on GGUF and efficient inference across CPU and accelerator backends.

IDllama-cpp
Versioningbuild-commit
Projecthttps://github.com/ggml-org/llama.cpp
Docshttps://github.com/ggml-org/llama.cpp

Deep tutorials stay on glukhov.org. This page is operational reference only.