llama.cpp
llama.cpp
Local and server inference runtime focused on GGUF and efficient inference across CPU and accelerator backends.
| ID | llama-cpp |
|---|---|
| Versioning | build-commit |
| Project | https://github.com/ggml-org/llama.cpp |
| Docs | https://github.com/ggml-org/llama.cpp |
Deep tutorials stay on glukhov.org. This page is operational reference only.