Llama.cpp is an open-source LLM inference project built on top of ggml, implemented in C/C++, and using GGUF as its primary model format. It mostly runs quantized models, using CPU, GPU, and system memory to run locally on ordinary personal computers and consumer-grade hardware.
Thanks to deep optimization around technologies such as ARM NEON, Accelerate, and Metal, it has good support for Apple Silicon and is especially well suited to Apple users.
| Name | What it is | Roughly responsible for in llama.cpp |
|---|---|---|
| ARM NEON | SIMD instruction set extension for ARM CPUs | CPU vector computation acceleration |
| Apple Accelerate | Apple's high-performance math computation framework | CPU math / matrix operation acceleration |
| Metal | Apple's GPU graphics and general-purpose compute API | GPU acceleration |
ARM NEON:
-
A SIMD (Single Instruction, Multiple Data) vector instruction technology provided by the ARM architecture;
-
An ordinary CPU instruction may process one number at a time, whereas SIMD lets a single instruction process a group of data at once. Matrix and vector operations, which are abundant in LLM inference, are very well suited to this approach;
-
ARM NEON = one of the vector computation capabilities at the heart of Apple Silicon CPUs;
Apple Accelerate:
-
Leverages the CPU's vector processing capability to provide high-performance, low-energy computation. It includes: BLAS / LAPACK, vDSP, vForce, BNNS, vImage, Sparse Solvers, and more;
-
Accelerate = Apple's official high-performance CPU math computation framework;
Metal analogy:
-
Apple Silicon: Metal β Apple GPU
-
NVIDIA: CUDA β NVIDIA GPU
NEON + Accelerate explain why the CPU can perform well; Metal handles how the GPU participates in acceleration.
llama.cpp supports macOS, Linux, Windows, iOS, and Android platforms, and provides both command-line and graphical interface tools.
Model Collection
Llama Models curates a selection of open-source models, all of which can run directly in llama.cpp, and lists the memory size each one requires, making it easy to choose a model that fits your hardware.
In addition, you can find many more models on Hugging Face (as of now, the number of GGUF models is 208,750) and run inference with llama.cpp.
When you open a Text Generation model repository page, for example ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF, click "Use this model" on the right, then select llama.cpp in the panel that appears. You'll see installation and run commands for different platforms β just follow them.
Installation
For Linux Server, the One-line install command is recommended, as shown below. This command automatically detects the platform architecture, fetches the latest binary, and installs it.
curl -LsSf https://llama.app/install.sh | sh
# Output similar to
Version: b11200
Probing CUDA...
Downloading cuda-probe...
Found: 89
Downloading llama...
Installation completed successfully
To make llama available in future sessions, add ~/.local/bin to your PATH by running:
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bash_profile
Then open a new terminal and run:
llama serve
To start it now without modifying your PATH, run:
~/.llama-app/llama serveAfter installation, check the version:
llama cli --version
# Output similar to
version: 0.5.0-dev (build 11200, commit 81bc6b83f)
built with Clang 19.1.7 for Linux x86_64On macOS, you can also install via the command line, but the graphical interface is recommended: download the dmg file and install it.
For platforms such as Windows and Android, see the releases page and download the corresponding zip file.
Tools
llama.cpp provides tools for model inference and output control. In day-to-day use:
-
Focus mainly on CLI and Server;
-
llama-completionand GBNF Grammar still exist and are maintained, but are used more for raw text generation, special-model compatibility, or lower-level generation control;
llama cli
llama cli is llama.cpp's current primary local command-line interaction entry point. It's suitable for loading a GGUF model directly and doing inference and chat in the terminal.
llama cli -m <model-path.gguf>For multimodal models, load the corresponding multimodal projector:
llama cli -m <model-path.gguf> --mmproj <mmproj-path.gguf>Good for:
-
Quickly testing models locally;
-
Interacting with a model directly in an SSH environment;
-
Testing whether a GGUF model loads correctly;
-
Tuning inference parameters such as context, GPU offload, and sampling;
-
Testing multimodal inputs such as images;
llama cli is one of the most direct interactive entry points for using llama.cpp today.
llama serve
llama serve starts a model as a long-running inference service, providing both an HTTP API and a Web UI.
llama serve -m <model-path.gguf>For multimodal models, load the corresponding multimodal projector:
llama serve -m <model-path.gguf> --mmproj <mmproj-path.gguf>Compared with llama cli, it no longer just interacts with the model directly in the current terminal session; instead, it turns the model into a service that other programs can call.
Typical uses:
-
Testing a model through the browser Web UI;
-
OpenAI-compatible API;
-
Calling from Python / JavaScript clients;
-
Integrating into your own FastAPI, Agent, or other application;
-
Serving the model over a LAN or remotely;
If the service is exposed to an untrusted network, you need to configure an API key or other access control β never leave the inference port exposed directly.
llama cli is for "a person using a model directly"; llama serve is for "exposing a model as a service."
Full llama serve documentation
llama-completion
Like cli, it can also load GGUF models for inference, but its level of abstraction and intended use differ. It doesn't need to be a focus of day-to-day use, but it can help you understand llama.cpp's lower-level generation mechanisms. It's a tool oriented toward raw Prompt β Completion text generation, for example:
llama-completion \
-m model.gguf \
-p "Once upon a time" \
-n 128 \
-no-cnvIt retains more capabilities for directly controlling the text interaction process, such as:
reverse prompt
input prefix
input suffix
interactive mode
chat template
sampling parametersAmong these, reverse prompt and input prefix / suffix are usage patterns from the early days of LLMs.
For example, early models simulated conversations through plain-text structures like this:
User: Hello
Assistant:input prefix / suffix can help a program automatically assemble this text; reverse prompt can stop generation when the model reaches a specific piece of text, such as User:, and hand control back to the user.
Modern Chat / Instruct models are usually already managed through a Chat Template:
System
User
Assistantwhich handles the message format and conversation boundaries between them, so ordinary users rarely need to manage this manually anymore.
But llama-completion is still valuable, for example for:
-
Raw text continuation for base models;
-
Models without a standard Chat Template;
-
Custom Prompt formats;
-
Debugging Chat Templates;
-
Studying a model's raw completion behavior;
-
Special fine-tunes / legacy model compatibility;
llama-completionis llama.cpp's retained low-level-oriented text generation tool; for everyday interaction with modern Chat / Instruct models,llama cliis usually preferred.
llama-completion documentation
GBNF Grammar
Strictly speaking, GBNF is not a standalone tool, but rather llama.cpp's Grammar format and constrained-generation mechanism.
GBNF: GGML BNF, based on the idea of BNF (BackusβNaur Form), used to describe what kind of text structure the model is allowed to generate. For example:
root ::= "yes" | "no"During generation, the model can then only produce content that conforms to that Grammar. The general flow is:
Model predicts Token
β
Grammar determines which Tokens are legal
β
Restrict Sampling candidates
β
Select a legal Token
β
Continue generatingThis falls under: Constrained Decoding / Constrained Generation.
GBNF can be used to constrain:
-
JSON;
-
Fixed text formats;
-
DSLs;
-
Specific command syntax;
-
Structured languages such as SQL;
-
Custom protocols or output formats;
Modern application development usually doesn't require writing GBNF by hand. For example, if a model needs to output:
{
"name": "Ocean",
"age": 18
}The application layer typically defines a JSON Schema, then constrains output through llama.cpp's Structured Output capability.
GBNF has gradually sunk from an interface ordinary users had to touch directly into a low-level capability within llama.cpp's constrained-generation system; modern applications typically use this kind of capability through higher-level interfaces such as JSON Schema / Structured Output.
If you need to describe a special language, DSL, or precise grammar beyond JSON, writing a Grammar directly is still worthwhile.
Relationship
| Name | Role | Current usage frequency |
|---|---|---|
llama cli | Local interactive inference | Common |
llama serve | Web UI + HTTP API service | Common |
llama-completion | Raw Prompt / Completion and low-level interaction control | Rarely used directly |
| GBNF Grammar | Grammar mechanism that constrains model output | Rarely hand-written, still important at a low level |
Everyday use
βββ llama cli
β βββ Interact with the model directly
β
βββ llama serve
βββ Turn the model into a service
Going deeper
βββ llama-completion
β βββ Prompt β Completion / raw text interaction control
β
βββ GBNF Grammar
βββ Constrained Decoding / output structure constraints