llama.cpp

Published 2026-10-04

Overview

Platform characteristics, GGUF model sources, and installation for llama.cpp, plus where llama cli, llama serve, llama-completion, and GBNF Grammar fit and when to use them.

Llama.cpp is an open-source LLM inference project built on top of ggml, implemented in C/C++, and using GGUF as its primary model format. It mostly runs quantized models, using CPU, GPU, and system memory to run locally on ordinary personal computers and consumer-grade hardware.

Thanks to deep optimization around technologies such as ARM NEON, Accelerate, and Metal, it has good support for Apple Silicon and is especially well suited to Apple users.

NameWhat it isRoughly responsible for in llama.cpp
ARM NEONSIMD instruction set extension for ARM CPUsCPU vector computation acceleration
Apple AccelerateApple's high-performance math computation frameworkCPU math / matrix operation acceleration
MetalApple's GPU graphics and general-purpose compute APIGPU acceleration

ARM NEON:

  • A SIMD (Single Instruction, Multiple Data) vector instruction technology provided by the ARM architecture;

  • An ordinary CPU instruction may process one number at a time, whereas SIMD lets a single instruction process a group of data at once. Matrix and vector operations, which are abundant in LLM inference, are very well suited to this approach;

  • ARM NEON = one of the vector computation capabilities at the heart of Apple Silicon CPUs;

Apple Accelerate:

  • Leverages the CPU's vector processing capability to provide high-performance, low-energy computation. It includes: BLAS / LAPACK, vDSP, vForce, BNNS, vImage, Sparse Solvers, and more;

  • Accelerate = Apple's official high-performance CPU math computation framework;

Metal analogy:

  • Apple Silicon: Metal β†’ Apple GPU

  • NVIDIA: CUDA β†’ NVIDIA GPU

NEON + Accelerate explain why the CPU can perform well; Metal handles how the GPU participates in acceleration.

llama.cpp supports macOS, Linux, Windows, iOS, and Android platforms, and provides both command-line and graphical interface tools.

Model Collection

Llama Models curates a selection of open-source models, all of which can run directly in llama.cpp, and lists the memory size each one requires, making it easy to choose a model that fits your hardware.

In addition, you can find many more models on Hugging Face (as of now, the number of GGUF models is 208,750) and run inference with llama.cpp.

When you open a Text Generation model repository page, for example ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF, click "Use this model" on the right, then select llama.cpp in the panel that appears. You'll see installation and run commands for different platforms β€” just follow them.

Installation

For Linux Server, the One-line install command is recommended, as shown below. This command automatically detects the platform architecture, fetches the latest binary, and installs it.

Bash
curl -LsSf https://llama.app/install.sh | sh
 
# Output similar to
Version: b11200
Probing CUDA...
Downloading cuda-probe...
Found: 89
Downloading llama...
Installation completed successfully
 
To make llama available in future sessions, add ~/.local/bin to your PATH by running:
 
  echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bash_profile
 
Then open a new terminal and run:
 
  llama serve
 
To start it now without modifying your PATH, run:
 
  ~/.llama-app/llama serve

After installation, check the version:

Bash
llama cli --version
 
# Output similar to
version: 0.5.0-dev (build 11200, commit 81bc6b83f)
built with Clang 19.1.7 for Linux x86_64

On macOS, you can also install via the command line, but the graphical interface is recommended: download the dmg file and install it.

For platforms such as Windows and Android, see the releases page and download the corresponding zip file.

Tools

llama.cpp provides tools for model inference and output control. In day-to-day use:

  • Focus mainly on CLI and Server;

  • llama-completion and GBNF Grammar still exist and are maintained, but are used more for raw text generation, special-model compatibility, or lower-level generation control;

llama cli

llama cli is llama.cpp's current primary local command-line interaction entry point. It's suitable for loading a GGUF model directly and doing inference and chat in the terminal.

Bash
llama cli -m <model-path.gguf>

For multimodal models, load the corresponding multimodal projector:

Bash
llama cli -m <model-path.gguf> --mmproj <mmproj-path.gguf>

Good for:

  • Quickly testing models locally;

  • Interacting with a model directly in an SSH environment;

  • Testing whether a GGUF model loads correctly;

  • Tuning inference parameters such as context, GPU offload, and sampling;

  • Testing multimodal inputs such as images;

llama cli is one of the most direct interactive entry points for using llama.cpp today.

Full llama cli parameters

llama serve

llama serve starts a model as a long-running inference service, providing both an HTTP API and a Web UI.

Bash
llama serve -m <model-path.gguf>

For multimodal models, load the corresponding multimodal projector:

Bash
llama serve -m <model-path.gguf> --mmproj <mmproj-path.gguf>

Compared with llama cli, it no longer just interacts with the model directly in the current terminal session; instead, it turns the model into a service that other programs can call.

Typical uses:

  • Testing a model through the browser Web UI;

  • OpenAI-compatible API;

  • Calling from Python / JavaScript clients;

  • Integrating into your own FastAPI, Agent, or other application;

  • Serving the model over a LAN or remotely;

If the service is exposed to an untrusted network, you need to configure an API key or other access control β€” never leave the inference port exposed directly.

llama cli is for "a person using a model directly"; llama serve is for "exposing a model as a service."

Full llama serve documentation

llama-completion

Like cli, it can also load GGUF models for inference, but its level of abstraction and intended use differ. It doesn't need to be a focus of day-to-day use, but it can help you understand llama.cpp's lower-level generation mechanisms. It's a tool oriented toward raw Prompt β†’ Completion text generation, for example:

Bash
llama-completion \
  -m model.gguf \
  -p "Once upon a time" \
  -n 128 \
  -no-cnv

It retains more capabilities for directly controlling the text interaction process, such as:

Plain text
reverse prompt
input prefix
input suffix
interactive mode
chat template
sampling parameters

Among these, reverse prompt and input prefix / suffix are usage patterns from the early days of LLMs.

For example, early models simulated conversations through plain-text structures like this:

Plain text
User: Hello
Assistant:

input prefix / suffix can help a program automatically assemble this text; reverse prompt can stop generation when the model reaches a specific piece of text, such as User:, and hand control back to the user.

Modern Chat / Instruct models are usually already managed through a Chat Template:

Plain text
System
User
Assistant

which handles the message format and conversation boundaries between them, so ordinary users rarely need to manage this manually anymore.

But llama-completion is still valuable, for example for:

  • Raw text continuation for base models;

  • Models without a standard Chat Template;

  • Custom Prompt formats;

  • Debugging Chat Templates;

  • Studying a model's raw completion behavior;

  • Special fine-tunes / legacy model compatibility;

llama-completion is llama.cpp's retained low-level-oriented text generation tool; for everyday interaction with modern Chat / Instruct models, llama cli is usually preferred.

llama-completion documentation

GBNF Grammar

Strictly speaking, GBNF is not a standalone tool, but rather llama.cpp's Grammar format and constrained-generation mechanism.

GBNF: GGML BNF, based on the idea of BNF (Backus–Naur Form), used to describe what kind of text structure the model is allowed to generate. For example:

Plain text
root ::= "yes" | "no"

During generation, the model can then only produce content that conforms to that Grammar. The general flow is:

Plain text
Model predicts Token
       ↓
Grammar determines which Tokens are legal
       ↓
Restrict Sampling candidates
       ↓
Select a legal Token
       ↓
Continue generating

This falls under: Constrained Decoding / Constrained Generation.

GBNF can be used to constrain:

  • JSON;

  • Fixed text formats;

  • DSLs;

  • Specific command syntax;

  • Structured languages such as SQL;

  • Custom protocols or output formats;

Modern application development usually doesn't require writing GBNF by hand. For example, if a model needs to output:

JSON
{
    "name": "Ocean",
    "age": 18
}

The application layer typically defines a JSON Schema, then constrains output through llama.cpp's Structured Output capability.

GBNF has gradually sunk from an interface ordinary users had to touch directly into a low-level capability within llama.cpp's constrained-generation system; modern applications typically use this kind of capability through higher-level interfaces such as JSON Schema / Structured Output.

If you need to describe a special language, DSL, or precise grammar beyond JSON, writing a Grammar directly is still worthwhile.

Relationship

NameRoleCurrent usage frequency
llama cliLocal interactive inferenceCommon
llama serveWeb UI + HTTP API serviceCommon
llama-completionRaw Prompt / Completion and low-level interaction controlRarely used directly
GBNF GrammarGrammar mechanism that constrains model outputRarely hand-written, still important at a low level
Plain text
Everyday use
β”œβ”€β”€ llama cli
β”‚   └── Interact with the model directly
β”‚
└── llama serve
    └── Turn the model into a service
 
 
Going deeper
β”œβ”€β”€ llama-completion
β”‚   └── Prompt β†’ Completion / raw text interaction control
β”‚
└── GBNF Grammar
    └── Constrained Decoding / output structure constraints