OZ
oceanz.site
HomeArticlesCheatsheetToolsExploreAbout
🇨🇳Chinese Simplified🇺🇸English
Cheatsheetllama.cpp

llama.cpp

llama.cpp is an open-source LLM inference project built on top of ggml, using GGUF as its primary model format, enabling quantized large language models to run locally on ordinary personal computers and consumer-grade hardware using CPU, GPU, and system memory.

1 items
llama.cppGitHubggmlggml-org at Hugging Face

Overview

Platform characteristics, GGUF model sources, and installation for llama.cpp, plus where llama cli, llama serve, llama-completion, and GBNF Grammar fit and when to use them.

  • llama.cpp has good support for Apple Silicon, leveraging ARM NEON, Accelerate, and Metal for CPU vector computation, math operations, and GPU acceleration respectively.
  • Day-to-day use mainly centers on llama cli and llama serve: the former for running and testing models directly in the terminal, the latter for providing a Web UI, HTTP API, and long-running inference service.
  • Beyond the model collection curated by the llama.cpp team, a huge number of GGUF models are available on Hugging Face, and you can pick a quantization that fits your available hardware memory.
About·
Contact·
Privacy Policy·
Usage & copyright

Explore curiosity. Collect ideas. Share the value.

Created with a blend of OceanZ insight and AI assistance, including CursorChatGPTChatGPTGeminiGrokClaudeCodexCodexClaude CodeClaude CodeDeepSeek etc.

Powered by curiosity. © 2026 oceanz.site. All rights reserved.