llama.cpp
llama.cpp is an open-source LLM inference project built on top of ggml, using GGUF as its primary model format, enabling quantized large language models to run locally on ordinary personal computers and consumer-grade hardware using CPU, GPU, and system memory.