llama.cpp
llama.cpp is an open C/C++ inference project with GGUF, many quantizations, CPU and GPU backends, a server API, and broad hardware support.
Why use it
Game tools and experimental NPCs can run text models offline or privately while controlling deployment cost and data flow.
Where it fits
Find engines, frameworks, code libraries, version control, and production automation.
What to check
Repository licensing does not cover downloaded models; govern model rights, prompt safety, latency, memory, and unreliable output separately.