Termux-LlamaCpp
Production-Grade GGUF LLM Runtime Utilizing Device Resources, Model Manager & OpenAI Server for Android ARM64
Install the official package directly into your runtime:
pip install termux-llamacpp && termux-llama install
# or Node.js:
npm install -g termux-llamacpp && termux-llama install
The Engineering Challenge
Running local LLMs on mobile Android typically requires multi-gigabyte compiler toolchains (Clang, CMake, Ninja), 20+ minute compilation times, fragile Bionic linker dependencies, and severe memory thrashing under default mmap allocations.
The Architectural Breakthrough
Ships verified Android ARM64 native binaries with bundled shared libraries and device resource optimization, enabling instant zero-compilation local inference and a robust OpenAI-compatible REST/SSE server.
Key Capabilities & Built-in Hardening
Zero-Compilation Instant Deployment
Installs verified Android Bionic ARM64 binaries and bundled shared libraries via cryptographic SHA-256 checks in under 3 seconds without local Clang/CMake.
Device Resource Optimization & Big-Core Tuning
Optimizes device resources with automated octa-core big.LITTLE cluster thread pinning (-t 4) and adaptive runtime routing.
Strict 2-Tier Execution Mode
Supports --device auto (SoC detection and deterministic routing) and --device cpu (zero-overhead direct forward pass).
OpenAI REST & SSE Streaming Supervisor
Built-in reverse proxy supervisor exposing /health, /v1/models, and /v1/chat/completions with real-time SSE streaming and loopback CORS isolation.
Automated Process Lifecycle & Cleanup
Guarantees 100% process termination on shutdown or error, with bounded health check polling and zero leftover orphaned daemon processes.
Supply-Chain Cryptographic Integrity
Enforces Ed25519 signed manifests, anti-downgrade policies, symlink traversal blocking, and local build receipts.
Supported Compute Kernels & Operations
| Subsystem Category | Operations & Kernels | Status |
|---|---|---|
| Inference Engine | llama.cpp (ARM64 NEON DotProd & Device Resource Optimization) | Production |
| Architecture Support | Android aarch64 / arm64-v8a (Bionic libc, Qualcomm Adreno, ARM Mali, Samsung Xclipse) | Production |
| Supported Quantizations | Q4_K_M, Q5_K_M, Q8_0, IQ4_XS, Q2_K, F16 GGUF v3 | Production |
| Protocols & APIs | OpenAI Chat Completions (JSON & SSE Stream), Models List, Health Check | Production |
| SDK Bindings | Python 3.8+ (termux-llamacpp), Node.js / TypeScript 16+ (termux-llamacpp) | Production |
Canonical Usage Example
from termux_llamacpp import LlamaRuntime, RuntimeConfig
# 1. Initialize Runtime with Vulkan GPU Acceleration
config = RuntimeConfig(
model_path="models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
device="auto",
threads=4,
context_size=2048
)
runtime = LlamaRuntime(config)
# 2. Synchronous or Streaming Text Generation
output = runtime.generate("한국의 사계절 중 가을의 매력에 대해 설명해줘.", max_tokens=256)
print(output.text)
print(f"Speed: {output.metrics.eval_tokens_per_sec:.2f} t/s (Prompt: {output.metrics.prompt_tokens_per_sec:.2f} t/s)")