Termux-LlamaCpp

Production-Grade GGUF LLM Runtime Utilizing Device Resources, Model Manager & OpenAI Server for Android ARM64

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install termux-llamacpp && termux-llama install
# or Node.js:
npm install -g termux-llamacpp && termux-llama install

The Engineering Challenge

Running local LLMs on mobile Android typically requires multi-gigabyte compiler toolchains (Clang, CMake, Ninja), 20+ minute compilation times, fragile Bionic linker dependencies, and severe memory thrashing under default mmap allocations.

The Architectural Breakthrough

Ships verified Android ARM64 native binaries with bundled shared libraries and device resource optimization, enabling instant zero-compilation local inference and a robust OpenAI-compatible REST/SSE server.

Key Capabilities & Built-in Hardening

Zero-Compilation Instant Deployment

Installs verified Android Bionic ARM64 binaries and bundled shared libraries via cryptographic SHA-256 checks in under 3 seconds without local Clang/CMake.

Device Resource Optimization & Big-Core Tuning

Optimizes device resources with automated octa-core big.LITTLE cluster thread pinning (-t 4) and adaptive runtime routing.

Strict 2-Tier Execution Mode

Supports --device auto (SoC detection and deterministic routing) and --device cpu (zero-overhead direct forward pass).

OpenAI REST & SSE Streaming Supervisor

Built-in reverse proxy supervisor exposing /health, /v1/models, and /v1/chat/completions with real-time SSE streaming and loopback CORS isolation.

Automated Process Lifecycle & Cleanup

Guarantees 100% process termination on shutdown or error, with bounded health check polling and zero leftover orphaned daemon processes.

Supply-Chain Cryptographic Integrity

Enforces Ed25519 signed manifests, anti-downgrade policies, symlink traversal blocking, and local build receipts.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Inference Engine llama.cpp (ARM64 NEON DotProd & Device Resource Optimization) Production
Architecture Support Android aarch64 / arm64-v8a (Bionic libc, Qualcomm Adreno, ARM Mali, Samsung Xclipse) Production
Supported Quantizations Q4_K_M, Q5_K_M, Q8_0, IQ4_XS, Q2_K, F16 GGUF v3 Production
Protocols & APIs OpenAI Chat Completions (JSON & SSE Stream), Models List, Health Check Production
SDK Bindings Python 3.8+ (termux-llamacpp), Node.js / TypeScript 16+ (termux-llamacpp) Production

Canonical Usage Example

from termux_llamacpp import LlamaRuntime, RuntimeConfig

# 1. Initialize Runtime with Vulkan GPU Acceleration
config = RuntimeConfig(
    model_path="models/Llama-3.2-1B-Instruct-Q4_K_M.gguf",
    device="auto",
    threads=4,
    context_size=2048
)
runtime = LlamaRuntime(config)

# 2. Synchronous or Streaming Text Generation
output = runtime.generate("한국의 사계절 중 가을의 매력에 대해 설명해줘.", max_tokens=256)
print(output.text)
print(f"Speed: {output.metrics.eval_tokens_per_sec:.2f} t/s (Prompt: {output.metrics.prompt_tokens_per_sec:.2f} t/s)")

Getting Started & Deep Guides