Termux-BitNet
Sovereign 1.58-Bit On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration
Install the official package directly into your runtime:
pip install termux-bitnet ameva-runtime
# or: npm install -g termux-bitnet @ameva/runtime
# or: curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash
The Engineering Challenge
Standard FP16/INT4 LLM inference on mobile architectures suffers from memory bandwidth saturation, severe battery drain, and thermal throttling. Furthermore, upstream 1.58-bit execution historically suffered from ternary bitfield misallocations (word salad), missing weight scales, rigid activation functions, and lack of mobile memory management.
The Architectural Breakthrough
Executes 1.58-bit ternary quantized weights {-1, 0, +1} directly via hand-vectorized ARM64 NEON assembly kernels and native Vulkan compute shaders powered by AMEVA-Runtime. Features hallucination-prevention prompt wrapping, a dynamic zero-overhead activation dispatcher supporting Squared ReLU (Microsoft 2B) and SwiGLU (Falcon-E-1B / Falcon3-7B), GPU watchdog fence chunking, and Zero-Copy mmap streaming for 7.45B model execution on 6GB RAM smartphones.
Key Capabilities & Built-in Hardening
Hallucination Prevention Prompt Wrapping & Architecture Control
Provides full user-configurable prompt formatting (--prompt-template chatml|falcon|raw), custom prefixes/suffixes, and explicit activation routing (--act-fn auto|relu2|swiglu) to permanently eradicate generation drift and word salad across diverse model weights.
Mobile GPU Memory Slicing & Watchdog Defense
Bypasses the 2.5s ARM Mali GPU watchdog fence timeout via chunked layer dispatch (--chunk-layers 4) while dynamically slicing LM Head FP16 projections (--vocab-slice 32768) to reduce GPU VRAM consumption by 576MB.
Dynamic Zero-Overhead Activation Dispatcher
Dynamically auto-dispatches between Squared ReLU (relu2) with Sub-LayerNorm (Microsoft BitNet 2B) and standard SwiGLU (SiLU) without Sub-LayerNorm (Falcon-E-1B, Falcon3-7B) based on tensor layout introspection.
Mathematical Elimination of Ternary Numerical Collapse
Canonical mathematical decoding w = (b & 3) - 1 eradicates 49.6% inactive neuron corruption, while integrating 32-byte GGUF tensor trailer weight scales into GEMV scaling for calibrated logit distributions.
Zero-Copy Mmap Streaming for 7.45B Edge Inference
Enables seamless on-device execution of 3.05GB 7.45B models (Falcon3-7B) on 6GB RAM devices without triggering Android Low Memory Killer (LMK) eviction.
Native Vulkan Compute GPU Acceleration
Powered by AMEVA-Runtime's modular SPIR-V compute pipeline (Ternary GEMV, RoPE, Attention Decode, SwiGLU, RMSNorm, FP16 LM Head), fusing 30 transformer layers into a single VkCommandBuffer submission.
Supported Compute Kernels & Operations
| Subsystem Category | Operations & Kernels | Status |
|---|---|---|
| Supported Architectures | Microsoft BitNet b1.58 (2B-4T), TII Falcon-E (1B Instruct), TII Falcon3 (7.45B Instruct), BitNet Embedding (270M) | Production (v2.0.1) |
| Hallucination Defense & Wrapping | ChatML, Falcon Instruct, Raw QA, Custom Prefix/Suffix, Dynamic EOS Overrides (--prompt-template) | Production (v2.0.1) |
| GPU Slicing & Watchdog Defense | Layer Chunking (--chunk-layers 4), Vocabulary Slicing (--vocab-slice 32768, -576MB VRAM) | Production (v2.0.1) |
| Activation Dispatch | Squared ReLU (relu2) with Sub-LayerNorm, SwiGLU (SiLU), RMSNorm, LayerNorm (--act-fn) | Production (v2.0.1) |
| CPU Kernel Core | ARM64 NEON DotProd (sdot / vdotq_s32), i2_s Ternary SIMD, Canonical Unpacking w = (b & 3) - 1 | Production |
| GPU Compute | Vulkan Compute SPIR-V (Qualcomm Adreno 830/700/600, ARM Mali-G78/G68/G77, Samsung Xclipse) | Production |
| Memory Governance | Zero-Copy Safetensors/GGUF mmap Streaming, Direct KV Cache, 6GB RAM Android LMK Defense | Production |
| Language Bindings | Python CFFI / ctypes (termux-bitnet), Node.js Thin Gateway (termux-bitnet-js), AMEVA Component SDK | Production |
Canonical Usage Example
from termux_bitnet import BitNetEngine, BitNetConfig
# 1. Initialize engine with model wrapping and GPU acceleration options
config = BitNetConfig(
model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
device="gpu", # "auto", "cpu", or "gpu" (Vulkan via ameva-runtime)
n_gpu_layers=24, # Offload 24 layers to GPU
chat_template="falcon", # Auto-wrapping: ChatML, Falcon, or Raw
chunk_layers=4, # Prevent Mali GPU watchdog timeout
vocab_slice=32768, # Save 576MB VRAM on LM Head
n_threads=4,
temperature=0.7,
top_p=0.95
)
# 2. Stream tokens in real time (up to 34.35 tok/s on S25, 4.46 tok/s on A53)
with BitNetEngine(config) as engine:
print("[Prompt]: What is the capital of France?")
print("[Response]: ", end="", flush=True)
for token in engine.generate_stream("What is the capital of France?"):
print(token, end="", flush=True)
print()
metrics = engine.get_last_metrics()
print(f"Speed: {metrics.tokens_per_second:.2f} tok/s | Prompt tokens: {metrics.prompt_tokens}")