Termux-BitNet

Sovereign 1.58-Bit On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install termux-bitnet ameva-runtime

# or: npm install -g termux-bitnet @ameva/runtime

# or: curl -sL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash

The Engineering Challenge

Standard FP16/INT4 LLM inference on mobile architectures suffers from memory bandwidth saturation, severe battery drain, and thermal throttling. Furthermore, upstream 1.58-bit execution historically suffered from ternary bitfield misallocations (word salad), missing weight scales, rigid activation functions, and lack of mobile memory management.

The Architectural Breakthrough

Executes 1.58-bit ternary quantized weights {-1, 0, +1} directly via hand-vectorized ARM64 NEON assembly kernels and native Vulkan compute shaders powered by AMEVA-Runtime. Features hallucination-prevention prompt wrapping, a dynamic zero-overhead activation dispatcher supporting Squared ReLU (Microsoft 2B) and SwiGLU (Falcon-E-1B / Falcon3-7B), GPU watchdog fence chunking, and Zero-Copy mmap streaming for 7.45B model execution on 6GB RAM smartphones.

Key Capabilities & Built-in Hardening

Hallucination Prevention Prompt Wrapping & Architecture Control

Provides full user-configurable prompt formatting (--prompt-template chatml|falcon|raw), custom prefixes/suffixes, and explicit activation routing (--act-fn auto|relu2|swiglu) to permanently eradicate generation drift and word salad across diverse model weights.

Mobile GPU Memory Slicing & Watchdog Defense

Bypasses the 2.5s ARM Mali GPU watchdog fence timeout via chunked layer dispatch (--chunk-layers 4) while dynamically slicing LM Head FP16 projections (--vocab-slice 32768) to reduce GPU VRAM consumption by 576MB.

Dynamic Zero-Overhead Activation Dispatcher

Dynamically auto-dispatches between Squared ReLU (relu2) with Sub-LayerNorm (Microsoft BitNet 2B) and standard SwiGLU (SiLU) without Sub-LayerNorm (Falcon-E-1B, Falcon3-7B) based on tensor layout introspection.

Mathematical Elimination of Ternary Numerical Collapse

Canonical mathematical decoding w = (b & 3) - 1 eradicates 49.6% inactive neuron corruption, while integrating 32-byte GGUF tensor trailer weight scales into GEMV scaling for calibrated logit distributions.

Zero-Copy Mmap Streaming for 7.45B Edge Inference

Enables seamless on-device execution of 3.05GB 7.45B models (Falcon3-7B) on 6GB RAM devices without triggering Android Low Memory Killer (LMK) eviction.

Native Vulkan Compute GPU Acceleration

Powered by AMEVA-Runtime's modular SPIR-V compute pipeline (Ternary GEMV, RoPE, Attention Decode, SwiGLU, RMSNorm, FP16 LM Head), fusing 30 transformer layers into a single VkCommandBuffer submission.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Supported Architectures Microsoft BitNet b1.58 (2B-4T), TII Falcon-E (1B Instruct), TII Falcon3 (7.45B Instruct), BitNet Embedding (270M) Production (v2.0.1)
Hallucination Defense & Wrapping ChatML, Falcon Instruct, Raw QA, Custom Prefix/Suffix, Dynamic EOS Overrides (--prompt-template) Production (v2.0.1)
GPU Slicing & Watchdog Defense Layer Chunking (--chunk-layers 4), Vocabulary Slicing (--vocab-slice 32768, -576MB VRAM) Production (v2.0.1)
Activation Dispatch Squared ReLU (relu2) with Sub-LayerNorm, SwiGLU (SiLU), RMSNorm, LayerNorm (--act-fn) Production (v2.0.1)
CPU Kernel Core ARM64 NEON DotProd (sdot / vdotq_s32), i2_s Ternary SIMD, Canonical Unpacking w = (b & 3) - 1 Production
GPU Compute Vulkan Compute SPIR-V (Qualcomm Adreno 830/700/600, ARM Mali-G78/G68/G77, Samsung Xclipse) Production
Memory Governance Zero-Copy Safetensors/GGUF mmap Streaming, Direct KV Cache, 6GB RAM Android LMK Defense Production
Language Bindings Python CFFI / ctypes (termux-bitnet), Node.js Thin Gateway (termux-bitnet-js), AMEVA Component SDK Production

Canonical Usage Example

from termux_bitnet import BitNetEngine, BitNetConfig

# 1. Initialize engine with model wrapping and GPU acceleration options
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
    device="gpu",           # "auto", "cpu", or "gpu" (Vulkan via ameva-runtime)
    n_gpu_layers=24,        # Offload 24 layers to GPU
    chat_template="falcon", # Auto-wrapping: ChatML, Falcon, or Raw
    chunk_layers=4,         # Prevent Mali GPU watchdog timeout
    vocab_slice=32768,      # Save 576MB VRAM on LM Head
    n_threads=4,
    temperature=0.7,
    top_p=0.95
)

# 2. Stream tokens in real time (up to 34.35 tok/s on S25, 4.46 tok/s on A53)
with BitNetEngine(config) as engine:
    print("[Prompt]: What is the capital of France?")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("What is the capital of France?"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s | Prompt tokens: {metrics.prompt_tokens}")

Getting Started & Deep Guides