Version Archive & Changelog
Changelog history and immutable releases
v1.3.13 - OpenCL Acceleration Engine & Unified 5-Backend Standard
Production Release (Latest)
2026-09-29
- Unified 5-Backend Standard: Standardized all CLI subcommands (
run,serve,benchmark) to enforce the unified backend choices:auto,gpu,vulkan,opencl,cpu. - Qualcomm Adreno OpenCL Acceleration: Added native
openclbackend routing via optimized OpenCL 2.0 kernels, successfully resolving driver-level SPIR-V Vulkan defect on Adreno 650 (Snapdragon 865). - Installer Symlink Hardening: Purged conflicting symlinks to prevent unintentional overwrite of native binary assets.
- Fleet Verification: Empirically verified on Galaxy S20 (OpenCL offload) and Galaxy S25 (Vulkan/Flash Attention).
v1.3.12 - Dynamic Linker Bionic Isolation & ChatML Repetition Mitigation
Stable Release
2026-09-29
- Dynamic Linker Bionic Isolation: Purged regressive $PREFIX/lib injection from fallback runtime environment, guaranteeing Android 15/16 Bionic namespace isolation and zero libunwindstack symbol collisions.
- ChatML Auto-Templating: Standardized <|im_start|> prompt encapsulation and automatic -r '<|im_end|>' reverse stop token injection, eliminating repetition loops on Qwen and Llama-3.
- Cross-SoC Mobile Inference Stability: Retained verified -fa 0 defense on Adreno 600 while maintaining Flash Attention on Adreno 830 and ARM Mali GPUs.
v1.3.10 - Dynamic Network Resume & Qualcomm Adreno Flash Attention Defense
Stable Release
2026-09-28
- Dynamic 5-Stage Network Resume: Exponential backoff with HTTP Range 206/416 self-healing for interrupted GGUF downloads.
- Adreno Flash Attention Defense: Injected -fa 0 by default to prevent Qualcomm driver assertion crashes during mobile GPU inference.
- Pure CPU Device Isolation: Set GGML_VK_VISIBLE_DEVICES='' under device='cpu' to prevent buggy vendor Vulkan driver crashes.
- Chat Template Auto-Resolution: Automatic chat template expansion for Qwen and Llama-3 models with stop token cleanup.
- Official ameva-runtime Adapter Integration: Deep binding with LlamaCppAdapter for Bionic HAL orchestration.
v1.2.0 - Unified AMEVA Vulkan HAL & Galaxy A35 Validation
Stable Release
2026-09-01
- Unified AMEVA Vulkan HAL Integration: Deep binding with ameva-vulkan-runtime v1.1.0 utilizing Android Bionic Vulkan ICD (/system/lib64/libvulkan.so).
- Strict 3-Tier Execution Mode: Added --device vulkan (Fail-Fast GPU compute), --device auto (transparent CPU recovery), and --device cpu (0ms direct NEON forward pass).
- Dynamic Octa-Core Thread Tuning: Automatic detection of ARM big.LITTLE architectures (Exynos 1380, Snapdragon) injecting optimal big-core threads (-t 4).
- Dual Package Distribution: Synchronously released to Python PyPI (termux-llamacpp v1.2.0) and Node.js npm (termux-llamacpp v1.2.0).
- Real-Device Galaxy A35 Validation: End-to-end verification across CLI, Python SDK, Node.js npm, and background OpenAI REST server daemon.
- Supply-Chain Receipt Gatekeeper: Added verified build receipt support (llama-server.build-receipt.json) for local native binaries.
v1.1.1 - Supply-Chain Cryptographic Verification & Multi-Channel Crawler
Stable Release
2026-08-30
- Ed25519 Public Key Verification: Cryptographic manifest verification with SHA-256 binary validation.
- Multi-Channel HuggingFace GGUF Crawler: Deep model discovery and automated quant downloader.
- Process Supervisor Resilience: Guaranteed zero-zombie process termination with tracked PID ledger.
v1.0.2 - Bionic Dynamic Linker Stabilization & Dual-Ecosystem Publishing
Stable Release
2026-08-28
- Bionic Shared Library Bundling: Packaged libgomp and libstdc++ to prevent missing symbol errors on raw Termux.
- NPM Binary Wrapper: Added standalone ESM binary resolver for Node.js environments.
v1.0.1 - Initial Production GGUF LLM Runtime Release
Initial Release
2026-08-27
- Prebuilt ARM64 llama.cpp Native Binaries: Direct extraction in under 3 seconds.
- Loopback Isolated OpenAI Server: /v1/chat/completions REST and SSE streaming supervisor.