Version Archive & Changelog

Changelog history and immutable releases

v2.8.1 - Context Shift on the GPU Route (ROPE on the f16 KV cache in ggml-ameva)

Patch Release 2026-10-05
  • Fix: with the backend of 2.8.0, text generation past the end of the context with context shifting enabled stopped with an abort on the GPU route (llama.cpp rotates the stored keys of its f16 KV cache in place; ggml-ameva declined ROPE on f16 tensors). ggml-ameva now runs it. ameva-run exec does not enable context shifting itself.
  • llama bundle: backend built from ameva-compute 7e99154; the four other bundles are unchanged. On devices without integer dot products in Vulkan (Galaxy S21, A53) token generation can use an integer dot product computed with 32-bit multiplications where measured faster; their default route stays the CPU.
  • Verified on all five fleet devices (Galaxy S25, S20, A35, A53, S21) with tools/fleet/lab/verify_kv.sh (context shift and KV cell reuse against the CPU, a 2048-token context, ggml test-backend-ops) and tools/fleet/lab/test_rel.sh (fresh installation, upgrade from the published 2.8.0).
  • Correction to the 2.8.0 notes: on the Galaxy S25 with all eight cores and no other load the CPU generates tokens faster than the GPU route (79-105 against 65-76 t/s for Qwen2.5-0.5B); the GPU route remains 5 to 9 times faster in prompt processing.

v2.8.0 - ggml-ameva as the LLM GPU Route (own Vulkan backend in the llama bundle, integer-dot kernels)

Feature Release 2026-10-04
  • LLM GPU route: the llama bundle carries the project's own Vulkan backend (lib/libggml-ameva.so with lib/libac_runtime.so.2, built against the bundle's ggml 0.9.8). ameva-run exec and ameva_runtime.run() run the Vulkan plan on it (--device AMEVA0, GGML_BACKEND_PATH of the started bundle). ggml-vulkan of the bundled llama.cpp returned corrupt text with flash attention and stopped without it on Mali and Turnip; it stays selectable with AMEVA_LLM_GPU_BACKEND=upstream.
  • Backend: weights are converted in a staging copy (Galaxy A35: Llama-3.2-1B loads in 2.6 s instead of 81 s); token generation with packed integer dot products (ameva_mmvq) where the driver offers them, measured per device and weight shape together with the f32 layouts and used where more than 10 % faster.
  • Routes: Adreno 600 series (Galaxy S20) takes the GPU by default; an explicit 'auto' request is the adaptive route (it used to fall through to the CPU route). Galaxy S21 and A53 keep the CPU route.
  • Measured with Qwen2.5-0.5B, CPU / GPU t/s: prompt processing Galaxy S25 30-58 / 1010-1081, S20 60 / 181-197, A35 109-110 / 190, S21 115-119 / 227, A53 55 / 67; token generation S25 16-49 / 36-77, S20 26-27 / 23.6-24.4, A35 35-37 / 21.5-21.6, S21 46-47 / 22.8-23.1, A53 22-24 / 14.5. CPU-relative KL divergence 0.0026 to 0.0064 (gate 0.01).
  • tools/fleet/divert_apt_llama.sh: dpkg diversions that keep the Termux apt package llama-cpp from overwriting the runtime's llama-cli, llama-server and llama-completion links.

v2.7.9 - Bundle Isolation (no bundle file in $PREFIX/lib, self-contained bundles, one Vulkan HAL shim)

Maintenance Release 2026-10-04
  • Installation layout: a bundle stays in its release directory and its commands in $PREFIX/bin are links into it. Up to 2.7.8 the libraries of all five bundles were copied to $PREFIX/lib, where the ggml libraries of llama.cpp and whisper.cpp share SONAMEs, the STT bundle replaced the libc++_shared.so of the Termux libc++ package, and llama.cpp and TTS found libomp.so only where the diffusion bundle had been installed.
  • Bundles: llama.cpp and TTS carry libomp.so (repacked, no ELF file rebuilt); the llama.cpp bundle stores each library once (archive 125 MB to 26 MB). The installer refuses a bundle that needs a library that neither the bundle, Android, nor a Termux package provides.
  • ameva-run migrate: for installations made by earlier versions; copied commands become links and library copies that nothing is found to load are moved to ~/.local/share/ameva/legacy-lib (not deleted).
  • Vulkan HAL shim: one build with its source in the repository, without the subgroup-size clamp that made ggml-vulkan abort on Adreno 830; used from inside the package, no longer copied to $PREFIX/lib.
  • LLM routes: ameva-run exec passes the router's options to llama-cli (it dropped them, among them -fa 0); the CPU route passes --device none (with -ngl 0 alone prompt batches still ran on the GPU backend); the Mali Vulkan plans turn flash attention off, with which the bundle's Vulkan backend returns corrupt text. The routing decisions are those of 2.7.8.

v2.7.8 - Release Integrity (STT bundle layout and pin, ELF RUNPATH, checksum sidecars)

Maintenance Release 2026-10-04
  • STT bundle: the installer pinned a bundle that is no longer published and required bin/whisper-cli, while the published bundle was flat, so STT provisioning failed in 2.7.2-2.7.7. The v2.7.8 bundle holds the same files byte for byte under bin/ and lib/; the pin matches it.
  • ELF RUNPATH corrected in termux-bitnet-cli, sherpa-ncnn-offline-tts and libggml-vulkan.so (home build directories, an empty entry and malformed entries removed; $ORIGIN first). Code unchanged; only the RUNPATH string differs from v2.7.3.
  • Checksum sidecars: one .sha256 per asset, without a byte order mark.
  • Published as a non-latest GitHub release so that 2.7.4-2.7.7, which fall back to the latest release, keep resolving the bundles their pins match.

v2.7.7 - Production Release & NPM Registry Alignment (Target-Aware 6-Modality Isolation)

Production Release 2026-09-29
  • Universal Registry Alignment: Synchronized PyPI, NPM (@ameva/runtime), and GitHub releases under unified v2.7.7 semantic versioning.
  • Target-Aware 6-Modality Library Isolation: Replaced monolithic for-loop library path injection with strict 1:1 MODALITY_LIBRARY_MAP in base.py and get_vulkan_env(), completely eliminating cross-contamination crashes between whisper.cpp and llama.cpp (libggml-vulkan.so).
  • Universal Modality Contract Enforcement: Standardized modality = "" and kwargs.setdefault("modality", cls.modality) across all 6 adapters (llamacpp, stt, vision, tts, diffusion, bitnet).
  • Qualcomm Adreno Smart Routing Architecture: Configured SmartRouter to auto-dispatch Snapdragon 865 (Adreno 650) to high-precision OpenCL backend (llama-cli-opencl) or CPU NEON, eliminating Vulkan subnormal FP16 repetition degeneration while retaining full GPU execution for Mali and modern Adreno 800+ silicons.
  • Adreno 600 Flash Attention Auto-Defense: Enforced -fa 0 injection for legacy Adreno GPUs in termux-llamacpp to prevent driver shader compilation deadlocks.

v2.7.6 - Target-Aware 6-Modality Library Isolation & Qualcomm Adreno OpenCL Smart Routing

Production Release 2026-09-29
  • Target-Aware 6-Modality Library Isolation: Replaced monolithic for-loop library path injection with strict 1:1 MODALITY_LIBRARY_MAP in base.py and get_vulkan_env(), completely eliminating cross-contamination crashes between whisper.cpp and llama.cpp (libggml-vulkan.so).
  • Universal Modality Contract Enforcement: Standardized modality = "" and kwargs.setdefault("modality", cls.modality) across all 6 adapters (llamacpp, stt, vision, tts, diffusion, bitnet).
  • Qualcomm Adreno Smart Routing Architecture: Configured SmartRouter to auto-dispatch Snapdragon 865 (Adreno 650) to high-precision OpenCL backend (llama-cli-opencl) or CPU NEON, eliminating Vulkan subnormal FP16 repetition degeneration while retaining full GPU execution for Mali and modern Adreno 800+ silicons.
  • Adreno 600 Flash Attention Auto-Defense: Enforced -fa 0 injection for legacy Adreno GPUs in termux-llamacpp to prevent driver shader compilation deadlocks.

v2.7.5 - Unified 5-Backend Architecture Standard & OpenCL Driver System Path Injection

Production Release 2026-09-29
  • Unified 5-Backend Architecture Standard: Formalized Vulkan, OpenCL, OpenCL-System, CPU NEON, and CPU Reference execution matrix.
  • OpenCL System Driver Injection: Added auto-detection and injection for vendor OpenCL ICD loaders across mobile chipsets.

v2.7.4 - Android 15 Linker Namespace Isolation, ChatML Auto-Templating & Adreno 600 Defense

Production Release 2026-09-29
  • Android 15/16 Bionic Dynamic Linker Namespace Isolation: Purged regressive $PREFIX/lib injection from get_vulkan_env(), preventing libunwindstack symbol collision crashes with Termux userland.
  • ChatML Prompt Template Auto-Encapsulation: Added automated ChatML encapsulation (<|im_start|> tags) and reverse stop token binding (-r "<|im_end|>") for Qwen and Llama-3 models, eliminating repetition loops.
  • Qualcomm Adreno 600 Series Flash Attention Defense: Implemented automatic -fa 0 injection for Snapdragon 865 (Adreno 600) to prevent closed-source driver segfaults, while preserving Flash Attention on Adreno 830 and Mali GPUs.

v2.7.3 - Termux-Diffusion Integration (Z-Image Turbo 6.0B DiT), Two-Track Supply Chain & Android 15 Zero-Collision

Production Release 2026-09-28
  • Termux-Diffusion SOTA Z-Image Turbo 6.0B DiT Integration: Verified on-device DiT compute pipeline with Flash Attention (--diffusion-fa) and VAE tiled decoding across mobile Qualcomm Adreno and ARM Mali GPUs.
  • Two-Track Enterprise Supply Chain Architecture: Separated proprietary core source code (100% Private) from public release binary hub (uno-km/ameva-runtime-releases), enabling anonymous, tokenless 1-Click native engine auto-provisioning.
  • Automated Production Release Fallback: Engineered resilient asset resolution defaulting to releases/latest/download when target version tag is unavailable.
  • Android 15 Bionic libc & Dynamic Linker Isolation: Reinforced symbol resolution in BionicDirectVulkanLoader against allocator and namespace collisions, eliminating Mesa llvmpipe CPU rasterizer traps.

v2.7.2 - Dual-Generation Qualcomm Adreno Vulkan Zero-NaN & Bionic Isolation

Maintenance Release 2026-09-18
  • Qualcomm Adreno Vulkan Zero-NaN Architecture: Resolved IEEE 754 denormal/subnormal flush-to-zero divergence and NaN token degeneration across Adreno 830 and Adreno 650.
  • Dynamic Asset Resolution SSOT: Refactored NativeAssetManager to resolve release assets dynamically with fail-fast checksum validation.
  • Unified Modality Adapters: Synchronized execution pipelines across LLM, STT, TTS, Diffusion, Vision, and BitNet.

v2.7.1 - Strict Model Resolution & Vulkan GPU Layer Offloading Expansion

Hardening Release 2026-09-15
  • Strict Model Resolution Guard (AmbiguousModelMatchError): Eliminated arbitrary model path guessing and enforced deterministic candidate resolution.
  • Vulkan Layer Offloading Expansion (requested_ngl=999): Enforced 100% VRAM offloading default across runtime execution routing.

v2.0.0 - Unified Single-Package Architecture, Mali Valhall MatMul Loop Elimination & STT 2.26x Acceleration

Major Milestone Release 2026-09-05
  • ARM Mali Valhall MatMul Zero-Stride Infinite Loop Elimination: Identified and resolved GLSL integer truncation in mul_mm.comp (loadstride_b = 16 * 1 / 32 = 0) on subgroup-16 devices, enforcing Medium kernels (_m, loadstride_b = 4) and achieving 4.44 tokens/sec on Galaxy A35 with 25/25 layers GPU offload (+26.9% faster than CPU).
  • Whisper STT 2.26x GPU Acceleration: Verified Whisper Large-v3-Turbo on Galaxy A35 completing in 360.60s (56% time reduction vs CPU 816.48s), pinning GPU at 100% (949 MHz) while lowering CPU load from 291% to 20~30%.
  • Qualcomm Adreno 830 JIT Bug Isolation: Handled Adreno JIT compiler crash on specialization constant NUM_COLS >= 3 by bounding mul_mat_vec_max_cols = 2, enabling stable GPU inference in 4,401 ms.
  • PyTorch-Style Single Package Architecture: Consolidated repository and packaging under ameva-runtime (v2.0.0), housing Vulkan acceleration in ameva_runtime.vulkan submodule with dynamic single-source-of-truth versioning.
  • Complete Sibling Ecosystem Migration: Updated termux-stt, termux-vision, termux-llamacpp, termux-diffusion, termux-bitnet, termux-tts, and termux-train to directly import ameva_runtime.vulkan.

v1.0.0 - Unified Production Release with Real Adreno & Mali Empirical Benchmarks

Official Production Release 2026-09-04
  • Snapdragon 8 Elite (Adreno 830) Validation: Verified 25/25 layer full VRAM offload achieving 35.80 tokens/sec.
  • ARM Mali Headless Fence Deadlock Isolation: Handled proprietary Mali driver power-management stalls.
  • Unified AmevaRuntime Python Execution Engine: Added ameva.run() and ameva.plan() top-level APIs.