Version Archive & Changelog

Changelog history and immutable releases

v2.0.1 - Hallucination Prevention Prompt Wrapping, Mali GPU Chunking & Multi-Device Hardening

Patch and Engineering Hardening Release 2026-10-01
  • Hallucination Prevention Prompt Wrapping: Introduced configurable --prompt-template (chatml, falcon, raw, none) alongside custom prefixes/suffixes and EOS overrides to eliminate conversational divergence across models.
  • Mali Watchdog Timeout Prevention: Introduced --chunk-layers (default 4) command-line option to slice GPU command dispatch buffers, eliminating the 2.5s Mali GPU watchdog fence timeout on Exynos devices.
  • Dynamic Vocabulary Slicing: Added --vocab-slice (default 32768) option to dynamically slice LM Head linear projection, reducing GPU VRAM allocation by 576MB.
  • Architecture Routing Control: Added explicit --act-fn (auto, relu2, swiglu) flag for user-directed mathematical activation kernel selection.
  • Full GPU Acceleration Engine Audit: Comprehensive inventory and validation of all 8 GLSL compute shaders, compiled SPIR-V bytecode binaries, embedded C++ array headers, and dynamic Vulkan loader.

v2.0.0 - Sovereign Ternary - Universal Multi-Model 1.58-Bit Engine & 7B Edge Inference

Major Architectural Generation Release 2026-10-01
  • Root-Cause Elimination of Word Salad: Corrected GGUF i2_s ternary bit unpacking from legacy (b & 1) - (b >> 1) to canonical w = (b & 3) - 1, preventing 49.6% zero neuron corruption from causing activation norm explosion.
  • 32-Byte Tensor Trailer Weight Scale Integration: Restores exact dequantization scaling S_W = mean(|W|) across 30 transformer layers, preserving calibrated logit distributions.
  • Zero-Overhead Dynamic Activation Dispatcher: Introspects layer weight tensors to dynamically auto-dispatch between Squared ReLU with Sub-LayerNorm (Microsoft BitNet 2B) and standard SwiGLU SiLU (Falcon-E-1B, Falcon3-7B).
  • 6GB RAM 7.45B Model Edge Inference: Verified Zero-Copy mmap streaming of 3.05GB Falcon3-7B model on Galaxy A53 and A35 without Android Low Memory Killer (LMK) eviction.
  • Real-Time Conversational Generation (9.69 tok/s): Achieved 9.69 tok/s on Galaxy A53 running Falcon-E-1B, delivering completely offline interactive dialog on mainstream mobile silicon.
  • Multi-Model Fleet Scorecard Published: Documented verified ground truth across Galaxy S25, Galaxy A53, and Galaxy A35 for 4 distinct 1.58-bit model families.
  • Foundation Research Whitepaper Linkage: Directly linked to official monograph AOSF-TR-2026-BITNET-TERNARY-02 and upstream PRs #551 and #624.

v1.4.0 - Native Vulkan Compute GPU Acceleration, Permanent VRAM Residency & Dynamic C ABI

Stable Release 2026-09-07
  • Native Vulkan Compute GPU Engine via AMEVA-Runtime with SPIR-V kernels.
  • Permanent Model VRAM Residency eliminating host-to-device bus traffic.
  • Full-Pipeline On-Chain Token Execution fusing 30 layers into single command buffer.