Version Archive & Changelog
Changelog history and immutable releases
v2.0.1 - Hallucination Prevention Prompt Wrapping, Mali GPU Chunking & Multi-Device Hardening
Patch and Engineering Hardening Release
2026-10-01
- Hallucination Prevention Prompt Wrapping: Introduced configurable --prompt-template (chatml, falcon, raw, none) alongside custom prefixes/suffixes and EOS overrides to eliminate conversational divergence across models.
- Mali Watchdog Timeout Prevention: Introduced --chunk-layers (default 4) command-line option to slice GPU command dispatch buffers, eliminating the 2.5s Mali GPU watchdog fence timeout on Exynos devices.
- Dynamic Vocabulary Slicing: Added --vocab-slice (default 32768) option to dynamically slice LM Head linear projection, reducing GPU VRAM allocation by 576MB.
- Architecture Routing Control: Added explicit --act-fn (auto, relu2, swiglu) flag for user-directed mathematical activation kernel selection.
- Full GPU Acceleration Engine Audit: Comprehensive inventory and validation of all 8 GLSL compute shaders, compiled SPIR-V bytecode binaries, embedded C++ array headers, and dynamic Vulkan loader.
v2.0.0 - Sovereign Ternary - Universal Multi-Model 1.58-Bit Engine & 7B Edge Inference
Major Architectural Generation Release
2026-10-01
- Root-Cause Elimination of Word Salad: Corrected GGUF i2_s ternary bit unpacking from legacy (b & 1) - (b >> 1) to canonical w = (b & 3) - 1, preventing 49.6% zero neuron corruption from causing activation norm explosion.
- 32-Byte Tensor Trailer Weight Scale Integration: Restores exact dequantization scaling S_W = mean(|W|) across 30 transformer layers, preserving calibrated logit distributions.
- Zero-Overhead Dynamic Activation Dispatcher: Introspects layer weight tensors to dynamically auto-dispatch between Squared ReLU with Sub-LayerNorm (Microsoft BitNet 2B) and standard SwiGLU SiLU (Falcon-E-1B, Falcon3-7B).
- 6GB RAM 7.45B Model Edge Inference: Verified Zero-Copy mmap streaming of 3.05GB Falcon3-7B model on Galaxy A53 and A35 without Android Low Memory Killer (LMK) eviction.
- Real-Time Conversational Generation (9.69 tok/s): Achieved 9.69 tok/s on Galaxy A53 running Falcon-E-1B, delivering completely offline interactive dialog on mainstream mobile silicon.
- Multi-Model Fleet Scorecard Published: Documented verified ground truth across Galaxy S25, Galaxy A53, and Galaxy A35 for 4 distinct 1.58-bit model families.
- Foundation Research Whitepaper Linkage: Directly linked to official monograph AOSF-TR-2026-BITNET-TERNARY-02 and upstream PRs #551 and #624.
v1.4.0 - Native Vulkan Compute GPU Acceleration, Permanent VRAM Residency & Dynamic C ABI
Stable Release
2026-09-07
- Native Vulkan Compute GPU Engine via AMEVA-Runtime with SPIR-V kernels.
- Permanent Model VRAM Residency eliminating host-to-device bus traffic.
- Full-Pipeline On-Chain Token Execution fusing 30 layers into single command buffer.