Version Archive & Changelog

Changelog history and immutable releases

v2.0.2 - Vulkan GPU Acceleration & Explicit OpenCL Declaration Protocol (SCRUM-512)

Production Release (Latest) 2026-10-06
  • Vulkan GPU Default & Explicit OpenCL Protocol: Enforced Vulkan GPU priority for mobile neural inference, routing to OpenCL exclusively when explicitly declared via --device opencl with zero silent fallback.
  • Diarization Multi-Thread Optimization: Injected dynamic multi-threading (--embedding.num-threads=4, --segmentation.num-threads=4) eliminating single-core bottlenecks and slashing latency from 109s to 64s (41% reduction).
  • Hybrid Engine GPU Acceleration: Propagated device and threads parameters directly into WhisperEngine and SherpaDiarizer.
  • Full Monorepo SemVer Parity: Synchronized version 2.0.2 across pyproject.toml, package.json, doc.config.yaml, and SDK manifests.

v2.0.1 - Timestamp Interval Display, Benchmark Normalization & Hot-Reload Validation (SCRUM-512)

Production Release 2026-10-06
  • Timestamp Interval Display: Added format_timestamp formatting to output [Speaker_X] [MM:SS.ss -> MM:SS.ss] for all diarization CLI and transcript exports.
  • Benchmark Sample Normalization: Standardized official Korean podcast sample to samples/kor_diarization.wav (45.00s 16kHz Mono PCM).
  • Real-Device 2-vs-3 Speaker & Hot-Reload Benchmark: Validated Galaxy A35 (Exynos 1380) with 2.48s hot-reload acceleration and precise cross-talk decoupling.
  • Version Alignment: Bumped version across pyproject.toml, package.json, doc.config.yaml, and SDK manifests to 2.0.1.

v2.0.0 - Major Upgrade: Next-Gen Neural Diarization & Vosk Deprecation (SCRUM-512)

Major Release 2026-10-06
  • Complete Vosk Legacy Deprecation: Fully purged legacy Vosk Kaldi bindings, xvector engine, and outdated dependencies from Python SDK, Node.js, CLI, and installer.
  • PyAnnote 3.0 & 3D-Speaker CAM++ 192d Integration: Integrated SOTA neural speaker segmentation and Alibaba 3D-Speaker CAM++ 192-dimensional d-vector embeddings.
  • TS-VAD Overlap Speech Resolution: Built OverlapResolver module to resolve simultaneous cross-talk and overlapping speech interruptions ('왜요?' backchannels) into clean multi-speaker streams.
  • Target-Aware Lightweight Provisioning: Default installer keeps base whisper STT lightweight, with neural diarization provisioned via 'termux-stt install --engine diarization' or interactive TTY [y/N] prompts.
  • Benchmark Sample Suite Expansion: Added real Korean podcast 2-speaker cross-talk benchmark ('samples/kor_diarization.wav') and 3-speaker benchmark ('samples/3_speaker_test.wav').
  • Full Monorepo SemVer Parity: Synchronized version 2.0.0 across pyproject.toml, package.json, doc.config.yaml, and SDK manifests.

v1.3.2 - SmartRouter Capability Filtering & CLI Flag De-Duplication

Production Release 2026-09-29
  • SmartRouter Dynamic Split-Mode Filtering: Automatically strips -sm flag when the native binary does not declare split mode support in its help output, eliminating unknown option aborts.
  • Thread Flag -t De-Duplication: Enforces single authoritative -t parameter passed to whisper-cli, eliminating duplicate argument conflicts.
  • Conditional Vulkan Request Logic: Restricts requested_backend='vulkan' strictly to explicit GPU/Vulkan target execution, ensuring pure CPU execution when --device cpu is selected.
  • Upstream HF Model Registry Alignment: Synchronized SHA-256 checksums and model endpoints for Whisper quantized GGML models in registry.py.

v1.3.1 - Asymmetric Mobile SoC Hybrid Pipeline (Encoder Vulkan GPU / Decoder CPU SIMD)

Stable Release 2026-09-29
  • Asymmetric Hardware Hybrid Pipeline: Offloads dense audio encoder GEMMs to mobile Vulkan GPUs while retaining sequential single-token autoregressive decoding on host CPU ARM NEON SIMD vector units.
  • Zero-Copy Cross-Attention DMA Transfer: Automatically syncs calculated KV projections across isolated GGML schedulers via wstate.kv_cross_cpu.
  • High-Performance Verification: Achieves 12.93s wall-clock latency on Whisper Small (Galaxy S25, 21.0% faster than Pure GPU, 16.5% faster than Pure CPU).
  • Dynamic CLI Controls: Added --split-mode (-sm, --hybrid), --optimize-1, and Fail-Fast non-GPU validation.

v1.3.0 - Bionic Zero-Collision Rule & Dual-Flagship Native Vulkan Acceleration

Stable Release 2026-09-28
  • Bionic Linker Zero-Collision Rule: Dynamically purges $PREFIX/lib from LD_LIBRARY_PATH during whisper-cli execution.
  • Dual-Flagship Native Vulkan GPU Universal Acceleration: Validated native Vulkan GPU compute offload across Qualcomm Adreno 830 and ARM Mali-G78.
  • Subprocess ABI Isolation: Standardized RPATH dynamic resolution guaranteeing standalone execution integrity.