Termux-STT
Unified On-Device Speech-to-Text & Neural Multi-Speaker Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD Overlap Resolver)
Install the official package directly into your runtime:
pip install --upgrade termux-stt && termux-stt install
# or Node.js:
npm install -g termux-stt && termux-stt install
The Engineering Challenge
Transcribing and diarizing multi-speaker conversations on edge devices typically demands heavy PyTorch runtimes (>2GB), external GPU servers, and severe cloud network latency with audio privacy risks.
The Architectural Breakthrough
Integrates Whisper.cpp (Native Vulkan GPU / NEON CPU) and Sherpa-ONNX with PyAnnote 3.0 neural segmentation, 3D-Speaker CAM++ 192-dim embeddings, and TS-VAD overlapped speech resolution with zero cloud egress.
Key Capabilities & Built-in Hardening
Asymmetric Vulkan GPU / CPU Hybrid Pipeline
Offloads dense audio encoder GEMMs to Vulkan mobile GPUs while retaining single-token sequential autoregressive decoding on host ARM NEON CPU SIMD vector units, achieving up to 21% lower latency.
Dual Modern STT Engine Architecture
Seamlessly switches between Whisper.cpp (high accuracy, Vulkan GPU/CPU) and Sherpa-ONNX SenseVoice (ultra-low latency on-device ASR) via a unified create_engine() factory.
Next-Gen Neural Diarization (PyAnnote 3.0 + CAM++ 192d)
Deploys PyAnnote 3.0 neural segmentation ONNX combined with Alibaba 3D-Speaker CAM++ 192-dimensional d-vector embeddings, achieving high speaker clustering accuracy without PyTorch.
TS-VAD Overlapped Speech Resolution
Resolves cross-talk and sudden interruptions (e.g., overlapping questions) into distinct speaker audio streams using target-speaker acoustic scanning.
Zero Cloud Egress Audio Privacy
Audio capture, acoustic feature extraction, neural diarization, and text transcription execute strictly on local mobile hardware without network egress.
Subprocess Crash Isolation
Isolates native C++ binaries within dedicated process pools, ensuring C++ segfaults never compromise the host Python or Node.js runtime.
Modular & Interactive Model Provisioning
Keeps standard installation lightweight while offering on-demand neural diarization installation (--engine diarization / --all) with interactive TTY [y/N] prompts.
Supported Compute Kernels & Operations
| Subsystem Category | Operations & Kernels | Status |
|---|---|---|
| Asymmetric Hybrid Pipeline | Vulkan GPU Audio Encoder + ARM NEON CPU Sequential Decoder (--split-mode / --optimize-1) | Production |
| Speech Recognition (STT) | Whisper.cpp (tiny/base/small/medium/large-v3-turbo), Sherpa-ONNX (SenseVoice) | Production |
| Speaker Diarization | PyAnnote 3.0 ONNX, 3D-Speaker CAM++ 192d, TS-VAD OverlapResolver, Cosine K-Means | Production |
| Audio Preprocessing | Zero-Subprocess Wave Parser, FFmpeg Auto-Fallback, Silero-VAD Silence Filter | Production |
| Subtitle & Data Formats | SubRip (SRT), WebVTT (VTT), NIST Rich Transcription (RTTM), Structured JSON | Production |
| Platform Runtimes | Android Termux (Bionic ARM64/aarch64), Python 3.8~3.13, Node.js 16~22 | Production |
Canonical Usage Example
from termux_stt import create_engine
# 1. Initialize Whisper Engine with Vulkan GPU / CPU Hybrid Acceleration
engine = create_engine("whisper", model="small", lang="ko", threads=4, split_mode=True)
# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())
# 3. Multi-Speaker Neural Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD)
hybrid = create_engine("hybrid", lang="ko", num_speakers=2)
diar_result = hybrid.diarize("samples/kor_diarization.wav")
for seg in diar_result.segments:
print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")