Termux-STT

Unified On-Device Speech-to-Text & Neural Multi-Speaker Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD Overlap Resolver)

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install --upgrade termux-stt && termux-stt install
# or Node.js:
npm install -g termux-stt && termux-stt install

The Engineering Challenge

Transcribing and diarizing multi-speaker conversations on edge devices typically demands heavy PyTorch runtimes (>2GB), external GPU servers, and severe cloud network latency with audio privacy risks.

The Architectural Breakthrough

Integrates Whisper.cpp (Native Vulkan GPU / NEON CPU) and Sherpa-ONNX with PyAnnote 3.0 neural segmentation, 3D-Speaker CAM++ 192-dim embeddings, and TS-VAD overlapped speech resolution with zero cloud egress.

Key Capabilities & Built-in Hardening

Asymmetric Vulkan GPU / CPU Hybrid Pipeline

Offloads dense audio encoder GEMMs to Vulkan mobile GPUs while retaining single-token sequential autoregressive decoding on host ARM NEON CPU SIMD vector units, achieving up to 21% lower latency.

Dual Modern STT Engine Architecture

Seamlessly switches between Whisper.cpp (high accuracy, Vulkan GPU/CPU) and Sherpa-ONNX SenseVoice (ultra-low latency on-device ASR) via a unified create_engine() factory.

Next-Gen Neural Diarization (PyAnnote 3.0 + CAM++ 192d)

Deploys PyAnnote 3.0 neural segmentation ONNX combined with Alibaba 3D-Speaker CAM++ 192-dimensional d-vector embeddings, achieving high speaker clustering accuracy without PyTorch.

TS-VAD Overlapped Speech Resolution

Resolves cross-talk and sudden interruptions (e.g., overlapping questions) into distinct speaker audio streams using target-speaker acoustic scanning.

Zero Cloud Egress Audio Privacy

Audio capture, acoustic feature extraction, neural diarization, and text transcription execute strictly on local mobile hardware without network egress.

Subprocess Crash Isolation

Isolates native C++ binaries within dedicated process pools, ensuring C++ segfaults never compromise the host Python or Node.js runtime.

Modular & Interactive Model Provisioning

Keeps standard installation lightweight while offering on-demand neural diarization installation (--engine diarization / --all) with interactive TTY [y/N] prompts.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Asymmetric Hybrid Pipeline Vulkan GPU Audio Encoder + ARM NEON CPU Sequential Decoder (--split-mode / --optimize-1) Production
Speech Recognition (STT) Whisper.cpp (tiny/base/small/medium/large-v3-turbo), Sherpa-ONNX (SenseVoice) Production
Speaker Diarization PyAnnote 3.0 ONNX, 3D-Speaker CAM++ 192d, TS-VAD OverlapResolver, Cosine K-Means Production
Audio Preprocessing Zero-Subprocess Wave Parser, FFmpeg Auto-Fallback, Silero-VAD Silence Filter Production
Subtitle & Data Formats SubRip (SRT), WebVTT (VTT), NIST Rich Transcription (RTTM), Structured JSON Production
Platform Runtimes Android Termux (Bionic ARM64/aarch64), Python 3.8~3.13, Node.js 16~22 Production

Canonical Usage Example

from termux_stt import create_engine

# 1. Initialize Whisper Engine with Vulkan GPU / CPU Hybrid Acceleration
engine = create_engine("whisper", model="small", lang="ko", threads=4, split_mode=True)

# 2. Transcribe Audio directly into Subtitles
result = engine.transcribe("samples/jfk_1min.wav")
print("Transcript:\n", result.text)
print("SRT Subtitles:\n", result.to_srt())

# 3. Multi-Speaker Neural Diarization (PyAnnote 3.0 + CAM++ 192d + TS-VAD)
hybrid = create_engine("hybrid", lang="ko", num_speakers=2)
diar_result = hybrid.diarize("samples/kor_diarization.wav")
for seg in diar_result.segments:
    print(f"[{seg.speaker}] ({seg.start:.1f}s -> {seg.end:.1f}s): {seg.text}")

Getting Started & Deep Guides