Quickstart & Execution Recipes

Standard usage patterns and rapid prototyping code

Included Sample Audio Benchmark Suite

termux-stt packages benchmark audio samples in samples/:

  • samples/jfk_1min.wav: English Speech Benchmark (60.00s, 16kHz Mono)
  • samples/kor_diarization.wav: Korean Real Podcast 2-Speaker Cross-Talk Benchmark (45.00s, 16kHz Mono)
  • samples/3_speaker_test.wav: 3-Speaker Multi-Conversation Benchmark (15.00s, 16kHz Mono)

Recipe 1: 1-Line Subtitle Transcription (SRT / WebVTT / JSON)

# 1. Transcribe audio and export to synchronized SRT subtitles
termux-stt transcribe samples/jfk_1min.wav --format srt > subtitles.srt

# 2. Export directly to WebVTT for HTML5 video players
termux-stt transcribe samples/jfk_1min.wav --format vtt > subtitles.vtt

# 3. Export to structured JSON for data pipelines
termux-stt transcribe samples/jfk_1min.wav --format json > output.json

Recipe 2: Neural Speaker Diarization ("Who Spoke When?")

Run high-precision multi-speaker diarization using PyAnnote 3.0 + CAM++ 192d neural embeddings:

termux-stt diarize samples/kor_diarization.wav --speakers 2 --format text

Sample Output:

[Speaker 0] (0.0s -> 5.2s): 주변에 얘기를 거의 안 했어요. 제가 사실 그 샀다고 얘기를 안 했어요.
[Speaker 1] (5.4s -> 6.1s): 왜요?
[Speaker 0] (6.5s -> 11.2s): 왜냐하면 그냥 뭐 얘기할 필요성을 잘 못 느껴서...

Recipe 3: Real-Time Microphone Streaming

Speak directly into your smartphone microphone with low latency and real-time Silero-VAD silence filtering:

# CLI real-time listener (Press Ctrl+C to terminate)
termux-stt listen --engine whisper --model tiny --lang ko

Recipe 4: Programmatic Python SDK Integration

from termux_stt import create_engine

# 1. Initialize Hybrid Engine
engine = create_engine("hybrid", model="small", lang="ko", num_speakers=2)

# 2. Transcribe and Diarize
result = engine.diarize("samples/kor_diarization.wav")

for seg in result.segments:
    print(f"[{seg.speaker}] {seg.start:.2f}s - {seg.end:.2f}s: {seg.text}")

Recipe 5: Programmatic Node.js / TypeScript API

const { createEngine } = require("termux-stt");

async function diarizeMeeting() {
  const engine = createEngine("hybrid", { model: "small", lang: "ko", numSpeakers: 2 });
  const result = await engine.diarize("samples/kor_diarization.wav");
  
  console.log("Speakers:", result.speakers);
  console.log("SRT Subtitles:\n", result.toSrt());
}
diarizeMeeting();