Complete API Reference
100% Full Class, Struct, and Method Documentation
1. Top-Level Engine Factory
| Function | Parameters | Return Type | Description |
create_engine |
engine="whisper"|"sherpa"|"hybrid", model=None, lang="ko", threads=None, vad=True, quantization="q5_1", num_speakers=0, **kwargs |
Engine |
Instantiates and returns the configured concrete STT or Hybrid engine instance. |
2. Core Engine Methods (Engine Abstract Base)
| Method | Signature | Return Type | Description |
transcribe |
(audio_path: str, **kwargs) |
TranscriptResult |
Transcribes audio file with subprocess crash isolation and returns full text with timestamped segments. |
diarize |
(audio_path: str, num_speakers: int = 2, **kwargs) |
DiarizedResult |
Runs the full neural diarization pipeline (PyAnnote 3.0 + CAM++ 192d + TS-VAD + Whisper STT). |
stream_mic |
(duration: Optional[float] = None) |
Iterator[Segment] |
Streams live audio from the device microphone and yields transcribed segments in real time. |
stream_file |
(audio_path: str, chunk_sec: float = 5.0) |
Iterator[Segment] |
Streams audio from a file in chunks and yields transcription segments sequentially. |
get_info |
() |
Dict[str, Any] |
Returns engine status metadata, active model name, language, and hardware thread configuration. |
3. Data Structures & Subtitle Formatters
TranscriptResult Struct:
text: str - Full reconstructed transcription text.
language: str - Detected or configured ISO 639-1 language code.
segments: List[Segment] - Timestamped utterance segments.
duration: float - Total audio duration in seconds.
to_srt() -> str - Serializes segments into valid SubRip (SRT) format.
to_vtt() -> str - Serializes segments into WebVTT format.
to_json() -> str - Serializes transcript and segment offsets into valid JSON.
to_rttm() -> str - Serializes segments into NIST Rich Transcription (RTTM) format.
DiarizedResult Struct:
- Extends
TranscriptResult.
speakers: Optional[List[str]] - Unique list of identified speaker labels (e.g. ["Speaker_0", "Speaker_1"]).
Segment Struct:
start: float - Start timestamp in seconds (precision: 0.001s).
end: float - End timestamp in seconds.
text: str - Transcribed text within the interval.
speaker: Optional[str] - Assigned speaker identifier (e.g. "Speaker_0").
confidence: Optional[float] - Transcription confidence score.