Advanced Parameters & Tuning

Kernel-level tuning, buffer pool sizing, and thread configuration

1. Whisper.cpp GGML Quantization Profiles

Quantization LevelMemory OverheadRelative SpeedRecommended Use Case
f16100% (Baseline)1.0xReference accuracy on desktop and high-end ARM nodes.
q8_055%1.3xNear-lossless precision with reduced RAM consumption.
q5_1 / q5_035%1.8xOptimal mobile default: High accuracy within minimal RAM.
q4_028%2.2xExtreme memory constraints (devices with <2GB total RAM).

2. Advanced Decoding & Search Parameters

CLI FlagSDK KeyDefaultDescription
--beam-size <N> beam_size 1 Expands beam search width for complex audio with overlapping speakers.
--temperature <T> temperature 0.0 Sampling temperature. 0.0 enables deterministic greedy search.
--prompt <TEXT> prompt None Provides vocabulary context hints (acronyms, domain terms, speaker names).
--translate / -tr translate False Translates multilingual foreign speech directly into English text.