Advanced Parameters & Tuning
Kernel-level tuning, buffer pool sizing, and thread configuration
1. Whisper.cpp GGML Quantization Profiles
| Quantization Level | Memory Overhead | Relative Speed | Recommended Use Case |
f16 | 100% (Baseline) | 1.0x | Reference accuracy on desktop and high-end ARM nodes. |
q8_0 | 55% | 1.3x | Near-lossless precision with reduced RAM consumption. |
q5_1 / q5_0 | 35% | 1.8x | Optimal mobile default: High accuracy within minimal RAM. |
q4_0 | 28% | 2.2x | Extreme memory constraints (devices with <2GB total RAM). |
2. Advanced Decoding & Search Parameters
| CLI Flag | SDK Key | Default | Description |
--beam-size <N> |
beam_size |
1 |
Expands beam search width for complex audio with overlapping speakers. |
--temperature <T> |
temperature |
0.0 |
Sampling temperature. 0.0 enables deterministic greedy search. |
--prompt <TEXT> |
prompt |
None |
Provides vocabulary context hints (acronyms, domain terms, speaker names). |
--translate / -tr |
translate |
False |
Translates multilingual foreign speech directly into English text. |