Benchmarks & Profiling
Deterministic latency and VRAM allocation statistics
Hardware Testbed Specifications
All benchmarks were conducted on physical hardware testbeds: Samsung Galaxy A35 (Exynos 1380, 6GB RAM, Android 14 Termux) and Samsung Galaxy S21 (Snapdragon 865, 8GB RAM, Android 15 Termux) using physical Korean podcast and JFK inauguration speech audio.
1. Real Korean Podcast 2-Speaker Cross-Talk Benchmark (45.00s Audio)
| Hardware Device | Pipeline Stage | Algorithm / Model | Execution Latency | Accuracy / DER | Peak RSS |
|---|---|---|---|---|---|
| Galaxy A35 (Exynos 1380) | Neural Diarization | PyAnnote 3.0 ONNX + CAM++ 192d | 16.91s | DER < 8.5% (High Precision) | 142.3 MB |
| Galaxy A35 (Exynos 1380) | Overlap Resolution | TS-VAD Multi-Label Cross-Talk Scanner | 0.42s | "왜요?" Interrupt Decoupled | 18.5 MB |
| Galaxy A35 (Exynos 1380) | Speech-to-Text | Whisper.cpp Small (Korean) | 34.12s | WER < 5.2% | 42.1 MB |
| Full End-to-End Pipeline | 51.45s (RTF 1.14x) | Production Validated | 158.2 MB | ||
2. Whisper Model Tier On-Device Performance Matrix (Galaxy S21 / S25)
| Hardware Device | Model Tier | Parameters | Model Size | Inference Time (60s Audio) | Real-Time Factor (RTF) | Peak RAM | WER Status |
|---|---|---|---|---|---|---|---|
| Galaxy S25 (SD 8 Elite) | Whisper Small (Hybrid -sm) | 244M | 466 MB | 12.93s | 0.2133x (4.7x Faster) | 48.5 MB | PASS (SOTA) |
| Galaxy S21 (SD865) | Whisper Base | 74M | 142 MB | 12.46s | 0.2077x (5x Faster) | 25.6 MB | PASS |
| Galaxy A35 (Exynos 1380) | Whisper Tiny | 39M | 75 MB | 22.36s | 0.3726x (2.7x Faster) | 5.75 MB | PASS |