Benchmarks & Profiling

Deterministic latency and VRAM allocation statistics

Hardware Testbed Specifications

All benchmarks were conducted on physical hardware testbeds: Samsung Galaxy A35 (Exynos 1380, 6GB RAM, Android 14 Termux) and Samsung Galaxy S21 (Snapdragon 865, 8GB RAM, Android 15 Termux) using physical Korean podcast and JFK inauguration speech audio.

1. Real Korean Podcast 2-Speaker Cross-Talk Benchmark (45.00s Audio)

Hardware DevicePipeline StageAlgorithm / ModelExecution LatencyAccuracy / DERPeak RSS
Galaxy A35 (Exynos 1380) Neural Diarization PyAnnote 3.0 ONNX + CAM++ 192d 16.91s DER < 8.5% (High Precision) 142.3 MB
Galaxy A35 (Exynos 1380) Overlap Resolution TS-VAD Multi-Label Cross-Talk Scanner 0.42s "왜요?" Interrupt Decoupled 18.5 MB
Galaxy A35 (Exynos 1380) Speech-to-Text Whisper.cpp Small (Korean) 34.12s WER < 5.2% 42.1 MB
Full End-to-End Pipeline 51.45s (RTF 1.14x) Production Validated 158.2 MB

2. Whisper Model Tier On-Device Performance Matrix (Galaxy S21 / S25)

Hardware DeviceModel TierParametersModel SizeInference Time (60s Audio)Real-Time Factor (RTF)Peak RAMWER Status
Galaxy S25 (SD 8 Elite) Whisper Small (Hybrid -sm) 244M 466 MB 12.93s 0.2133x (4.7x Faster) 48.5 MB PASS (SOTA)
Galaxy S21 (SD865) Whisper Base 74M 142 MB 12.46s 0.2077x (5x Faster) 25.6 MB PASS
Galaxy A35 (Exynos 1380) Whisper Tiny 39M 75 MB 22.36s 0.3726x (2.7x Faster) 5.75 MB PASS