Benchmarks & Real-Device Profiling
Strict ground truth performance measurements captured directly on physical Samsung Galaxy S20+ hardware.
1. Hardware & System Testbed Specifications
| Property |
Hardware Specification |
Details |
| Target Device | Samsung Galaxy S20+ 5G (SM-G986N) | Physical smartphone hardware |
| SoC / Processor | Qualcomm Snapdragon 865 (SM8250) | 1x Kryo 585 Prime @ 2.84GHz, 3x Gold @ 2.42GHz, 4x Silver @ 1.80GHz |
| OS & Runtime | Android 13 / Termux Bionic ARM64 | Linux Kernel 4.19.113-android11-9-27083162 |
| SIMD Extensions | NEON, FP16, DotProd (ARMv8.2-A) | Hardware accelerated INT8 dot product and FP16 arithmetic |
| RAM Footprint | 12GB LPDDR5 Physical Memory | 3.88 GB free at cold start |
2. Empirical Inference Throughput (Llama 3.2 3B Instruct Q4_K_M)
| Benchmark Phase |
Latency / Duration |
Throughput |
System Status |
| Cold Model Loading |
1,840 ms |
1.04 GiB/s Direct Sequential Load |
1.92 GiB weights loaded with --no-mmap |
| Prompt Evaluation (Cold) |
2,347.19 ms (38 tokens) |
16.19 tokens / sec (61.77 ms/token) |
4 threads pinned on big.LITTLE cores |
| Token Generation (Eval) |
293.26 ms (3 tokens) |
10.23 tokens / sec (97.75 ms/token) |
Real-time interactive generation |
| Cached Prefix Request |
804.70 ms (10 tokens total) |
11.08 tokens / sec (90.24 ms/token) |
Prompt Cache Reuse (Similarity: 0.816) |
3. Comparison: Native Prebuilt vs Scratch Compilation
| Evaluation Metric |
Legacy Scratch Compilation |
Prebuilt Zero-Compilation (v1.0.0) |
Improvement Factor |
| Installation Time |
18 ~ 25 minutes |
< 3.0 seconds |
> 400x Faster |
| Compiler Disk Space |
~ 2.8 GB (Clang, CMake, Ninja, Git) |
0.0 MB (No build toolchains required) |
100% Space Saved |
| Thermal Load |
Heavy CPU throttling (>55°C) |
Zero compile thermal overhead (<32°C) |
Battery & Thermal Safe |