Performance & Hardware Benchmarks
Empirical latency, throughput, and memory consumption across mobile CPUs.
1. 0-Point Baseline Production Audit Scorecard
| Pillar | Metric Target | Pure Python Backend | NumPy (OpenBLAS NEON) | Score |
| Pillar 1: Autograd & Math | < 5.0 ms | 1.13 ms | 0.99 ms | 20.0 / 20.0 pts |
| Pillar 2: Transformer & RoPE | < 2000 ms | 9179 ms | 1072 ms | 20.0 / 20.0 pts |
| Pillar 3: Memory Efficiency | < 100 ms | 65.7 ms | 20.1 ms | 20.0 / 20.0 pts |
| Pillar 4: Performance Latency | < 1000 ms | 6253 ms | 622.5 ms | 20.0 / 20.0 pts |
| Pillar 5: Checkpoint Resilience | < 50 ms | 40.5 ms | 21.2 ms | 20.0 / 20.0 pts |
| TOTAL SCORE | 100.0 / 100.0 | Grade A+ (PERFECT) |
2. LoRA Adapter Compression vs Full Checkpoint
| Model Architecture | Full PyTorch Checkpoint | termux-train SafeTensors LoRA | RAM & Storage Reduction |
| Tiny Whisper Speech-to-Text | 557.25 KB | 21.01 KB | 96.2% Reduction |
| Tiny Transformer LM (2-Layer) | 1.20 MB | 42.50 KB | 96.5% Reduction |