Multimodal On-Device Training Manual
Comprehensive engineering manual for multimodal LoRA fine-tuning
Multimodal On-Device Training Pipeline Manual
Termux-Train operates through a deterministic 4-stage pipeline designed to protect mobile system integrity while maximizing mathematical convergence.
Stage 1: Zero-Leak Data Ingestion & Persistent Feature Caching
To avoid repeated CPU/GPU bottlenecks across epochs, compute-intensive preprocessing operations are cached deterministically on disk using SHA-256 hashes as keys:
- Diffusion VAE Latents: Source images are passed through the VAE encoder once; latent representations ($C \times \frac{H}{8} \times \frac{W}{8}$) are stored in
.cache/latents/<sha256>.safetensors. - Audio Mel Filterbanks: Raw waveforms are transformed into 80-channel log-Mel spectrograms and cached on disk.
Stage 2: DAG Autograd Forward-Backward Dispatch
The pure Python/C DAG Autograd dynamically constructs computational graphs. Tensors maintain backward gradient closure functions without heap allocation overheads.
Stage 3: 300MB Guard-Band Memory Protection
During every forward-backward iteration, the runtime queries /proc/meminfo via native Bionic syscalls. If available physical RAM drops below 300MB, the engine immediately suspends batch ingestion, triggers garbage collection, and pages layer weights to prevent Android LMK termination.
Stage 4: Target-Aware SafeTensors Serialization
Weights are serialized into industry-standard SafeTensors files with explicit little-endian byte ordering. A companion _config.json sidecar is automatically generated to enable zero-configuration import into downstream inference engines.