Multimodal On-Device Training Manual

Comprehensive engineering manual for multimodal LoRA fine-tuning

Multimodal On-Device Training Pipeline Manual

Termux-Train operates through a deterministic 4-stage pipeline designed to protect mobile system integrity while maximizing mathematical convergence.

Stage 1: Zero-Leak Data Ingestion & Persistent Feature Caching

To avoid repeated CPU/GPU bottlenecks across epochs, compute-intensive preprocessing operations are cached deterministically on disk using SHA-256 hashes as keys:

Stage 2: DAG Autograd Forward-Backward Dispatch

The pure Python/C DAG Autograd dynamically constructs computational graphs. Tensors maintain backward gradient closure functions without heap allocation overheads.

Stage 3: 300MB Guard-Band Memory Protection

During every forward-backward iteration, the runtime queries /proc/meminfo via native Bionic syscalls. If available physical RAM drops below 300MB, the engine immediately suspends batch ingestion, triggers garbage collection, and pages layer weights to prevent Android LMK termination.

Stage 4: Target-Aware SafeTensors Serialization

Weights are serialized into industry-standard SafeTensors files with explicit little-endian byte ordering. A companion _config.json sidecar is automatically generated to enable zero-configuration import into downstream inference engines.