Termux-Train
Unified Multimodal On-Device Deep Learning & LoRA Training Framework for Android Termux with 6-Modality Adapters (LLM, Diffusion, VLM, STT, TTS, BitNet), GPU Slicing, and 44GB Disaggregated Cluster Virtual RAM Pooling
Install the official package directly into your runtime:
pip install termux-train
# or via npm:
npm install -g termux-train
# or standalone bootstrap:
bash <(curl -sSL https://raw.githubusercontent.com/uno-km/termux-train/main/install.sh)
The Engineering Challenge
Standard deep learning training systems (PyTorch, DeepSpeed, Megatron) require multi-gigabyte build toolchains unavailable on Android Bionic, lack native mobile GPU autograd dispatch, and rapidly trigger the Android Low Memory Killer (LMK) during backward activation accumulation.
The Architectural Breakthrough
Termux-Train v2.0.1 delivers a dependency-free C-vectorized DAG Autograd engine, 6-Modality target-aware LoRA/DoRA training (LLM, Image Diffusion, Vision VLM, Whisper STT, TTS, BitNet 1.58-bit), Mali/Adreno GPU Slicing, 300MB Guard-Band memory protection, and 44GB disaggregated AMEVA Cluster virtual RAM pooling across heterogeneous mobile fleets.
Key Capabilities & Built-in Hardening
6-Modality Target-Aware Adapter Hub
Directly trains and exports native SafeTensors and _config.json sidecars for LLM (llama.cpp/GGUF), Image Diffusion (ComfyUI/Diffusers), Vision VLM (termux-vision), Whisper STT (termux-stt), TTS (termux-tts), and BitNet 1.58-bit ternary models without format conversion steps.
Single-Device GPU Slicing (Mali & Adreno)
Prevents mobile unified memory allocation spikes via Vocab Slicing (--vocab-slice), Chunked Layer Dispatch (--chunk-layers), and Layer Streaming (--stream-layers), preventing GPU watchdog resets.
44GB Disaggregated Cluster Virtual RAM Pooling
Aggregates physical RAM across heterogeneous mobile devices (S25, S21, S20, A53, A35) into a unified 44GB virtual memory pool with 300MB Guard-Band memory protection against Android LMK.
On-Device Reinforcement Learning (GRPO / DPO / PPO)
Executes client-side DeepSeek-R1 style Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), and PPO directly on mobile hardware.
Persistent Latent & Feature Caching
Avoids redundant VAE decoding and Mel filterbank recalculation across training epochs through deterministic SHA-256 keyed .safetensors on-device disk caching.
100% Dual-Engine Parity (Python & Node.js)
Guarantees identical subcommands, parameter names, and exit codes across Python API, Python CLI, Node.js SDK, and Node.js Global CLI.
Supported Compute Kernels & Operations
| Subsystem Category | Operations & Kernels | Status |
|---|---|---|
| LLM / Text PEFT | LoRA, DoRA, TinyTransformerLM, RoPE Rotary Embeddings, GGUF PEFT Export | Production Verified |
| Image Diffusion | DDPMScheduler, Sinusoidal Timestep MLP, CrossAttentionLoRA, ComfyUI SafeTensors | Production Verified |
| Vision VLM | MultimodalProjectorLoRA, Visual Cross-Attention, LLaVA/Qwen2-VL QA Dataset | Production Verified |
| Audio STT | Whisper Cross-Attention LoRA, Mel Spectrogram Caching, Acoustic Fine-Tuning | Production Verified |
| Audio TTS | Speaker Style Adaptation LoRA, Phonetic-to-Mel Alignment, Voice Synthesis | Production Verified |
| BitNet 1.58-bit | Ternary Quantization-Aware Training, Dynamic Activation Scaling, bitnet.cpp Export | Production Verified |
| Cluster Virtual RAM | 44GB Virtual RAM Pooling, Proportional Sharding, Distributed Pipeline Session | Production Verified |
Canonical Usage Example
import termux_train as tt
# 1. Image Diffusion LoRA Training (ComfyUI / Diffusers compatible)
tt.diffusion.train_diffusion_lora(
image_dir="./training_images",
output_path="./adapter_diffusion.safetensors",
prompt="high quality technical schematic",
resolution=512,
epochs=5,
lr=0.0001,
rank=8,
backend="vulkan"
)
# 2. Vision Multimodal VLM LoRA Training (LLaVA / Qwen2-VL format)
tt.vision.train_vision_vlm_lora(
data_source="./vlm_dataset.jsonl",
output_path="./adapter_vision.safetensors",
epochs=3,
lr=0.0002,
backend="vulkan"
)
# 3. Whisper STT LoRA Acoustic Adaptation
tt.stt.train_stt_lora(
data_source="./audio_transcripts.jsonl",
output_path="./adapter_stt.safetensors",
epochs=4,
backend="vulkan"
)