Supported Models & Architecture Zoo
Architecture compatibility, adapter layers, and memory specifications
Supported Model Architectures & Target Adaptations
| Modality | Base Architecture | Trainable Adaptation Layers | Target Runtime Compatibility | Memory Footprint (Training) |
|---|---|---|---|---|
| LLM / Text | TinyTransformerLM, LLaMA, Qwen, RoPE | Attention $W_q, W_k, W_v, W_o$, MLP Gate/Up (LoRA / DoRA) | llama.cpp, termux-llamacpp, GGUF PEFT |
120 MB - 1.8 GB |
| Image Diffusion | Stable Diffusion 1.5 / 2.1 / SDXL UNet | Cross-Attention Key/Value projections ($d_k=64/128$) | ComfyUI, Diffusers, termux-diffusion |
450 MB - 2.2 GB |
| Vision Multimodal | LLaVA 1.5, Qwen2-VL, ViT Multimodal Projector | Visual Projection MLP, Cross-Attention Adapter | termux-vision, HuggingFace Transformers |
380 MB - 1.6 GB |
| Audio STT | Whisper (tiny, base, small) Encoder-Decoder | Decoder Cross-Attention Blocks, Mel Projector | termux-stt, whisper.cpp |
290 MB - 1.2 GB |
| Audio TTS | FastSpeech / VITS / StyleTTS Acoustic Backends | Style Conditioning Linear Layers, Mel Decoder | termux-tts, Piper TTS |
180 MB - 850 MB |
| BitNet 1.58-bit | Ternary Weight Transformer $\{-1, 0, +1\}$ | Dynamic Scale Factor, Ternary Quantized Kernel | bitnet.cpp, termux-bitnet |
90 MB - 600 MB |