Supported Models & Architecture Zoo

Architecture compatibility, adapter layers, and memory specifications

Supported Model Architectures & Target Adaptations

Modality Base Architecture Trainable Adaptation Layers Target Runtime Compatibility Memory Footprint (Training)
LLM / Text TinyTransformerLM, LLaMA, Qwen, RoPE Attention $W_q, W_k, W_v, W_o$, MLP Gate/Up (LoRA / DoRA) llama.cpp, termux-llamacpp, GGUF PEFT 120 MB - 1.8 GB
Image Diffusion Stable Diffusion 1.5 / 2.1 / SDXL UNet Cross-Attention Key/Value projections ($d_k=64/128$) ComfyUI, Diffusers, termux-diffusion 450 MB - 2.2 GB
Vision Multimodal LLaVA 1.5, Qwen2-VL, ViT Multimodal Projector Visual Projection MLP, Cross-Attention Adapter termux-vision, HuggingFace Transformers 380 MB - 1.6 GB
Audio STT Whisper (tiny, base, small) Encoder-Decoder Decoder Cross-Attention Blocks, Mel Projector termux-stt, whisper.cpp 290 MB - 1.2 GB
Audio TTS FastSpeech / VITS / StyleTTS Acoustic Backends Style Conditioning Linear Layers, Mel Decoder termux-tts, Piper TTS 180 MB - 850 MB
BitNet 1.58-bit Ternary Weight Transformer $\{-1, 0, +1\}$ Dynamic Scale Factor, Ternary Quantized Kernel bitnet.cpp, termux-bitnet 90 MB - 600 MB