Model Hub & GGUF Presets
Curated collection of quantized diffusion models optimized for Android ARM64 architecture, Samsung Galaxy hardware, and Termux memory budgets.
π― Model Selection Strategy for Mobile Devices
For instant interactive applications, use sdxs (generates in ~4s on Galaxy A35/S21). For high-fidelity artwork and portraits, use anime (LCM) or realistic (Realistic Vision) with the --vae-tiling flag enabled.
1. Official Model Presets (곡μ μ¬μ νμ¬ ν리μ )
| Preset ID | Architecture / Base | GGUF Checkpoint | Size | Optimal Steps | Optimal CFG | A35 Latency | VAE Tiling |
|---|---|---|---|---|---|---|---|
z-image-turbo Flagship DiT |
6.0B Diffusion Transformer + Qwen3-4B LLM | z_image_turbo-Q2_K.gguf (DiT)Qwen3-4B-Instruct-2507-Q2_K.gguf (LLM)taef1.safetensors (TAESD) |
2.41 GB (1.0 GB VRAM Cap) |
8 | 1.0 (Fixed) | Vulkan GPU (S21 5G Verified) | TAESD Flux |
sdxs (Default) |
Tiny-SD (1-Step Distilled) | sdxs-512-tinySDdistilled_Q8_0.gguf |
651 MB | 1 ~ 2 | 1.0 (Fixed) | 4.08s (256p) 23.2s (512p) |
Optional (TAESD) |
anime |
DreamShaper 8 (LCM Accelerated) | DreamShaper8_LCM_q4_0.gguf |
1,550 MB | 4 ~ 6 | 1.5 ~ 2.0 | 422s (6-step) | Required |
realistic |
Realistic Vision V6.0 B1 | realisticVisionV60B1_v51HyperVAE-Q4_k.gguf |
1,547 MB | 6 ~ 8 | 3.5 ~ 4.5 | 447s (6-step) | Required |
speed |
SD 1.5 Turbo Base | sd15-turbo-q4_0.gguf |
1,420 MB | 4 ~ 5 | 2.0 ~ 3.0 | ~380s (5-step) | Required |
2. Golden Rules by Model Architecture (μν€ν μ²λ³ νμ μ§μΉ¨)
2.0 Z-Image Turbo (6.0B DiT + Qwen3-4B LLM) β Sovereign Mobile Diffusion Transformer
- Tri-Engine Principle: Combines a 4.0B LLM for deep natural language prompt understanding, a 6.0B Diffusion Transformer for global optical coherence, and a 10MB TAESD for 1-second latent reconstruction.
- Golden Rule: Enable Layer Streaming and VRAM Capping: Always invoke with
--stream-layers,--max-vram vulkan0=1, and--params-backend diffusion=cpu. This guarantees the 2.41GB DiT model stays strictly within 1.0 GB VRAM, preventing Android LMK kills (Zero-OOM). Use--steps 8with Euler ODE solver and CFG 1.0 for physical PBR fidelity. - CLI Example:
sd-cli-vulkan -p "A cinematic photo of a neon cybernetic tiger walking in Seoul street at night" -W 512 -H 512 -t 4 --steps 8 --cfg-scale 1.0 --sampling-method euler --backend clip=cpu,diffusion=vulkan0,vae=cpu --max-vram vulkan0=1 --stream-layers --params-backend diffusion=cpu --diffusion-fa --mmap --vae-tiling --vae-format flux -o cyber_tiger.png - Research Whitepaper: Detailed micro-kernel analysis and Samsung Galaxy S21 5G verification traces are documented in the AMEVA Labs Research Portal.
2.1 SDXS (512-tinySDdistilled / 651 MB) β Ultra-Fast Mobile Distillation
- Distillation Principle: Recovers full latent distribution in 1~2 steps via advanced score distillation. TAESD decoder completes VAE stage in under 0.2s.
- Golden Rule: Never set CFG Scale above 2.0. CFG 1.0 with Euler A sampler yields the sharpest, uncorrupted output. 2nd-order solvers (Heun, DPM2) destroy the distilled manifold.
- CLI Example:
termux-diffusion generate "a small robot on workbench" -m sdxs -W 256 -H 256 -s 1 --cfg-scale 1.0
2.2 DreamShaper 8 LCM (1.55 GB) β 2D/2.5D Illustration
- LCM Principle: Latent Consistency Model formulation reduces standard 20+ steps down to 4~6 steps with vibrant cel-shading and particle effects.
- Golden Rule: Use
--sampler lcmwith--schedule karrasordefault. Always enable--vae-tilingto prevent Android low-memory kills. - CLI Example:
termux-diffusion generate "anime warrior girl, cyber armor, 8k" -m anime -s 4 --cfg-scale 1.5 --sampling-method lcm --vae-tiling
2.3 Realistic Vision V6.0 B1 (1.55 GB) β Photorealism & Portraits
- Burnout Prevention: High CFG (6.5+) causes color burn and dark vein artifacts. Set CFG to 3.5 ~ 4.5 for pristine, translucent skin tones.
- Prompt Strategy: Avoid literal keyword
'skin pores'(causes clustered hole artifacts). Instead, use lighting and lens terms:85mm f1.4 lens, natural skin texture, soft morning window light. - Negative Defense: Always include
(deformed iris, deformed pupils:1.3), plastic skin, doll, cgi, 3d, renderto eliminate uncanny valley effects. - CLI Example:
termux-diffusion generate "portrait of 24yo woman, cozy sweater, 85mm f1.4" -n "low quality, cgi, doll" -m realistic -s 6 --cfg-scale 4.0 --vae-tiling
3. GGUF Quantization Formats & Memory Matrix
| Quantization Type | Bits Per Weight (bpw) | Model RAM Footprint | Perplexity / Quality Loss | Recommended Device Tier |
|---|---|---|---|---|
Q8_0 |
8.50 bpw | ~650 MB (SDXS) / ~2.2 GB (SD1.5) | < 0.05% (Near Lossless) | All Devices (SDXS) / 8GB+ RAM |
Q4_K |
4.50 bpw | ~1.55 GB (SD1.5 Full) | < 0.80% (High Visual Fidelity) | 6GB RAM (Galaxy A35) & 8GB RAM |
Q4_0 |
4.00 bpw | ~1.50 GB (LCM Full) | < 1.20% (Fast Computation) | 6GB RAM (Galaxy A35) & 8GB RAM |