AmfyUI: ComfyUI Mobile DAG Studio & Sovereign DiT Showcase
Official user manual, architecture specification, and Galaxy S20 1280×720 HD 4-Season empirical benchmarks for AmfyUI — the native Android ComfyUI DAG runtime.
AmfyUI eliminates the multi-gigabyte Python/PyTorch desktop footprint. Built directly on top of Android Bionic libc and native ARM64 NEON / Qualcomm Adreno OpenCL compute kernels, AmfyUI executes full ComfyUI JSON DAG node graphs natively on your smartphone with deterministic VRAM caps and Zero-LMK stability.
1. Architectural Nomenclature: What is AmfyUI?
AmfyUI is an engineered recursive backronym and structural design protocol tailored for sovereign edge computing:
--max-vram) and dynamic layer streaming preventing Android Low Memory Killer (LMK) aborts.2. ComfyUI Compatibility Matrix & Operational Boundaries
AmfyUI interprets and executes official ComfyUI JSON workflows directly without running heavy desktop Python daemons.
| ComfyUI Node Class | AmfyUI Native Mapping | Execution Backend | Status |
|---|---|---|---|
CLIPLoader / DualCLIPLoader |
Qwen3-4B / CLIP ViT-L / T5-XXL (GGUF / FP8) | CPU NEON Multi-threading (4-8 Cores) | 100% Native |
UNETLoader / DiffusionModelLoader |
Z-Image Turbo 6.0B DiT, SDXS, SD 1.5, SDXL | Adreno OpenCL 2.0 / Mali Vulkan 1.1 / CPU SDOT | 100% Native |
VAELoader / TAESDLoader |
TAESD FLUX.1 (10MB) & Standard SD VAE (FP16) | CPU NEON Vectorized Decode (Zero-VRAM Spike) | 100% Native |
CLIPTextEncode |
Prompt Conditioning & Cross-Attention Vectorizer | Host RAM Shared Context | 100% Native |
EmptyLatentImage |
Spatial Latent Tensor Allocator (up to 1280×720) | Zero-Copy Buffer Pool | 100% Native |
KSampler / KSamplerAdvanced |
Res_Multistep, Euler, Euler_A, DPM++ 2M, LCM | Tiled Flash Attention ODE Solver | 100% Native |
VAEDecode / VAEEncode |
Latent-to-RGB Reconstruction / Img2Img Tiler | Spatial Tiled Decoder (--vae-tiling) |
100% Native |
SaveImage |
PNG Export & Android MediaStore Broadcaster | Samsung Gallery Auto-Sync (Direct DCIM/Pictures) | 100% Native |
LoraLoader |
Dynamic Weight Additive Patching | GGML Layer Merging | Supported |
Custom Python C++ Nodes |
External desktop extensions requiring GCC/CUDA | Fallback to AMEVA-Cluster RPC Offloading | Cluster Offload |
3. Visual Node Orchestration & User Interface
AmfyUI runs an interactive web studio right inside Android Chrome or Samsung Internet. Launch the server from Termux terminal:
# 1. Launch AmfyUI Mobile Studio (Default: host=0.0.0.0, port=11553)
# Automatically binds to 0.0.0.0:11553 (both local and external LAN/Tailscale access enabled)
termux-diffusion ui
# Access URLs:
# - On phone: http://localhost:11553
# - From PC/LAN: http://<phone-ip>:11553 (e.g. http://100.106.99.81:11553)
# 2. Optional: Custom port & host binding:
termux-diffusion ui --port 11554 --host 0.0.0.0
4. Galaxy S20 (Snapdragon 865) 1280×720 HD 4-Season Master Benchmark
Physical test run (Jira SCRUM-481) executed on Samsung Galaxy S20 5G (Qualcomm Snapdragon 865 KONA, 12GB LPDDR5 RAM, Adreno 650 GPU OpenCL 2.0). Generating wide 16:9 high-definition 1280×720 portrait photography across four seasonal conditions.
- 2.35× Hardware Acceleration via OpenCL Flash Attention: GPU execution reduced 8-step inference latency from 9,022s (CPU baseline) down to 3,835s.
- Cortex-A77 SDOT Hardware Advantage: On CPU, Q8_0 (8,534s) executed 488 seconds faster than Q4_0 (9,022s) because ARMv8.2-A hardware
SDOTinstructions execute 4 FMAs per cycle without software bit-shifting overhead. - Thermal Efficiency: OpenCL GPU execution operated at 41.4°C–41.8°C, which is ~6°C cooler than full CPU saturation (47.8°C).
- Tensor Forensic Integrity: 32 individual checkpoint tensors across 4 stages recorded exactly NaN=0, Zero=0.
| Season & Stage | Weight Quant | Execution Backend | Total Elapsed | Step Latency | GPU Speedup | Peak Thermal | Forensic Sanity |
|---|---|---|---|---|---|---|---|
| 🌸 1. Spring (봄) | Q4_0 (3.53 GB) | CPU (ARM NEON) | 9,022s (2h 30m) | ~1,180s / step | 1.00× (Baseline) | 47.8°C | NaN=0 / Zero=0 |
| ☀️ 2. Summer (여름) | Q4_0 (3.53 GB) | Adreno OpenCL 2.0 | 3,835s (1h 03m) | ~468s / step | 2.35× Accelerated | 41.4°C | NaN=0 / Zero=0 |
| 🍂 3. Autumn (가을) | Q8_0 (6.13 GB) | CPU (ARM SDOT) | 8,534s (2h 22m) | ~1,060s / step | 1.06× (vs CPU Q4) | 47.1°C | NaN=0 / Zero=0 |
| ❄️ 4. Winter (겨울) | Q8_0 (6.13 GB) | Adreno OpenCL 2.0 | 3,895s (1h 04m) | ~475s / step | 2.19× Accelerated | 41.8°C | NaN=0 / Zero=0 |
5. 4-Season 1280×720 HD Portrait Gallery & Prompt Engineering
Directly rendered on Samsung Galaxy S20 without upscaling. All images are 1280×720 24-bit TrueColor PNG artifacts generated using 8-step Z-Image Turbo with TAESD VAE.
6. ComfyUI Workflow JSON Download & Direct CLI Reproduction
Directly load these workflow files into AmfyUI Mobile Studio or execute them unattended via CLI:
Executing Workflow via Terminal CLI
# Run workflow file end-to-end with real-time progress
termux-diffusion workflow run workflows/test2_s20_opencl_winter.json \
--device opencl \
--output-dir /sdcard/Pictures/TermuxDiffusion
# Validate and inspect workflow topology before execution
termux-diffusion workflow inspect workflows/test1_s20_cpu_spring.json
Executing Workflow via Python SDK
from termux_diffusion.workflow import WorkflowExecutor
# Execute ComfyUI DAG programmatically
executor = WorkflowExecutor()
result_image = executor.execute_file(
"workflows/test2_s20_opencl_winter.json",
device="opencl",
output_dir="/sdcard/Pictures/TermuxDiffusion"
)
print(f"Generated 1280x720 TrueColor PNG: {result_image}")