AMEVA-Runtime

Next-Gen Unified On-Device Hardware Orchestration & 6-Modality Acceleration Runtime for Mobile & Edge (with 1:1 Modality Library Isolation & Adreno OpenCL Smart Routing)

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install ameva-runtime
# or: npm install @ameva/runtime

The Engineering Challenge

Deploying generative multi-modal AI on mobile hardware encounters severe vendor-specific compiler bugs and driver anomalies. Qualcomm Adreno GPUs suffer from JIT compiler crashes during specialization constant dispatch, while ARM Mali drivers encounter compute shader integer truncation infinite loops and headless fence deadlocks without active display swapchains.

The Architectural Breakthrough

AMEVA Runtime is a unified on-device hardware orchestration and AI acceleration engine engineered for mobile and edge systems. It dynamically inspects SoC topology and driver environments, routing compute graphs between Qualcomm Adreno, ARM Mali, and ARM Cortex CPU-NEON. ### 1. Measured LLM Inference on Devices (Qwen2.5-0.5B-Instruct, GGUF Q4_K_M) Measured 2026-10-04 with `llama-completion` of the v2.8.0 llama bundle, 4 threads; the GPU route is the bundle's ggml-ameva backend with all layers on the GPU. After a warm-up run, two prompt runs and three generation runs alternate with the CPU. Prompt processing is a prompt of about 370 tokens evaluated at once; token generation is 64 tokens, one at a time. The text the GPU route generates is the CPU's on all five devices. | Device & Processor | Default route | Prompt processing, CPU / GPU (t/s) | Token generation, CPU / GPU (t/s) | | :--- | :---: | :---: | :---: | | **Galaxy S25** (Snapdragon 8 Elite, Adreno 830) | GPU | 30-58 / 1010-1081 | 16-49 / 36-77 | | **Galaxy S20** (Snapdragon 865, Adreno 650) | GPU | 60 / 181-197 | 26-27 / 23.6-24.4 | | **Galaxy A35** (Exynos 1380, Mali-G68) | GPU | 109-110 / 190 | 35-37 / 21.5-21.6 | | **Galaxy S21** (Exynos 2100, Mali-G78) | CPU | 115-119 / 227 | 46-47 / 22.8-23.1 | | **Galaxy A53** (Exynos 1280, Mali-G68) | CPU | 55 / 67 | 22-24 / 14.5 | Galaxy S25 was running other work during the measurement, hence its wide ranges. Galaxy S20 was measured at a GPU clock cap of 441 MHz; at 587 MHz its token generation was 28-31 on the CPU and 32.6-32.9 on the GPU. On the Mali devices token generation is faster on the CPU, and Galaxy S21 and A53, whose drivers lack integer dot products, take the CPU route by default. Larger models and the measurement conditions are in the [v2.8.0 release notes](https://github.com/uno-km/ameva-runtime-releases/releases/tag/v2.8.0). ### 2. Empirical Real-Device STT Benchmarks (Whisper Large-v3-Turbo Q5_0, 548MB) - **Test Device**: Samsung Galaxy A35 5G (Exynos 1380, ARM Mali-G68 MP5, 8GB RAM, Android 16 Termux) - **Audio Source**: John F. Kennedy 1-minute speech sample (`jfk_1min.wav`) | Execution Mode | Target Hardware | Elapsed Time | GPU Clock / Load | CPU Utilization | Accuracy | Speedup | | :--- | :--- | :---: | :---: | :---: | :---: | :---: | | **CPU NEON Mode** (`-dev -1`, 4 threads) | Cortex-A78 x4 cores | **816.48s (13m 36s)** | 0% (Idle) | 291% (Active) | Standard | Baseline | | **Vulkan GPU Mode** (`-dev 0`, Mali Quirk) | Mali-G68 MP5 | **360.60s (6m 00s)** | **949 MHz (100%)** | **20~30% (Low)** | Standard | **2.26x (56% time reduction)** | ### 3. Root-Cause Defect Resolution (Ground Truth) #### (1) ARM Mali-G68 Valhall Integer Truncation Infinite Loop Elimination - **Defect**: Executing `mul_mm.comp` on Mali-G68 (subgroup size 16) caused GPU hangs and hardware watchdog TDR resets (`VK_ERROR_DEVICE_LOST`). - **Root Cause**: The stride calculation `loadstride_b = gl_WorkGroupSize.x * LOAD_VEC_B / BK = 16 * 1 / 32 = 0` truncated to zero in integer division, producing an infinite loop `for (uint l = 0; l < BN; l += 0)`. - **Resolution**: Enforced Medium MatMul kernels (`_m`, workgroup size 128, `loadstride_b = 4 > 0`) via `enforce_medium_matmul: true`, enabling stable 25/25 layer GPU offloading. #### (2) Qualcomm Adreno 830 JIT Compiler Bug Isolation - **Defect**: Whisper STT pipeline compilation failed on Snapdragon 8 Elite with `VK_ERROR_UNKNOWN (-13)` during `mul_mat_vec` dispatch. - **Root Cause**: Qualcomm's Adreno JIT compiler failed register allocation when Specialization Constant `NUM_COLS >= 3`. - **Resolution**: Bound `mul_mat_vec_max_cols = 2` for Adreno 830, achieving stable GPU inference in 4,401 ms on speech input. ### 4. Single Package Architecture - Consolidated under `pip install ameva-runtime` and `npm install @ameva/runtime`. - Specialized Vulkan acceleration exposed via `from ameva_runtime import vulkan`.

Key Capabilities & Built-in Hardening

SmartRouter Silicon Topology Dispatch

Analyzes device nodes (/dev/kgsl-3d0, /dev/mali0) and cgroups to dispatch between Qualcomm Adreno and ARM Mali.

Native Adreno 830 Full Offloading

Runs all LLM layers on the GPU of Snapdragon 8 Elite through the bundled ggml-ameva backend: Qwen2.5-0.5B prompts at 1010-1081 tokens/sec (CPU 30-58) and generation at 36-77 tokens/sec (CPU 16-49) with the CPU's text (measured with v2.8.0 while the device ran other work).

Mali Valhall MatMul Quirk Engine

On Mali the GPU route runs on ggml-ameva with kernel layouts measured per device and weight shape. Galaxy A35 (Mali-G68): Qwen2.5-0.5B prompts at 190 tokens/sec (CPU 109-110) and generation at 21.5 tokens/sec (CPU 35-37) with the CPU's text; Llama-3.2-1B and Gemma-2-2B run as well (v2.8.0 release notes).

Whisper STT 2.26x Hardware Acceleration

Accelerates Whisper Large-v3-Turbo on mobile GPU (360s vs 816s CPU, 56% time reduction) while unburdening CPU.

Z-Image Turbo 6.0B DiT Vulkan Compute

Dispatches SOTA Diffusion Transformer compute with Flash Attention and VAE tiling directly on mobile Qualcomm Adreno and ARM Mali GPUs.

Two-Track Enterprise Supply Chain Architecture

Decouples proprietary source code into private repositories while providing verified release binaries via public registry for tokenless 1-Click provisioning.

Unified Multi-Modality Architecture

Drives GGUF LLM, Whisper STT, LLaVA Vision, Stable Diffusion / DiT, Piper TTS, and BitNet from a single runtime.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Galaxy S25 (Qualcomm Adreno 830) Vulkan through ggml-ameva, all LLM layers on the GPU: prompt 1010-1081 t/s (CPU 30-58), generation 36-77 t/s (CPU 16-49) for Qwen2.5-0.5B Production
Galaxy A35 (ARM Mali-G68) Vulkan through ggml-ameva: prompt 190 t/s (CPU 109-110), generation 21.5 t/s (CPU 35-37) for Qwen2.5-0.5B; Whisper Large-v3-Turbo (2.26x) Production
Mobile DiT Diffusion (Z-Image Turbo 6.0B) Vulkan DiT Pipeline, Flash Attention (--diffusion-fa), VAE Tiling Production (v2.7.4)
Two-Track Public Registry Anonymous 1-Click Auto-Provisioning via uno-km/ameva-runtime-releases Active
Multi-Modality Engines GGUF LLM, Whisper STT, LLaVA Vision, Stable Diffusion / DiT, Piper TTS, BitNet Production

Canonical Usage Example

import ameva_runtime as ameva
from ameva_runtime import vulkan

# 1. Execute LLM inference directly with optimal on-device hardware dispatch
result = ameva.run(
    model="qwen2.5-0.5b",
    prompt="Space in Korean is:",
    max_tokens=32
)
print(f"Generated text: {result.text}")
print(f"Hardware backend: {result.backend_used} ({result.tokens_per_second:.2f} t/s)")

# 2. Hardware diagnostic inspection via Vulkan engine
doc = vulkan.Doctor()
report = doc.run_self_test(verbose=False)
print(f"GPU Target: {report.device_name} (Passed: {report.passed_stages}/{report.total_stages})")

Getting Started & Deep Guides