AMEVA-Cluster
Symmetric Disaggregated Mobile RAM Pooling & On-Device AI Acceleration Runtime with Compute-Memory Decoupling
Install the official package directly into your runtime:
pip install ameva-cluster
# or
npm install -g @ameva/cluster
The Engineering Challenge
Executing frontier 70B scale models on standalone mobile devices inevitably triggers Android Low Memory Killer (LMK) termination and thermal throttling. Conventional distributed inference forces compute on all nodes, overheating lower-tier battery-powered devices.
The Architectural Breakthrough
AMEVA-Cluster unifies idle LPDDR memory across multiple heterogeneous mobile smartphones into a single disaggregated memory pool using a symmetric master/worker binary. Compute is fully isolated on a single flagship GPU, while edge phones serve strictly as passive RAM caches, completely preventing thermal throttling and Android Low Memory Killer (LMK) SIGKILL terminations. Features built-in storage-assisted hybrid failover to local UFS 3.1 flash mmap streaming and fast-path HMAC mutual authentication Guard proxies.
Key Capabilities & Built-in Hardening
Symmetric Single Binary CLI
All-in-one architecture supporting both master and worker roles via single CLI flag.
Compute-Memory Decoupling
Worker nodes execute zero GPU passes, acting solely as RAM caches to prevent heat and failure.
300MB Safety Guard-Band
Automatically deducts buffer memory from usable pool to protect against Android LMK kills.
Mutual Authentication Guard Proxy
0.001s Fail-Fast HMAC-SHA256 verification rejecting unauthorized RPC traffic and DoS attempts.
Storage-Assisted Hybrid Failover
Gracefully degrades to local UFS 3.1 flash streaming upon worker disconnects without crashing.
6-Modality Pluggable Engine
Standardized distributed tensor sharding across LLM, Diffusion, Vision, STT, TTS, and Training.
Supported Compute Kernels & Operations
| Subsystem Category | Operations & Kernels | Status |
|---|---|---|
| Compute Engine | WebGPU Compute Shaders (WGSL), FP16/FP32 | Production |
| Memory Subsystem | Zero-Copy Ring Buffers, Weakref GC Pooling | Production |
| Platform Runtimes | Node.js, Chromium WebGPU, Android Termux Bionic | Production |
Canonical Usage Example
from ameva_cluster import AmevaCluster, ClusterMaster
# 1. Start worker daemon with Guard protection on remote phone:
# ameva-cluster worker --port 50052
# 2. Execute distributed inference across phone fleet:
AmevaCluster.run_master(
model_path="Qwen2.5-7B-Instruct-Q4_K_M.gguf",
workers=["100.106.251.21:50052", "100.77.47.37:50052"],
tensor_split="40,30,30",
prompt="Explain disaggregated neural memory architecture."
)