AMEVA-Cluster

Symmetric Disaggregated Mobile RAM Pooling & On-Device AI Acceleration Runtime with Compute-Memory Decoupling

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install ameva-cluster
# or
npm install -g @ameva/cluster

The Engineering Challenge

Executing frontier 70B scale models on standalone mobile devices inevitably triggers Android Low Memory Killer (LMK) termination and thermal throttling. Conventional distributed inference forces compute on all nodes, overheating lower-tier battery-powered devices.

The Architectural Breakthrough

AMEVA-Cluster unifies idle LPDDR memory across multiple heterogeneous mobile smartphones into a single disaggregated memory pool using a symmetric master/worker binary. Compute is fully isolated on a single flagship GPU, while edge phones serve strictly as passive RAM caches, completely preventing thermal throttling and Android Low Memory Killer (LMK) SIGKILL terminations. Features built-in storage-assisted hybrid failover to local UFS 3.1 flash mmap streaming and fast-path HMAC mutual authentication Guard proxies.

Key Capabilities & Built-in Hardening

Symmetric Single Binary CLI

All-in-one architecture supporting both master and worker roles via single CLI flag.

Compute-Memory Decoupling

Worker nodes execute zero GPU passes, acting solely as RAM caches to prevent heat and failure.

300MB Safety Guard-Band

Automatically deducts buffer memory from usable pool to protect against Android LMK kills.

Mutual Authentication Guard Proxy

0.001s Fail-Fast HMAC-SHA256 verification rejecting unauthorized RPC traffic and DoS attempts.

Storage-Assisted Hybrid Failover

Gracefully degrades to local UFS 3.1 flash streaming upon worker disconnects without crashing.

6-Modality Pluggable Engine

Standardized distributed tensor sharding across LLM, Diffusion, Vision, STT, TTS, and Training.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Compute Engine WebGPU Compute Shaders (WGSL), FP16/FP32 Production
Memory Subsystem Zero-Copy Ring Buffers, Weakref GC Pooling Production
Platform Runtimes Node.js, Chromium WebGPU, Android Termux Bionic Production

Canonical Usage Example

from ameva_cluster import AmevaCluster, ClusterMaster

# 1. Start worker daemon with Guard protection on remote phone:
# ameva-cluster worker --port 50052

# 2. Execute distributed inference across phone fleet:
AmevaCluster.run_master(
    model_path="Qwen2.5-7B-Instruct-Q4_K_M.gguf",
    workers=["100.106.251.21:50052", "100.77.47.37:50052"],
    tensor_split="40,30,30",
    prompt="Explain disaggregated neural memory architecture."
)

Getting Started & Deep Guides