Advanced Parameters & Hardware Tuning

Hardware Compute Optimization, KV Cache Quantization, Sampling Controls & Mobile Thermal Budgets (Termux-AIChain)

WHY HARDWARE-AWARE PARAMETERS MATTER ON MOBILE

Thermal Budgets & Memory Bounds vs Cloud Assumptions

Cloud agent frameworks like LangChain assume server-grade machines with constant cooling, redundant power supplies, and hundreds of gigabytes of RAM. Running unconstrained models on a smartphone without strict thread affinity and KV cache limits causes severe thermal throttling (SoC temperature > 45°C), rapid battery depletion, and instant kernel OOM kills.

termux-aichain exposes precise hardware-level control knobs: You can tune CPU big-core allocation, micro-batch sizes, KV cache quantization (saving up to 75% RAM), and default-deny tool sandboxing to run sustained autonomous workloads safely on battery power.

1. Hardware Compute & Engine Flags (`LocalServerConfig`)

Parameter Name Type Default Recommended Range Physical System Impact & Engineering Rationale
threads int CPU-1 4 ~ 8 Binds inference to ARM big cores (e.g. 8 threads for Snapdragon 8 Elite Oryon; 6 for Exynos 1280) to maximize NEON SIMD throughput without starving OS tasks.
n_ctx int 2048 1024 ~ 4096 Context window token limit. Reducing from 8192 to 2048 saves up to 75% of static DRAM allocation, crucial for devices with ≤8GB RAM.
n_batch int 512 128 ~ 512 Prompt evaluation batch size. Larger batch speeds up TTFT (Time To First Token) during long prompt processing.
n_ubatch int 256 32 ~ 256 Physical micro-batch compute size. Setting to 128 or 256 prevents burst memory allocation spikes on mobile LPDDR5 RAM.
n_gpu_layers int 0 0 ~ 99 GPU layer offload count. Set to 0 for CPU NEON inference (stable on Oryon), or 99 for Adreno/Mali Vulkan GPU compute.
cache_type_k str "f16" "q8_0" / "q4_0" Key cache quantization format. "q8_0" saves 50% KV cache memory with zero perceptible perplexity loss; "q4_0" saves 75%.
cache_type_v str "f16" "q8_0" / "q4_0" Value cache quantization format. Match with cache_type_k for symmetrical memory reduction.
mlock bool False True / False Locks model weights in physical RAM to prevent Android OS zRAM swap thrashing, eliminating sudden token generation stalls.

2. Generation & Sampling Controls (`OpenAICompatibleChat`)

Sampling Parameter Type Default Tuning Spectrum Behavioral Characteristics
temperature float 0.7 0.0 ~ 1.5 0.0: Deterministic mathematical/JSON output. 0.3: Technical analysis. 0.7: Conversational balance. 1.2+: Creative exploration.
top_p float 0.95 0.1 ~ 1.0 Nucleus probability cutoff. Filters candidate tokens to the top cumulative probability mass, preventing low-probability hallucinations.
top_k int 40 1 ~ 100 Limits candidate pool to the top K most likely tokens before softmax sampling.
min_p float 0.05 0.01 ~ 0.2 Minimum relative probability cutoff against highest-probability token. Eliminates tail hallucinations far better than top_p alone.
repeat_penalty float 1.18 1.0 ~ 1.3 Penalizes repeated tokens. Set to 1.15~1.20 to prevent models from falling into infinite loops during long text generation.
max_tokens int 128 16 ~ 2048 Hard ceiling on generated tokens per turn. Critical on mobile to prevent battery-draining runaway generations.
stop list[str] None ["<|eot_id|>", "\n\n"] Explicit delimiter tokens that immediately halt generation.
timeout float 20.0 5.0 ~ 120.0 Socket timeout in seconds. Triggers fail-fast error rather than hanging indefinitely during mobile sleep states.

3. Tool Execution Governance & Sandboxing (`ToolPolicy`)

Mobile smartphones carry sensitive personal data (GPS, SMS, camera). termux-aichain enforces strict default-deny tool policies:

from termux_aichain import ToolPolicy, ToolRule

# Enforce zero-trust default-deny security policy
policy = ToolPolicy(
    default="deny",
    rules=[
        # Read-only sensor inspection: Allowed without limits
        ToolRule(name="get_battery_status", allow=True),
        # Physical actuation: Allowed with rate limiting
        ToolRule(name="vibrate_device", allow=True, max_calls_per_minute=10),
        # High-privilege shell execution: Explicitly denied
        ToolRule(name="execute_shell", allow=False)
    ]
)

4. Mobile Thermal Throttling & Battery Safety Protocol

To ensure long-term device longevity, agents should check battery temperature before executing compute-intensive tasks:

import json
import time
from termux_aichain import get_battery_status

def check_thermal_budget(max_temp_celsius=40.0) -> bool:
    data = json.loads(get_battery_status())
    current_temp = data.get("temperature", 30.0)
    if current_temp > max_temp_celsius:
        print(f"[THERMAL THROTTLE] Temperature {current_temp}C exceeds safe limit {max_temp_celsius}C. Pausing...")
        time.sleep(10.0)
        return False
    return True

Next Steps