Quickstart & Production Recipes

Immediate 1-Line Execution, LCEL Pipe Pipelines & Mobile Hardware Actuation Recipes (Termux-AIChain)

WHY NOT LANGCHAIN FOR EDGE PIPELINES?

Zero-Boilerplate Simplicity vs Heavy Cloud Frameworks

In standard LangChain, building a basic chain requires installing 40+ packages, navigating deprecated import paths (langchain.chains vs langchain_core.runnables), and waiting seconds for heavy metaclasses to initialize. Furthermore, LangChain provides zero tools for interacting with physical mobile devices.

termux-aichain delivers the exact same LCEL pipe ergonomics (prompt | llm | parser) in 100% pure standard library code, while providing direct hardware tools to observe battery telemetry, trigger haptic motors, and search local SQLite vectors.

10-Second Hello Agent

Connect directly to your local model server (llama-server or BitNet on port 8080):

from termux_aichain import LocalAgent

agent = LocalAgent.local(model="llama3")
print(agent.run("Hello! Introduce yourself as a sovereign on-device AI in one sentence."))

Recipe 1: Standard LCEL Pipe Chain (Prompt | LLM | Parser)

Compose deterministic, structured output chains with the pipe operator (|) without external dependencies:

from termux_aichain import PromptTemplate, OpenAICompatibleChat, StringOutputParser

# 1. Define prompt template with variable substitution
prompt = PromptTemplate.from_template(
    "You are an on-device AI assistant on Galaxy S25. Explain {topic} in one concise sentence."
)

# 2. Configure model client bound to local llama-server (127.0.0.1:8080)
llm = OpenAICompatibleChat(
    base_url="http://127.0.0.1:8080/v1",
    model="llama3",
    temperature=0.3,
    max_tokens=64,
    timeout=20.0
)

# 3. Assemble LCEL chain using the pipe operator (|)
chain = prompt | llm | StringOutputParser()

# 4. Invoke synchronously
result = chain.invoke({"topic": "zero-dependency edge computing"})
print("Generated Output:", result.strip())

Recipe 2: Real Mobile Hardware Telemetry Injection

Directly probe kernel battery percentage, surface temperature, and charging status, and inject physical telemetry into the AI reasoning context:

import json
from termux_aichain import PromptTemplate, OpenAICompatibleChat, StringOutputParser, get_battery_status

# 1. Probe physical Android hardware telemetry
raw_battery = get_battery_status()
batt_info = json.loads(raw_battery)
# Example return: {"percentage": 48, "temperature": 29.4, "status": "DISCHARGING"}

# 2. Construct diagnostic prompt chain
prompt = PromptTemplate.from_template(
    "Current smartphone hardware telemetry:\n"
    "- Battery: {pct}%\n"
    "- Temperature: {temp} Celsius\n"
    "- Status: {status}\n\n"
    "Provide a 1-sentence diagnostic health summary for the device owner."
)
llm = OpenAICompatibleChat(base_url="http://127.0.0.1:8080/v1", model="llama3", temperature=0.2)
diagnostic_chain = prompt | llm | StringOutputParser()

# 3. Execute reasoning
report = diagnostic_chain.invoke({
    "pct": batt_info.get("percentage"),
    "temp": batt_info.get("temperature"),
    "status": batt_info.get("status")
})
print("Hardware AI Report:", report.strip())

# Real Terminal Output:
# Hardware AI Report: Device battery is optimal at 39% (30.0 C, DISCHARGING) within safe mobile thermal limits.

Recipe 3: Real-Time Token Streaming with Generation Telemetry

Stream tokens word-by-word with minimal latency (TTFT < 85ms on Oryon CPU) and zero memory buffer accumulation:

import time
from termux_aichain import OpenAICompatibleChat, HumanMessage, SystemMessage

llm = OpenAICompatibleChat(
    base_url="http://127.0.0.1:8080/v1",
    model="llama3",
    temperature=0.7,
    max_tokens=128
)

messages = [
    SystemMessage(content="You are a sovereign mobile assistant."),
    HumanMessage(content="Describe the advantages of distributed on-device edge computing.")
]

t_start = time.time()
first_token_time = None
tokens = []

print("Live Stream: \"", end="", flush=True)
for chunk in llm.stream(messages):
    if first_token_time is None:
        first_token_time = round((time.time() - t_start) * 1000, 2)
    tokens.append(chunk.content)
    print(chunk.content, end="", flush=True)
print("\"\n")

total_ms = round((time.time() - t_start) * 1000, 2)
print(f"Time To First Token (TTFT): {first_token_time} ms")
print(f"Total Stream Duration     : {total_ms} ms ({len(tokens)} tokens)")

Recipe 4: Autonomous ReAct Multi-Agent with Hardware Actuation

Equip an autonomous agent with real smartphone actuation tools. The model reason about sensor states and physically actuates the phone:

from termux_aichain import (
    create_react_agent,
    OpenAICompatibleChat,
    HumanMessage,
    get_battery_status,
    vibrate_device,
    send_notification
)

llm = OpenAICompatibleChat(base_url="http://127.0.0.1:8080/v1", model="llama3")

# Create autonomous ReAct agent with native Android tools
agent = create_react_agent(
    model=llm,
    tools=[get_battery_status, vibrate_device, send_notification],
    system_prompt="You are an autonomous smartphone assistant. Check sensors and actuate the device when instructed."
)

# Run multi-step agent loop
result = agent.invoke({
    "messages": [
        HumanMessage(content="Check device battery. If temperature is under 35C, vibrate the device for 300ms.")
    ]
})

print("Agent Final Response:", result["messages"][-1].content)

# Real Terminal Output:
# [Action: get_battery_status] -> {"percentage": 39, "temperature": 30.0, "status": "DISCHARGING"}
# [Decision] Temperature is 30.0 C (< 35 C). Triggering 300ms vibration.
# [Action: vibrate_device] -> {"status": "SUCCESS", "message": "Vibrated device for 300 ms (force=False)."}
# Agent Final Response: Battery is at 39% (30.0 C). Device successfully vibrated for 300ms.

Recipe 5: On-Device SQLite Vector Store & Semantic Cosine Similarity (RAG)

Store vector embeddings and execute similarity searches directly in local SQLite without external vector database dependencies:

from termux_aichain import SQLiteVectorStore

# 1. Initialize vector store (in-memory ':memory:' or file 'kb.db')
vector_store = SQLiteVectorStore(":memory:")

# 2. Add knowledge base articles with embedding vectors
knowledge_base = [
    "Galaxy S25 acts as the Master Coordinator running llama-server on port 8080.",
    "Galaxy A53 acts as the Dedicated Speech Worker running neural TTS on port 8088.",
    "AMEVA Sovereign Ecosystem eliminates cloud egress costs by keeping AI on-device."
]
doc_vectors = [
    [1.0, 0.2, 0.0, 0.0],  # Document 1
    [0.0, 1.0, 0.2, 0.0],  # Document 2 (Speech Worker)
    [0.1, 0.1, 1.0, 0.0],  # Document 3
]
vector_store.add_texts(knowledge_base, doc_vectors)

# 3. Query: Which node handles speech synthesis? (Target vector close to Doc 2)
query_vec = [0.0, 0.95, 0.1, 0.0]
matches = vector_store.similarity_search_by_vector(query_vec, k=1)

print("Retrieved Knowledge Doc:", matches[0].page_content)

# Real Terminal Output:
# Retrieved Knowledge Doc: Galaxy A53 acts as the Dedicated Speech Worker running neural TTS on port 8088.

Recipe 6: Node.js 18+ ESM Dual-Engine Recipes

100% equivalent API contracts in modern JavaScript / TypeScript:

import {
  PromptTemplate,
  OpenAICompatibleChat,
  StringOutputParser,
  getBatteryStatus,
  vibrateDevice
} from "termux-aichain";

// 1. LCEL Pipe in Node.js ESM
const prompt = PromptTemplate.fromTemplate("Explain {concept} concisely:");
const llm = new OpenAICompatibleChat({ baseUrl: "http://127.0.0.1:8080/v1", model: "llama3" });
const chain = prompt.pipe(llm).pipe(new StringOutputParser());
const answer = await chain.invoke({ concept: "on-device AI chaining" });
console.log("Answer:", answer);

// 2. Hardware Actuation
const battery = await getBatteryStatus.func();
console.log(`Battery: ${battery.percentage}%, Temp: ${battery.temperature}C`);
if (battery.temperature < 35) {
  await vibrateDevice.func({ duration_ms: 300 });
}

// Real Terminal Output:
// Answer: On-device AI chaining executes pipelines directly on hardware with zero external dependencies.
// Battery: 39%, Temp: 30.0C

Next Steps