Quickstart & Execution Recipes

Standard usage patterns and rapid prototyping code

1. Instant CLI Chat Execution

# Run interactive CLI inference with auto-detected Vulkan GPU / CPU NEON
termux-llama run -m models/Llama-3.2-1B-Instruct-Q4_K_M.gguf -p "안녕하세요!" --device auto

2. Background OpenAI-Compatible Server Daemon

# Start background server daemon on port 8080 (zero-zombie supervisor)
termux-llama serve -m models/Llama-3.2-1B-Instruct-Q4_K_M.gguf --port 8080 -d

# Send standard OpenAI Chat Completion request
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Llama-3.2-1B",
    "messages": [{"role": "user", "content": "로컬 엣지 AI의 장점은?"}],
    "stream": true
  }'

# Gracefully terminate daemon
termux-llama serve --stop