Quickstart & Execution Recipes
Standard usage patterns and rapid prototyping code
1. Instant CLI Chat Execution
# Run interactive CLI inference with auto-detected Vulkan GPU / CPU NEON
termux-llama run -m models/Llama-3.2-1B-Instruct-Q4_K_M.gguf -p "안녕하세요!" --device auto
2. Background OpenAI-Compatible Server Daemon
# Start background server daemon on port 8080 (zero-zombie supervisor)
termux-llama serve -m models/Llama-3.2-1B-Instruct-Q4_K_M.gguf --port 8080 -d
# Send standard OpenAI Chat Completion request
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Llama-3.2-1B",
"messages": [{"role": "user", "content": "로컬 엣지 AI의 장점은?"}],
"stream": true
}'
# Gracefully terminate daemon
termux-llama serve --stop