Termux-Vision

Zero-Dependency On-Device Computer Vision & Multimodal VLM Inference Engine with UltraFace SSD Neural Face Detection (11.94ms), Pure Vulkan GPU (0.23ms), and ARM64 NEON Acceleration for Android Termux

PyPI Version npm Version License Platform
1-Line Quick Installation

Install the official package directly into your runtime:

pip install termux-vision
# or: termux-vision install
# or: npm install -g termux-vision
# or: bash <(curl -sSL https://raw.githubusercontent.com/uno-km/termux-vision/main/install.sh)

The Engineering Challenge

Standard vision frameworks (OpenCV, TorchVision) suffer from massive binary sizes (>150MB), complex C++ compilation bottlenecks on ARM64 Termux, and legacy 2001-era Haar Cascade filters that produce dozens of false-positive ghost boxes on indoor lighting and white walls.

The Architectural Breakthrough

Provides UltraFace SSD Deep Learning Neural Face Detection (11.94ms on Snapdragon 8 Elite, 100% False-Positive Elimination), 100% End-to-End Vulkan GPU Compute Canny Edge Detection (0.23ms), ultra-optimized ARM64 NEON C++ kernels (3.02ms), 5-stage classical filters, on-device SmolVLM/Qwen2-VL Multimodal Vision-Language inference, and automated ONNX asset provisioning.

Key Capabilities & Built-in Hardening

UltraFace SSD Deep Learning Neural Face Detector (11.94 ms)

Directly binds UltraFace RFB Single-Shot Detector (SSD) with ONNX Runtime ARM64 NEON, eliminating 2001-era Haar Cascade false positives (0 ghost boxes) and achieving 11.94ms ultra-low latency on Snapdragon 8 Elite with 100.0% confidence.

100% End-to-End Pure Vulkan GPU Compute Canny (0.23 ms)

Chains Sobel 3x3, NMS, and Hysteresis compute shaders entirely within VRAM using vkCmdPipelineBarrier, accelerating edge detection by 834x (0.23ms on Snapdragon 8 Elite Adreno 830).

Ultra-Fast ARM64 NEON C++ Kernel (3.02 ms)

Eliminates atan2f via tangent ratio bit quantization and 1-byte direction buffers, running Canny filtering in 3.02ms on Snapdragon 865 and 4.36ms on Exynos 1380.

4-Axis Idempotent Installer & Asset Provisioner

Provisions precompiled ARM64 native binaries, Vulkan GPU probes, multimodal VLM CLI, and verified UltraFace ONNX weights automatically with 4-axis smoke tests.

Supported Compute Kernels & Operations

Subsystem Category Operations & Kernels Status
Classical CV Filters 5-Stage Canny Edge Detector, Sobel 3x3, Gaussian Blur 5x5, 2D Integral Images, Morphology Production
Multimodal VLM Moondream2 1.8B, SmolVLM-500M, Qwen2-VL-2B, Supervised llama-cli Bridge, Dynamic Resolution Scaling Production
ARM Mali GPU Acceleration Vulkan SPIR-V Compute (-ngl 99), 0.00 MiB CPU VRAM, MMVQ Kernel Tuning (--tune-mali) Production Verified
Qualcomm Adreno Acceleration Snapdragon 8 Elite Adreno 830 Vulkan SPIR-V JIT Patch (mul_mat_vec_max_cols = 2), KGSL Watchdog Safe Prefill (-b 64 -ub 64) Production Verified (정식 지원 및 실기기 검증 완료)
Samsung Xclipse Acceleration Exynos 2200/2400 AMD RDNA Mobile SPIR-V Instruction Scheduling In Development (개발 진행 중)

Canonical Usage Example

import termux_vision as tv

# 1. 100% Vulkan GPU Canny Edge Detection (0.23ms) or NEON CPU (3.02ms)
img = tv.io.load_image("photo.jpg")
edges = tv.cv.canny(tv.transforms.to_grayscale(img), low_threshold=40, high_threshold=120, backend="auto")
tv.io.save_image(edges, "edges.png")

# 2. UltraFace SSD Neural Face Detection (11.94ms, 100% False-Positive Elimination)
detections = tv.detect.detect_faces(img, score_threshold=0.70, backend="auto")
for d in detections:
    print(f"Face Box: {d.bbox.to_xywh()}, Confidence: {d.score * 100:.1f}%")

# 3. On-Device VLM Multimodal Inference with 5-Backend Hardware Acceleration
with tv.vlm.load("smolvlm-500m", device="gpu") as engine:
    res = engine.describe("photo.jpg", prompt="Describe this scene in detail.", quality="optimal")
    print(f"Generated ({res.metrics.tokens_per_second:.1f} t/s via {res.metrics.backend}): {res.text}")

Getting Started & Deep Guides