Model Hub & Pre-Quantized GGUF Registry

Verified 1.58-bit GGUF models optimized for Android Termux on-device inference.

Verified Pre-Quantized Model Catalog

Model Alias Hugging Face Upstream Repository Parameters & Size Recommended Hardware / Target Download CLI Command
falcon-e-1b tiiuae/Falcon-E-1B-Instruct-GGUF 1.0B / 635 MB (SwiGLU) Real-Time Mobile Dialog (Galaxy A53: 9.69 tok/s) termux-bitnet download falcon-e-1b
bitnet-2b microsoft/bitnet-b1.58-2B-4T-gguf 2.0B / 1.13 GB (Squared ReLU) General Reasoning (Galaxy S25: 3.95 tok/s, A53: 5.91 tok/s) termux-bitnet download bitnet-2b
bitnet-embed-270m microsoft/bitnet-embedding-270m 268M / 367 MB (Embedding) On-Device Vector Search & Local RAG (Galaxy A53: 30.68 tok/s) termux-bitnet download bitnet-embed-270m
falcon3-7b tiiuae/Falcon3-7B-Instruct-1.58bit-GGUF 7.45B / 3.05 GB (SwiGLU) High-Capacity Reasoning on 6GB RAM Phones (Galaxy A53: 2.00 tok/s) termux-bitnet download falcon3-7b
bitnet-large RichardErkhov/1bitLLM_-_bitnet_b1_58-large-gguf 0.7B / 404 MB Ultra-Constrained Legacy Devices termux-bitnet download bitnet-large

Model Cache Location & Storage Architecture

Downloaded models are cached locally in ~/.cache/termux-bitnet/models/. The built-in downloader validates SHA-256 checksums and automatically resumes interrupted transfers using HTTP Range headers.