Model Hub & Pre-Quantized GGUF Registry
Verified 1.58-bit GGUF models optimized for Android Termux on-device inference.
Verified Pre-Quantized Model Catalog
| Model Alias | Hugging Face Upstream Repository | Parameters & Size | Recommended Hardware / Target | Download CLI Command |
|---|---|---|---|---|
falcon-e-1b |
tiiuae/Falcon-E-1B-Instruct-GGUF |
1.0B / 635 MB (SwiGLU) | Real-Time Mobile Dialog (Galaxy A53: 9.69 tok/s) | termux-bitnet download falcon-e-1b |
bitnet-2b |
microsoft/bitnet-b1.58-2B-4T-gguf |
2.0B / 1.13 GB (Squared ReLU) | General Reasoning (Galaxy S25: 3.95 tok/s, A53: 5.91 tok/s) | termux-bitnet download bitnet-2b |
bitnet-embed-270m |
microsoft/bitnet-embedding-270m |
268M / 367 MB (Embedding) | On-Device Vector Search & Local RAG (Galaxy A53: 30.68 tok/s) | termux-bitnet download bitnet-embed-270m |
falcon3-7b |
tiiuae/Falcon3-7B-Instruct-1.58bit-GGUF |
7.45B / 3.05 GB (SwiGLU) | High-Capacity Reasoning on 6GB RAM Phones (Galaxy A53: 2.00 tok/s) | termux-bitnet download falcon3-7b |
bitnet-large |
RichardErkhov/1bitLLM_-_bitnet_b1_58-large-gguf |
0.7B / 404 MB | Ultra-Constrained Legacy Devices | termux-bitnet download bitnet-large |
Model Cache Location & Storage Architecture
Downloaded models are cached locally in ~/.cache/termux-bitnet/models/. The built-in downloader validates SHA-256 checksums and automatically resumes interrupted transfers using HTTP Range headers.