Complete API Reference
100% Full Class, Struct, and Method Documentation
Core Engine & Runtime API
| Class / Function | Parameters | Description |
|---|---|---|
LlamaRuntime |
config: RuntimeConfig |
High-level local inference supervisor with auto backend selection. |
LlamaRuntime.generate |
prompt: str, max_tokens: int = 512, stream: bool = False |
Generates tokens with latency profiling and GenerationMetrics. |
ServerManager |
config: ServerConfig |
Manages lifecycle of OpenAI-compatible reverse proxy daemon. |
ModelManager.download |
model_id: str, destination: Path |
Downloads curated GGUF models with SHA-256 verification. |