Complete API Reference

100% Full Class, Struct, and Method Documentation

Core Engine & Runtime API

Class / FunctionParametersDescription
LlamaRuntime config: RuntimeConfig High-level local inference supervisor with auto backend selection.
LlamaRuntime.generate prompt: str, max_tokens: int = 512, stream: bool = False Generates tokens with latency profiling and GenerationMetrics.
ServerManager config: ServerConfig Manages lifecycle of OpenAI-compatible reverse proxy daemon.
ModelManager.download model_id: str, destination: Path Downloads curated GGUF models with SHA-256 verification.