MemVanta
Low-memory C++20 LLM inference runtime for quantized GGUF models on CPU — mmap-backed weights, paged KV cache, Q4/Q8 kernels, and reproducible llama.cpp benchma
What is MemVanta?
Low-memory C++20 LLM inference runtime for quantized GGUF models on CPU — mmap-backed weights, paged KV cache, Q4/Q8 kernels, and reproducible llama.cpp benchmarks.