imp

An LLM server for a single RTX 5090, built for agent workloads: tool calls, long conversations, reasoning, and many requests at once. One of the fastest engines

What is imp?

An LLM server for a single RTX 5090, built for agent workloads: tool calls, long conversations, reasoning, and many requests at once. One of the fastest engines on this card, at batch 1 and at dozens of concurrent streams, with the numbers in the repo.