llm-d

Achieve state of the art inference performance with modern accelerators on Kubernetes

What is llm-d?

Achieve state of the art inference performance with modern accelerators on Kubernetes