llm-d — Achieve state of the art inference performance with modern accelerators on Kubernetes
GitHub Repo
Apache License 2.0
July 3, 2026 at 08:22 AM
0 views

llm-d — Achieve state of the art inference performance with modern accelerators on Kubernetes

@llm-dProject Author

Achieve SOTA Inference Performance On Any Accelerator: A Detailed Portrait of llm-d

llm-d Logo

In the fast-moving world of large language models, llm-d positions itself as a high-performance distributed inference serving stack designed for production deployments on Kubernetes. Born from collaboration among Red Hat, Google Cloud, IBM Research, CoreWeave, and NVIDIA—with support from a broad ecosystem including AMD, Cisco, Hugging Face, Intel, Lambda, Mistral AI, UC Berkeley, and the University of Chicago—llm-d is a Cloud Native Computing Foundation (CNCF) sandbox project. Its mission is straightforward: help you achieve the fastest “time to state-of-the-art (SOTA) performance” for open-source LLMs across a wide range of accelerators and infrastructure providers through well-tested guides and real-world benchmarks. This post unpacks what llm-d offers, how it works, and how you can begin using it to drive production-grade inference at scale.

llm-d at a Glance

llm-d is a purpose-built orchestration layer layered atop proven model servers such as vLLM and SGLang. While model servers excel at running very large models on accelerators, llm-d adds production-grade orchestration, optimization, and operational features that make real-world traffic feel reliably fast and scalable. The project’s focus is not merely to run models; it is to run them at scale with stability, predictability, and efficiency under multi-tenant workloads.

What llm-d delivers to production inference

llm-d’s offerings are organized around four core themes—each designed to address distinct pain points in production LLM serving. Together, they form a comprehensive stack that helps teams maximize throughput, minimize latency, and maintain reliability as demand grows.

1) Intelligent Routing

  • Prefix-cache aware routing accelerates the path from user query to model response, reducing repeated computation and serving latency.

  • Load-aware balancing distributes traffic across replicas to prevent hot spots and ensure consistent service levels.

  • Experimental predicted latency-based scheduling aims to further lower latency and raise throughput by anticipating response times before requests are dispatched.

    2) Advanced KV-Cache Management

  • Tonal cache strategies improve the effectiveness of the local KV store, enabling better handling of multi-turn conversations.

  • Tiered offloading to CPU or disk expands the functional working set, allowing systems to serve more complex dialog histories without overwhelming memory.

  • Precise global indexing of the KV cache state keeps state consistent and searchable across nodes, easing synchronization challenges in large deployments.

    3) Serving Large Models

  • Disaggregation strategies for prefill and decode unlock the potential of very large models, enabling efficient utilization of fast interconnects on accelerators.

  • Wide expert-parallelism across accelerator networks helps scale inference across multiple devices, maintaining throughput as model size grows.

  • The approach is designed to be compatible with broad hardware ecosystems, supporting models such as DeepSeek-R1 and GPT-OSS on a variety of accelerators.

    4) Operational Excellence

  • Intelligent flow control supports multi-tenant serving, balancing fairness and performance.

  • Proactive, SLO-aware autoscaling uses real-time inference signals to adjust capacity before latency penalties accrue.

  • The architecture emphasizes reliability, observability, and predictable performance under production load.

    5) Batch Processing

  • Experimental features for batch processing align with OpenAI-compatible Batch APIs, enabling offline inference at scale.

  • Asynchronous processing enhances hardware utilization by smoothing workloads and reducing idle times.

    Well-lit path guides and practical recipes

    llm-d’s guides are designed to be practical and reproducible. The well-lit path guides compile benchmarked recipes and Helm charts to help operators start serving quickly while following production best practices. The intent is to eliminate the heavy lifting traditionally required to tune and deploy generative AI inference on modern accelerators. Whether you’re a newcomer setting up a baseline or an experienced engineer optimizing complex deployments, these guides provide validated patterns you can trust.

    Performance Highlights: Real-world Gains

    The llm-d project highlights substantial performance gains observed in production deployments and partner benchmarks. These figures illustrate how adopting llm-d can transform latency, throughput, and overall efficiency in real-world settings.

  • 3x higher output throughput and 2x faster time-to-first-token (TTFT) with prefix-cache-aware routing versus round-robin, demonstrated on Llama 3.1 70B across 4× AMD MI300X GPUs in a Red Hat environment. The underlying result is a substantial reduction in latency and an increase in sustained throughput, enabling more requests to be served per second.

  • 40% reduction in TTFT and initial token latency (ITL) with predicted-latency scheduling compared to heuristic-based approaches on NVIDIA GPUs and Google Cloud. This is achieved through smarter request placement that anticipates completion times and adjusts routing accordingly.

  • Up to 70% higher tokens per second with prefill/decode disaggregation versus standard vLLM. This was shown with GPT-OSS on NVIDIA B200 (p6-b200) hardware on AWS, reflecting how disaggregated inference can unlock significant throughput gains for large models.

  • 10–30% throughput improvement with disaggregated serving on identical infrastructure for GPT-OSS-120B and Llama 3.3 70B on AMD MI300X (Oracle), illustrating cross-hardware benefits of the disaggregation approach.

  • 50k tokens/second cluster throughput with Wide Expert-Parallelism on a 16×16 NVIDIA B200 setup, delivering about 3.1k tokens/second per GPU. This showcases the scalability of wide EP architectures for high-volume inference.

  • 13.9x throughput improvement with hierarchical KV offloading at 250 concurrent users versus GPU-only setups on 4× NVIDIA H100, underscoring the value of cache-aware designs in multi-tenant environments.

    For those who want more depth, detailed, reproducible benchmarks are available on Prism, where you can explore a range of configurations and results across accelerators and workloads.

    Getting Started: Quickstart and Baselines

    Ready to reach SOTA performance? The llm-d Quickstart Guide provides a clear path to deploying your first optimized inference service on Kubernetes. It walks you through:

  • Setting up the llm-d stack and dependencies

  • Configuring the intelligent router

  • Validating performance using production-ready benchmarks

    A recommended starting point is the Optimized Baseline, which provides a high-performance foundation suitable for a wide range of LLM serving use cases. This baseline streamlines the journey toward production-grade inference by giving you a robust, well-tested starting point that you can customize for your workloads.

    Latest News and Milestones

    llm-d has seasoned a trajectory of rapid development and community growth in recent years. Key milestones in 2026 illustrate the project’s maturation and expanding footprint in the AI infrastructure ecosystem.

  • May 2026 (v0.7): The release introduces an optimized baseline that’s renamed and stabilized, kustomize-first migrated guides, expanded nightly CI for OpenShift, GKE, and CoreWeave, predicted-latency scheduling GA, batch gateway (experimental), and revamped project-wide documentation. This release consolidates the effort to make production-grade deployment more accessible and reliable.

  • March 2026: llm-d joins the CNCF as a Sandbox project. This milestone marks formal recognition of llm-d as a critical component of SOTA AI infrastructure, supported by major players and a commitment to open collaboration across the ecosystem.

  • February 2026 (v0.5): The v0.5 release introduces reproducible benchmark workflows, hierarchical KV offloading, cache-aware LoRA routing, active-active high availability, UCCL-based transport resilience, and scale-to-zero autoscaling. It validates approximately 3.1k tokens/second per B200 decode GPU (wide-EP) and up to 50k output tokens per second on a 16×16 B200 topology. The improvements reflect substantial reductions in TTFT and latency compared with earlier baselines.

  • December 2025 (v0.4): Demonstrates 40% reductions in per-output-token latency for DeepSeek V3.1 on H200 GPUs, Intel XPU, and Google TPU disaggregation. It also introduces a new well-lit path for prefix cache offload to vLLM-native CPU memory tiering and previews a workload variant autoscaler to improve model-as-a-service efficiency.

    The roadmap captures llm-d’s emphasis on reproducibility, resilience, and performance as it scales to more accelerators and deployment scenarios.

    Architecture: How llm-d Fits into the Inference Stack

    llm-d accelerates distributed inference by combining best-in-class open technologies with Kubernetes-based orchestration. The architecture emphasizes modularity, scalability, and compatibility with industry-standard components. The core objective is to provide a production-ready platform that can run on a variety of accelerators and cloud environments while delivering repeatable performance gains.

    The Architecture story centers on:

  • Integration with vLLM and related model servers to handle the heavy lifting of model execution on accelerators.

  • Intelligent routing, cache management, and KV store innovations that improve latency and throughput for multi-turn conversations.

  • Disaggregated deployment patterns that separate prefill and decode workloads across fast interconnects, enabling scalable and efficient inference even for massive models.

  • Robust operational capabilities, including multi-tenant flow control, autoscaling, resilience patterns, and observability that support practical production use.

    See the llm-d architecture diagram for a visual summary of these components. The diagram (llm-d Arch) provides a concise view of how the pieces connect and interact to deliver high-performance, scalable inference.

    [llm-d Arch]

    Releases, Guides, and How to Contribute

    llm-d maintains an active set of guides that are living documents designed to stay current with the project’s rapid evolution. For details about Helm charts, component releases, and updated configurations, visit the GitHub Releases page. The release notes capture the nuances of each version, including new features, performance refinements, and known issues. Accelerators-specific documentation outlines tested configurations, networks, and best practices, ensuring operators have the right starting point for their hardware.

    To stay aligned with best practices and governance, llm-d adheres to CNCF standards of conduct and community collaboration. The project overview provides deeper context on development processes and governance. Contributing guidelines outline how developers can participate, from code contributions to documentation and community engagement. Special Interest Groups (SIGs) offer focused avenues to collaborate on specific aspects of the project, ranging from architecture to performance benchmarking. The llm-d community uses Slack as the primary channel for ongoing discussions and coordination, in addition to a Google Groups mailing list for sharing architectural diagrams and other content. For community calendars and scheduling, the shared llm-d calendar provides a transparent view of standups and SIG meetings.

    The Release and Guides ecosystem at a glance:

  • Guides: Live documentation that evolves with the project, including getting-started tutorials and deployment patterns.

  • Releases: The GitHub Releases page hosts release notes detailing new features, fixes, and improvements.

  • Accelerators Documentation: Accelerator-specific guides outline configurations, networks, and validation procedures.

    Architecture and Implementation Highlights

    The llm-d stack is designed to be pragmatic and production-ready. It does not reinvent existing, battle-tested technologies but instead focuses on orchestration, optimization, and reliability on top of those technologies. By integrating with vLLM and Kubernetes, llm-d provides:

  • A proven foundation for multi-tenant, scalable inference deployments.

  • Optimizations that reduce latency and increase throughput for diverse workloads and model families.

  • A pathway to adapt to different accelerators and cloud providers with tested configurations and documented best practices.

    Contribute and Community Practices

    llm-d emphasizes open collaboration and responsible governance. The CNCF Code of Conduct governs how contributors interact within the project, ensuring a welcoming and inclusive environment for all participants. Contributors are encouraged to review the project overview to understand development processes and governance structures. The contributing guidelines detail how to submit changes, how reviews are handled, and how to participate in the project’s development cycle. Community engagement is supported through multiple channels:

  • Slack for developer discussions and cross-organization collaboration.

  • Bi-weekly contributor standups (every other Wednesday at 12:30 PM ET) and various SIG meetings.

  • Google Groups for sharing architecture diagrams and related content.

  • Public calendars and shared resources to keep the community aligned and informed.

    License

    The llm-d project is licensed under the Apache License 2.0. The license page provides the legal terms and conditions for using, modifying, and distributing the software. This permissive license supports broad adoption and integration into a variety of production environments.

    Why llm-d: The Value Proposition for Production Inference

    llm-d is designed for teams who want to deploy performant, reliable, and scalable model-serving infrastructure in real-world environments. Its architecture and feature set address the complexities that typically hinder production inference:

  • Multi-tenant scenarios: Intelligent flow control and autoscaling help balance competing workloads while maintaining predictable latency.

  • Large models and advanced interconnects: Prefetch, decode disaggregation, and wide EP enable scaling to very large models without sacrificing throughput.

  • Cache-aware design: Global KV indexing and tiered offloading expand the effective working set, reducing bottlenecks caused by memory constraints.

  • Production-ready tooling: Guides, benchmarks, and validated recipes reduce the time to deployment, enabling teams to move quickly from theory to operational service.

    If you’re evaluating options for high-performance inference, llm-d offers a structured path to achieve SOTA results across accelerators and infrastructures. It pairs a robust core with practical, well-documented patterns that translate to tangible improvements in latency, throughput, and reliability.

    How to Use llm-d in Practice

    A practical approach to adopting llm-d typically follows a sequence of steps aligned with its guides and documented patterns:

  • Start with the Optimized Baseline: Use the baseline as a solid, high-performance foundation suitable for a wide array of LLM serving scenarios.

  • Set up the Quickstart: Deploy the llm-d stack on Kubernetes using the quickstart guide to validate end-to-end operation and performance.

  • Configure Intelligent Routing: Implement the prefix-cache and load-aware routing to optimize traffic distribution and latency.

  • Enable KV Cache Management: Activate tiered offloading and global KV state indexing to maximize working set size and state coherence.

  • Integrate Batch Processing: If offline or batch workloads are part of your use case, experiment with the batch gateway and asynchronous processing to maximize hardware utilization.

  • Benchmark and Iterate: Use the Prism benchmarks and the well-lit path guides to reproduce results, compare configurations, and converge on the best performing deployment for your model family and hardware.

    Conclusion: A Path to SOTA Inference on Any Accelerator

    llm-d represents a concerted effort to bridge the gap between cutting-edge research and real-world production. It leverages established model servers, Kubernetes-based orchestration, and a suite of production-oriented optimizations to deliver state-of-the-art performance across diverse accelerators and deployment scenarios. From intelligent routing and aggressive cache management to disaggregated serving and proactive autoscaling, llm-d is designed to help teams push inference performance toward SOTA levels while maintaining stability and predictability.

    If you’re seeking a robust, scalable, and well-supported path to production-grade LLM inference, llm-d provides a compelling architecture and a clearly defined set of practices. The project’s ongoing evolution—as reflected in its v0.7 release cadence, CNCF sandbox status, and expanding ecosystem—signals a growing community effort to shape the future of AI infrastructure. For organizations that want faster time-to-SOTA outcomes and a reliable path to production-grade inference at scale, llm-d offers a coherent, battle-tested approach grounded in real-world benchmarks and field-tested recipes.

    To learn more, start with the Quickstart Guide, explore the well-lit path guides, and consult the release notes and accelerator documentation. The llm-d project is designed to be approachable for operators while remaining powerful enough to satisfy the demands of large-scale, production AI workloads. With ongoing community engagement, registries of benchmarks, and ongoing improvements in automation and resilience, llm-d stands as a compelling option for teams pursuing the fastest, most reliable inference at scale.

    llm-d Architecture Diagram Image

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
achieve-sota-inference-performance-on-any-accelerator
Created
July 3
Last Updated
July 3, 2026 at 08:22 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.