Raja Koduri Lays Out the Design Priorities That Will Define Inference Infrastructure

March 24, 2026 contact@asapdigitalmarketing.org

Raja Koduri’s Keynote Sets the Stage for a New Design Era

At a recent industry event centered on the theme AI Everywhere, Raja Koduri took the stage and grounded that phrase in something harder to ignore than a slogan: actual numbers. His talk, Chiplet Quilting for the Age of Inference, addressed a challenge that every silicon architect is now facing head-on. AI inference demand is growing faster than traditional design assumptions can accommodate, and the decisions being made in chip architecture today will determine performance, cost, and energy use for years to come.

The central argument was not abstract. It was urgent.

Token usage currently sits in the quadrillions per month. Using conservative growth projections, inference demand is expected to reach roughly 10¹⁸ tokens per month by 2030. Even under best-case efficiency conditions, that trajectory points to hundreds of gigawatts of infrastructure. The demand signal is not ambiguous. The open question is how architects respond to it.

First Principles Still Decide Outcomes

The Fundamentals That Matter When Building Silicon

Raja Koduri framed the challenge through a set of fundamentals that cut to the core of what silicon architects actually need to get right:

  • Performance per dollar
  • Performance per watt
  • Flexibility across future workloads
  • Packaging cost
  • Energy to compute
  • Energy to move data
  • Energy to access memory

Physics defines the limits. Economics determines what scales.

Compute operations now cost femtojoules per bit. Data movement costs considerably more. Off-chip memory access dominates the energy budget, and the variables that matter most are deceptively simple: distance, memory placement, and packaging. Each of these has real consequences at scale. A design that ignores any one of them pays for it in efficiency losses that compound over time.

The throughline of this section of the keynote was that chiplets are no longer optional. Post-Dennard scaling has forced hard tradeoffs into the open. Advanced nodes cost more per square millimeter. Power efficiency gains have flattened. Not every function belongs on the most advanced process, and pretending otherwise leads to systems that are expensive, inefficient, or both.

Why Chiplet Quilting Changes the Equation

Chiplet architectures make those tradeoffs explicit rather than leaving them buried in monolithic design assumptions. Chiplet quilting takes this further by treating the entire system as a configurable fabric rather than a fixed layout.

Compute, memory, and interconnect elements become modular. Architects gain the ability to tune for cost, power, bandwidth, and latency based on actual workload needs, not worst-case assumptions baked in at the start of a design cycle. The benefit, as the keynote described it, is flexibility without sacrificing rigor.

This is a meaningful distinction. Flexibility in hardware design has historically come with a cost in precision or predictability. The argument here is that chiplet quilting, when paired with the right tooling, can deliver both.

Tooling Turns Theory Into Practice

How Oxmiq Labs Models Chiplet-Based Systems

At Oxmiq Labs, the work centers on building tools that allow teams to model chiplet-based systems using real parameters. Raja Koduri walked through what that modeling involves in practice, covering a range of inputs that matter at the system level:

  • Die size and geometry
  • Memory bandwidth
  • Interconnect bandwidth
  • Power consumption
  • Cost inputs
  • Picojoules per bit
  • Picojoules per operation
  • Inference workload profiles

The reference system for the demonstration was based on a current flagship GPU platform, the NDGX B200. That was compared with a hypothetical quilted system designed around tighter memory coupling and higher bandwidth interconnects.

The results illustrated the direction clearly: order-of-magnitude gains in throughput and order-of-magnitude reductions in energy per token. The point was not a product claim. It was a demonstration of architectural leverage. When memory sits closer to compute and interconnect bottlenecks shrink, inference economics shift in ways that matter at the scale the industry is now approaching.

The Takeaway From the Keynote

Raja Koduri closed with three ideas that summarized the position he had built across the talk. The age of inference has arrived. Details at the physics level decide winners. Chiplets are fun, and they offer a practical path forward.

That last point, the idea that chiplets are fun, is worth pausing on. It signals something about the spirit of the talk. This was not a presentation built on alarm or abstraction. It was an invitation to engage with a design problem that is genuinely interesting, technically demanding, and consequential in ways that will be felt across the industry for the rest of the decade.

For teams designing silicon for inference, the message from the keynote was clear: the architecture choices that get made now are not just technical preferences. They are bets on what the infrastructure of AI will look like at scale. The physics, the economics, and the tooling are all pointing in the same direction.

About Raja Koduri

Raja Koduri is a semiconductor and graphics architecture executive with extensive experience across chip design and AI computing infrastructure. He delivered the keynote Chiplet Quilting for the Age of Inference at an industry event themed AI Everywhere and is associated with Oxmiq Labs, where tooling for chiplet-based system modeling is being developed. His work focuses on the intersection of silicon architecture, energy efficiency, and the infrastructure demands of large-scale AI inference.