Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI

At 100,000-accelerator scale, something fails multiple times an hour, which is why Google's Chief Technologist for AI Infrastructure Amin Vahdat thinks FLOPS is a vanity metric. The metric that matters is what he calls goodput: the useful work a workload actually delivers through real failures. Amin explains why long-horizon agents are sending demand for CPUs and storage through the roof alongside accelerators, how optical circuit switches reroute light to a spare rack in milliseconds, and why Google would rather wait on a utility than build its own gigawatt. We also cover orbital compute and the multi-megawatt rack of 2036.
Watch Now
Transcript
Chapters
Introduction
Amin Vahdat: In hardware, as you know, there is this opportunity where the more you specialize to a particular workload, the less flexible it is, but the faster and the more power-efficient the hardware is going to be. So it is this art, and it's this projection of what you're designing to and how persistent that workload is. In other words, if it's going to go away after a month or two months or three months, even if it's big for those three months, you've got a really narrow window to intercept it. So it has to be somewhat durable, and you have to be able to project exactly what you can get for specializing to it.
Main conversation
Sonya Huang: Thrilled to welcome Amin Vahdat to the show. Amin, thank you for joining us today.
Amin Vahdat: It's really exciting to be here. I'm really looking forward to it.
Sonya Huang: I am very excited for today's topic, because we are in the middle of the biggest CapEx buildout in human history, and you are in the middle of it. Google alone is expected to spend more than $200 billion on CapEx this year, most of which is going into building data centers. You're at the center of it all. You were named the head of Google's AI infra at the end of last year, so you are the man spearheading the efforts of one of the most capital-intensive buildouts in human history. I am really excited to get into it with you today.
Amin Vahdat: It's a huge year, for sure—across the industry, including at Google. I don't think we've ever seen anything like this, honestly. Certainly not at Google. But as you said, I think in the history of humanity, in terms of the buildout and the pace of transformation, it's really unbelievable.
Sonya Huang: Before we get into it, maybe just a crash course for the audience. What is an AI data center, and how is it different from a non-AI data center?
Amin Vahdat: It's a very good question, in that an AI data center and a non-AI data center do have a lot of similarities, actually. They're not radically different. It consists of concrete, an enclosure. It consists of electrical yards, mechanical yards, cooling. You have row after row of power being distributed. There's a huge amount of network infrastructure there—in other words, we connect large amounts of compute to one another. A fair amount of storage infrastructure goes in there.
I think the big difference that we're seeing with AI infrastructure is specialization. In the past, when we were building a data center, it really was a 25-, 30-year building investment, and we were thinking about how it's going to be evolving over that 25-, 30-year period. You might have servers go into it, you might have networking, storage, you might have some accelerators—GPUs, TPUs, whatever they might be—going into it. But it's a 25-, 30-year planning horizon, and the lifetime of the hardware might be six years. So there are many generations that we have to plan for.
An AI data center oftentimes is going to be more purpose-built. In other words, we are oftentimes co-designing the building with the hardware that might go into it. We might be saying, "You know what, actually, we're not going to put a lot of storage into this building." Why? Because a storage rack might have 10, 20, 30, 40 kilowatts of power. You put that next to a CPU rack or a GPU rack that easily is hitting hundreds of kilowatts today, and might hit megawatts—people are talking about that in the next two years. Designing a building that might take 30 storage racks in a row versus one or two AI racks in that same row is a very, very different design. Just think of it in terms of the size. Think of it in terms of the power and how that would be distributed across the building. Networking would be the same: the amount of networking that you would need for a storage rack is tiny, especially if it's hard drives, compared to an AI rack.
So if you're trying to make it fully fungible, you'll probably make it too big and too overbuilt—fungible over, certainly, a 30-year period. An AI data center likely is going to be much more purpose-built, co-designed with the hardware, even to the point of cooling and power distribution. So a lot more co-optimization.
Sonya Huang: You all delivered a big Vera Rubin cluster to one of my portfolio companies, Ineffable Intelligence, and I saw the photographs of it as it went out. That is a thing of beauty.
Amin Vahdat: It is. It is.
Sonya Huang: It's one of those things where, similar to looking at the biggest construction projects in human history, you look at that data center and it's like, wow, that is a monument to what mankind can do.
Amin Vahdat: We've got a nice picture of it. And this is just a couple of racks. The cabling, the fiber distribution for it, is really beautiful. Beauty is in the eye of the beholder, but for folks like me and probably yourself and others in the audience as well, it is a thing of beauty. We put this picture up on social media, and people love the fact that it's Ineffable—they're huge fans of what they're doing, fantastic team—and that got a lot of attention. "Okay, great rack for Ineffable." But actually, people just loved seeing the pictures of the fiber and the sort of fractal nature of the fiber. One of the most popular posts that we've ever had, actually, was around that. So yeah, absolutely, we're really, really excited about that.
Sonya Huang: Yeah, I had goosebumps seeing the picture. How do you measure and hold yourself accountable in the middle of this big buildout? We were talking right before the show about how FLOPS is a vanity metric and you prefer an alternate metric. Can you say more?
Amin Vahdat: Yeah. So one thing to note is that whether it's FLOPS or your other favorite chip-centric metric, these are in theory. In other words, under some conditions, for whatever chip you have, this is the maximum amount of FLOPS that you can deliver. What we really care about in the end is the performance delivered by workload. That's what we're looking at. And it happens, but actually relatively rarely, that that performance is determined by a single chip—its FLOPS, its HBM, its SRAM capacity. Those are all super important, but it might be how two, four, eight, 16, 1,000, 10,000 or more of these chips compose together—not just the accelerators, whether CPUs or GPUs, but then the CPUs that might feed them the data, the network that connects them all together. So what is the workload that you're running, and what is the performance of that workload?
One measure that is interesting is your FLOPS utilization. So, for example, for a particular workload, if you have in theory a teraflop or petaflop capable for your workload, what fraction of that are you delivering? Now, that's a measure of goodput. What is goodput? You can think of throughput, which is a well-known term. That's the throughput that is possible. But now let's consider some other considerations. One is the slowdown of the workload—again, just enhance the workload. But other aspects of it really hit our reliability.
So in terms of how we hold ourselves accountable: if we have a chip failing for a synchronous workload—and many of these workloads, whether training or serving or agentic workloads, are synchronous—there are many, many components that are working together simultaneously. Now, if you have, let's say, 1,000, 10,000, 100,000 of these components working together simultaneously, and they really need to coordinate at microsecond and millisecond granularity, one of them fails, and it might actually bring the whole thing to a stop. Because everyone is counting on everyone else to do their part of the job in order to come up with the answer to a really tough question. One of them stops. Okay, now we have to figure out what happened. Which one stopped? What's a checkpoint that we have of the computation at some previous state and time? How do we restore that checkpoint? How do we restart? Worst case, we have to restart from the beginning. That would be really bad, but that can happen in certain cases, especially more on the inference side.
The point here is that if you then have to go back and redo a bunch of computation, if you have to pause and wait to figure out what happened with the failure, all that work is work. It's not actually helping you get the answer. Let's say that you're working out a problem on paper. Step one, step two, step three, step four. If you have to go back to step one because you made a mistake, yes, you're still doing work. That's throughput. You're doing work. But what is the goodput in delivering your answer? It's basically the total amount of time it took you to solve the problem. That's what matters. And if you have failures, if you have failure recovery, if you have whatever it is that's interrupting your work, that's part of the issue.
So how do we hold ourselves accountable? It's delivered goodput. It's not theoretical benchmark throughput, or goodput in theory, but actually, for a workload, for real failure conditions, what's happening. And the unfortunate reality is that at 100,000-accelerator scale—
Sonya Huang: I was going to say, at that scale something's failing all the time, right?
Amin Vahdat: All the time. And each one of these chips, as you're aware, is a wonder of nature. They're at the very bleeding edge of what's possible to manufacture. To me, it's stunning. And now, actually, these chips aren't just one chip. They're packages made up of, in many cases, two, four, eight, maybe more chiplets that are composing together. And then of course HBM off to the side, network connectivity, maybe things like co-packaged optics. Not a criticism—lots of things can fail. And if you've got 100,000 of them, something is going to fail. You have to be prepared for that. Detect that in near real time, recover from that in near real time. Doing all that—the telemetry problem is massive. It's like finding a needle in the haystack continuously, across certainly seconds and minutes, but for some of these jobs, hours and days and even weeks. Just continuous and online. How we hold ourselves accountable is the goodput that we deliver for the workloads that actually matter in the data center.
Sonya Huang: Is "goodput" a Google term, or is it an industry term?
Amin Vahdat: It's a Google term. But I think more and more people across the industry are starting to pick it up.
Sonya Huang: And then just to calibrate me, at 100,000-accelerator scale, are we talking it's going to fail once a minute, once a day? How often?
Amin Vahdat: 100,000 accelerators. Let me see. I think it is definitely going to be, at that scale, multiple times a day, and perhaps, depending on the exact configuration, multiple times an hour something is going to fail.
Sonya Huang: And what are the most common failure reasons?
Amin Vahdat: This is the problem, actually. It's a very good question. If there were a common failure reason, we would have figured it out and fixed it. It is a long tail of constant discovery. Whenever we have a new product that's being introduced, there's going to be something that hits us—in many cases, frankly, because it's at the very bleeding edge. It might be network-related. It might be how we connect these things together at super high speed. That might be something to do with the hardware. Those things we work through. But then a lot of the issues could be software.
So this is the other aspect of how we hold ourselves accountable. The chip might be capable of a certain level of FLOPS, but if you have a compiler bug, a runtime bug, a model issue, an operating system issue, it doesn't matter—that's going to impact the end-to-end system performance. So you might have perfect hardware, fully reliable, but then you might have software issues that hurt you.
Sonya Huang: Is there a standard reference stack from the accelerator companies—from Nvidia or from the TPU team—of "here is the optimal system to build around our accelerators, and as long as you build to that system, you're good"? Or how much of your own data center design do you have to do above and beyond what the semiconductor companies give you?
Amin Vahdat: Right. So there's a reference stack. And I would say that Nvidia is an incredible whole-systems company. They are obviously a semiconductor company, but they're not just a semiconductor company. They give you a really, really strong reference stack. But the way I'd put it is, most of our customers leverage that reference stack, but many also specialize. In other words, they might find that, as natural for their particular use case, they have an optimization opportunity or something different that they need to do, and they're going to do that. Similarly, for the TPU side, we have a reference stack. Most people will leverage that, but many will specialize it as well.
Sonya Huang: Got it. And then in terms of overall capacity, you've told your teams that Google has to roughly double the serving capacity every six months or so. Is that right?
Amin Vahdat: Well, to be clear, this is in terms of effective available capacity. The way in the end that we look at it, from a serving perspective, is token generation capability. So capacity—and this is where I would say it's a combination of software and hardware. The hardware might have some, quote unquote, inherent level of FLOPS. What I'm saying here is not necessarily that you have to double the number of FLOPS every six months. That's one path. You have to double the capability of that hardware to generate tokens every six months. And as much or more of that is going to come from software as it is from hardware. In other words, it could be a model optimization that delivers that. It could be some runtime optimization that delivers that. It's really probably going to be dozens or hundreds of individual optimizations that are just landing again and again and again to make it all possible. But yes, the rate of capacity improvement is incredible.
Sonya Huang: We've had a few years of data center buildout now and a few years of model progress and software progress. What has been the empirical breakdown of how much capacity—maybe as measured by intelligence per watt—how much capacity increase has come from the silicon itself versus the models versus other software, and any other kind of big components?
Amin Vahdat: Yeah, it's a really good question. I don't have the exact breakdown, but I would say that in our experience, most of the benefits come from model-side improvements, in terms of intelligence per watt. And by the way, intelligence per watt is a fantastic metric, really focusing. For us, it's goodput per watt, and we can come back to why the watts need to be the denominator. Goodput could well be a measure of the intelligence delivered per watt—in other words, goodput as a workload-specific metric.
But I would say most of the gains often come from the model side, the software system. Quite a bit. Why? Because they're able to actually make sure that the hardware that you have is being used effectively. And then the hardware—it is pretty stunning. In other words, we're living in a world right now where 2x or more year-over-year performance improvement is absolutely possible. And so the hardware does support it. It's, if you want, a free multiplier—not free, but everyone else above it can count on it in terms of that being a multiplier that lifts everyone else up year over year.
Sonya Huang: I'd love to talk a bit about the TPU program and co-design. Google started building custom silicon more than a decade ago. It was a contrarian call at the time. How has the TPU program evolved?
Amin Vahdat: Quite a lot. When the program started in 2013, it was really a contrarian call. And I think that at the time—it's hard to put yourself back into the moment—2013 conventional wisdom, all the smartest, wisest people would say, you don't build a custom-built accelerator for a single workload. Why? Because Moore's Law is there. Doubling of performance every, whatever it is, 18 or 24 months is there. You get to leverage standard programming models, all your C++ code, Java, Python, whatever it might be.
Sonya Huang: The bitter lesson of chips.
Amin Vahdat: Yes, exactly. The bitter lesson of chips is that specialization never wins. But in this particular case, we had this one application, or a few, a small number of them, that would benefit tremendously, and that would require an unimaginable amount of general-purpose CPU to support. So really, in 2013, it was a bet. And there were quite a few people, even within the company, I would say, who were not sure that it would work out. So it was a bet. It turned out to be a massively successful bet.
The first chip was all about inference. The second chip was, hey, actually, we can take the same idea and build a training chip. From there, it got picked up for more and more use cases. Around the time the second chip came out, transformers were invented. This was a major, major moment, and that actually completely shifted the TPU program. We of course learned that recommender systems could run really, really well on the TPUs, and so ads and related use cases came along.
So I would say that the evolution has been one of expansion in scope and impact. In other words, it started with a really impactful use case of a couple of applications for inference serving—this was language translation primarily, and voice recognition—to then training, to then transformers, recommender systems, and continuing to generalize as the GenAI moment took hold, to ever larger, ever more scalable systems as well.
Sonya Huang: You mentioned the moment the transformer came through being a major moment, and I'm curious to explore—the TPU was a specific architecture, but it wasn't specific to the transformer. How do you straddle the fine line of how specific you want your chip to be for the workload?
Amin Vahdat: This is a great question, and I think it really comes down to the applicability of the chip to a particular workload. In the end, we've been thinking about this for every generation: would we further specialize? The question that we were facing, let's say two years ago or a little bit more, was: in 2026, should we have two chips or one? We could have one chip that could do inference and training simultaneously and do both quite well. Or we could have two chips, one that was further specialized for inference and a second one that was further specialized for training.
This analysis and this work led to the release of two chips this year: TPU 8i for inference, TPU 8t for training. And in the end, what we realized is that it wasn't the case necessarily a couple of years ago, but by '26 we saw inference and serving really taking off. And so having a chip that would be significantly faster for serving—that we thought might be 30, 40, 50, 60% of the market in its lifetime—started making a lot of sense for us. Relative to, if we projected inference to be 2% or 5% of the market, even if that specialized chip is, let's say, 2x faster, it might not make sense, right? Because you'd actually go for the one general-purpose chip that's not fully optimized, but that's okay, because you now have the uniformity.
So really it does come down to a calculus question of: okay, would you further specialize? How big is that workload? How big is that workload projected to be in a couple of, three or four, years? Is this sustained growth for a particular workload? In hardware, as you know, there is this opportunity where the more you specialize to a particular workload, the less flexible it is, but the faster and the more power-efficient the hardware is going to be. So it is this art, and it's this projection of what you're designing to and how persistent that workload is. In other words, if it's going to go away after a month or two months or three months, even if it's big for those three months, you've got a really narrow window to intercept it. So it has to be somewhat durable, and you have to be able to project exactly what you can get for specializing to it.
Sonya Huang: So the rough tradeoff is there's a large fixed cost to support a new program—
Amin Vahdat: Yes.
Sonya Huang: —and you basically have to think that there's going to be enough demand for that specific program to justify the cost.
Amin Vahdat: Exactly. And to some extent—for example, for our 8i and 8t chips, both chips can do the other workload. This is also key. If 8i could only do inference and couldn't do training at all—in other words, it had zero performance for training—and then vice versa, if 8t was amazing at training and had zero performance for inference, that would have also been a key limitation. Why? Because we would have to predict ahead of time, over a six-year period, the lifetime of the hardware, exactly how much we need of each. In this case, it was nice because both chips are better at what they're specialized for, but if you had leftover capacity in one place or the other, both can actually do the other's job. Depending on the specialization you do, you might get so specialized that you actually limit yourself from being flexible, fungible.
Sonya Huang: And if the rough tradeoff is the size of the program—on the other end, it seems to me that so much of the modern AI market is transformer-based. So maybe just to ask a provocative question, why not just burn the transformer architecture into the chip?
Amin Vahdat: Yeah. So I think it's transformer-based, but then the next level of question would be: transformers, in the end, are about vector and matrix multiply operations, and a softmax. There's a range of essentially linear algebra primitives. We have those, and others do as well, roughly baked into the hardware. It's all transformer-based, but then there's your model architecture. Okay, how many layers do you have? How do you go through the layers? How do you go across the layers? For each dimension, what exact shape of matrices and vectors are you applying? You could go further and specialize to not just the transformer but your model. That's the next level of specialization. I think there are a number of companies out there that are thinking about that. I think it's a very, very interesting direction as well.
Sonya Huang: Okay. So at this point in time, do your customers view TPU and GPU as roughly fungible equivalents, or is there a certain set of problems that is better suited for one or the other?
Amin Vahdat: There's for sure a set of problems that are better suited for one or the other. GPUs are more general-purpose than TPUs, for one. That is clear. They're incredible products. At Google and Google Cloud, we sell a lot of GPUs. We use GPUs internally. But I think it really then comes down to the specifics of your problem. There is overlap between the two, but our customers then evaluate the workload and evaluate their options. What we like to do at Google is give our customers choice. In other words, we want to have the right solution for their needs, and of course provide the solution that best meets their needs for them.
Sonya Huang: What is the case for co-design, from the chip to the network to the software? And then what is the case against co-design?
Amin Vahdat: Yeah. So there are huge optimization opportunities with co-design. What you can imagine is that if you have an N-layer stack and you want to be able to pick and choose whatever component you want—let's say you're running across many clouds, or you're running across many pieces of hardware, many pieces of software—you could design abstraction layers that would basically say, I can run on any hardware, I can run on any software, I can run on any network topology, any amount of network that you give me, no problem. And my system is going to be fully adaptive.
So now you have a great capability. You can move anywhere. New capacity becomes available overnight, and you're up and running, because you've designed your system that way, to actually be able to take advantage of anything. You have not hard-coded or specialized at all to anyone's particular infrastructure. The downside is you probably leave a lot of performance on the table if you become fully fungible, fully flexible to anybody's hardware, software, network, storage, compute stack.
So the case for co-designing is that between each one of those layers, there's a big impedance mismatch if you try to get full generality. There might be 10%, 20%, 2x across each of these layers. You start multiplying those optimization opportunities through, and all of a sudden you're left with a big end-to-end opportunity in terms of, whatever you want to say, intelligence per watt or output per watt, that you can leverage—all the way down to power delivery and power availability, software optimizations, et cetera. So the pros are you can run anywhere, anytime. You have no lock-in. The cons are you're leaving, in all likelihood, significant amounts of performance on the table.
Sonya Huang: So my understanding is that if you take OpenAI and Anthropic, OpenAI was primarily building on a homogeneous compute stack, and Anthropic was building a roughly more heterogeneous compute stack. Do you think that co-design is part of the reason why they converged on what is rumored to be very different architectures for their models?
Amin Vahdat: I don't want to speculate in terms of what OpenAI and Anthropic are doing. It's one possibility, but without knowing the details of what they're doing, I would imagine that there could be many reasons for why they wind up with different architectures.
Sonya Huang: Well, what about for Google, then? I'm curious what the working relationship looks like between your team and DeepMind. Who's in the room at what stage of model development, and how are you making decisions together on co-design?
Amin Vahdat: It's one of the most fun and frankly gratifying parts of being at Google: the opportunity to work really shoulder to shoulder with the DeepMind team in terms of co-design of our hardware and models. There's a third element to it, in that we also get to extend that with the consumer services and Cloud. I'll put that aside for a moment; I can come back to it. But with respect to DeepMind, it really is a deep partnership.
From the past, I can give you examples where they have come up with model optimizations, let's say to transformers or to particular math that they might want to do, and we might have a chip in progress, not quite done. But then they say, "Oh my gosh, if we had hardware support for this, our end-to-end training or serving might get significantly faster, more efficient. What would it take to actually change the hardware definition that we're in execution on to accommodate this?" And this then leads to our engineers and their researchers getting together intensely in the same room for a few days, a week, two weeks, saying, "Okay, yes, we can do this." More likely: "We couldn't quite do what you wanted, but we can do this other thing. And then maybe you can change your model architecture in this other direction. That gives you 98%. That gives us 90% of what we're looking for." And yes, we can then go back and intercept the hardware. Or we can say, "You know what? We're going to delay the tape-out by a week or two weeks, because wow, for this level of benefit, it's totally worth it."
Similarly, when we're projecting our roadmap out—we have many generations of chips that are essentially in progress at any point in time. We have the chips that are in production. That's one. We have the chips that we're working on getting into production; they're already back from the manufacturer, and we're debugging them, making them work. We have the chips that are in implementation that are about to tape out and go to the manufacturer. We have the chips that are in design phase. And then we have the chips that are in concept phase. So it's really this five-, six-stage, many-year pipeline from introduction to end of life.
The collaboration with DeepMind for in-production is significant, because we get to work together in maximizing delivered intelligence, or delivered goodput, per watt. We know exactly what's happening in the model, and we know exactly what's happening in the hardware and everything in between. We also get to collaborate deeply on the chips that are just about to tape out. Why? Because we can intercept, and we can make changes to the chip—literally the chip architecture—in flight, which would be somewhere between hard and impossible to do if we were working across company boundaries. Not impossible, but it would be much harder for us to say, "Oh my gosh, we're a few weeks or a few months away from getting this chip done, and now let's get in the same room, shoulder to shoulder, and figure out if we should disrupt the program." It's possible, but harder, is what I would say.
And then of course, for the chips that are in design, we have many architectures we can evaluate together. We can ask our colleagues in DeepMind, where do you see model architectures going in two or three years' time? Here's the Pareto of things that we can do. Here's the Pareto of where model architecture is going. We actually have deep and significant simulation infrastructure that can predict how the workloads are going to map to different hardware architectures. Deep and fast iteration. It really isn't the teams working separately. It's in the same building, same rooms, many times. Deep daily interaction. I talk to, whether it's Koray or Demis, multiple times a week. So it really is a super fun aspect of the work.
Sonya Huang: That's awesome. And eventually their models will help with chip design.
Amin Vahdat: We're using Gemini to design hardware for future Geminis as well.
Sonya Huang: That's really cool. I want to come back to that a little bit later. On the collaboration itself—is there an impedance mismatch still? I figure the cycle times are probably just slower when you're dealing with hardware than your DeepMind folks get to deal with on the software side. And I figure your planning cycles are much further ahead. So how much room do you actually have to adjust?
Amin Vahdat: Yeah, it's a really good question. We are planning hardware two, three, four, five years in advance, no question. In other words, we talked about TPU 8i and 8t that we just announced, but you can imagine that 9, 10, and maybe some others are in concept, in execution, in something else, and they might be years out. If you're working with model architecture, by default you're not going to be thinking years out.
But I think this is also the great thing about how the company has grown up together. In other words, Google Research and DeepMind and transformers being invented—all of this was happening in these overlapping rooms as well. So there's a whole generation of researchers who've grown accustomed to being able to influence the hardware, and knowing that the hardware is operating on multi-year cycles. They also know that, look, if they have a small tweak that's going to deliver 1% or 0.5% or something like that for a chip that's about to tape out, they're probably not going to come to us, because they know enough to know that it's not like software, where you can just do a change list and it's going to ship to production in two weeks' time. Stopping a tape-out is a big deal. But they also know if they've got a really good opportunity, yes, we're absolutely going to work together to figure out if we can get it in there.
So as I said, it's multiplicative in terms of where the benefits come from. The hardware does lift all the tides, right? And so significant portions of the DeepMind team—and we're so grateful for it; it's an amazing team—significant portions of it are thinking about, how do I influence the roadmap? Because it's actually a pipeline. The idea I had two years ago is in production now, and it's helping all the workloads at Google go faster. That's a good feeling.
Sonya Huang: Totally. It seems like an impossible task, though, to predict in five years what workloads will be most common and what algorithmic breakthroughs will have occurred. Doesn't it seem like an impossible task?
Amin Vahdat: I hear you that it would seem like an impossible task. But here's the awesome thing—and actually, we're working on writing this up in great, great detail, and it's a lot of fun to do it. The stunning thing is that the TPU architecture, at a medium level of detail—not at a super high level of detail, at a medium level of detail—hasn't really changed since TPU v1.
One way to look at it is the instruction set architecture for a CPU. You've got loads and you've got stores and you've got adds and you've got subtracts and branches. And yet what has the software on top of it done over that period of time? Same thing with TPUs. We have some fundamental instructions and fundamental primitives. Of course we've extended it—it's not like the instruction set hasn't changed at all—but the primitives are there, in terms of specializing the numerics to very big, very large matrix multiply units; a SparseCore that manages vector operations and scatter-gather operations. There are five or six things that really define it. Remote load-store—that's another one, actually. We can read and write remote memory through our ICI network. The fundamentals have been there, and they've extended to many generations of models, many generations of even deep neural network algorithms and model structures.
Sonya Huang: Okay. Speaking of shifting workloads, it seems like one of the biggest changes in workloads over the last year or so—I think it really started at the beginning of this calendar year—was the rise of the long-horizon agent.
Amin Vahdat: Yes.
Sonya Huang: And I would guess that's a very different shape of workload than the quick-turn LLM conversations of years past. What does that mean in terms of data center needs?
Amin Vahdat: Yeah. So I think there are two huge aspects to this. One is it's no longer human-to-accelerator interaction. In other words, when you are typing at a prompt in a web browser or on your phone or whatever, of course in response to your prompt there's going to be a bunch of work that happens. But then when the response comes back, you've got to read it. You've got to think about it. And then maybe you have a follow-up. That's going to be multiple seconds of interaction time. Now, in long-horizon agents, there's no human in the loop that is going to naturally rate-limit how quickly requests are going to go to the model. So what went from seconds, maybe tens of seconds, in terms of interaction time is now going into perhaps milliseconds. As soon as I get a response back, I can parse it, I can reason about it perhaps a bit, and I can figure out what my next request is going to be. So that's big change one.
Big change two is all that reasoning and all that parsing is going to probably happen on a CPU, and that CPU is then going to have to think about, okay, what other state do I need to go gather before making my next prompt back into the model? In other words, I've learned something from this response. I'm going to take another step, but I actually need to go grab some context—maybe from DRAM local to me, maybe from someone else's DRAM on another CPU, maybe on SSD, or maybe on HDD somewhere else. So now a huge amount of orchestration has to take place as well. So the design actually is changing pretty significantly, where the demand for accelerated compute is going up, but the demand for CPU and networking and storage—traditional data center CPU—is also going through the roof.
Sonya Huang: So does that mean you're putting more CPUs alongside your GPU racks, then?
Amin Vahdat: GPU, TPU racks—this is the key question, going back to this question of optimization and specialization. If we start putting lots of CPU racks—and it is a question—next to the GPUs and TPUs, that means that actually we can't fully specialize to the density and network requirements, let's say, of a TPU rack relative to a CPU rack. A TPU rack is going to be more dense than a CPU rack. It's probably going to need more networking than a CPU rack. So in other words, now our building design is going to change.
Another option is to maintain your uniformity. Put your GPUs or TPUs all in one building, and in the building next door, maybe put your CPUs. And maybe, by the way, the hard drives have to be in another building on the other side, because they have yet another set of requirements. Now you need networking between these buildings at a pretty significant level. Once you leave a building, the networking complexity goes up significantly, from a reliability perspective, from a cost perspective. Latency goes up—probably acceptably, but still, now you might go to hundreds of microseconds, potentially more, with queuing between the components. So the considerations do change in pretty interesting ways.
Sonya Huang: Your job is hard.
Amin Vahdat: It's fun. Yeah.
Sonya Huang: What's happening on the networking side? I've heard that Google's always been at the forefront of the newest in networking, including optical. Can you say a word on the state of optical networking?
Amin Vahdat: We at Google, this was probably 15, 16 years ago, were among the first to bring essentially what's called wavelength-division multiplexing—where you could put multiple signals on a single fiber—within the data center. And we actually leverage that for all of our communication between racks. At the same time, we introduced a technology, along with the wavelength-division multiplexing, called optical circuit switching. Essentially what optical circuit switching does, in contrast with traditional electrical packet switching, is it transmits and moves data entirely in the optical domain.
Let me describe the technology that underpins it. In a packet switch, you would take a packet. It has a header. You would look at it in the electrical domain, figure out where the packet is headed. In its header, it might have an IP address that says, okay, where am I headed? You look up, for that destination, in a table: what port do I forward it along? So basically billions and billions, perhaps trillions, of packets coming through at super high speed per second, and you're forwarding them along. Optical circuit switching says, I'm not touching these bits in the electrical domain. What I'm going to do is figure out, for an input port, which output port to send the light to. There are multiple ways to do this. The one that we started with is called MEMS switches—micro-electrical motors that basically control mirrors in 3D. So we can now programmatically take a box that might have, you pick, 128 ports, 256 ports, some number like that, and we can configure every input port for a fiber that comes into it, map it to an output port, and then change the rotation of mirrors, where the light literally shines down on these mirrors and gets reflected to the right output port.
Initially, the reason we did this was twofold. One was to create locality between groups of racks. Let's say I had a compute cluster and a storage cluster, and both of them were in support of—again, making it up—search. We knew that these two clusters would talk to each other a lot, so we would configure the mirrors to create shortcuts, purely optically, between those two clusters of racks. That was reason number one. Reason number two was we wanted to be able to expand the network and contract the network without actually moving any fiber. Without going into the details—I could draw this on a board—the optical circuit switch would allow you to actually reconfigure the spine of the network, to expand it or shrink it, without a human being having to do anything other than be, literally, a controller that would manage that.
Now fast-forward to TPUs. A TPU has a torus topology that connects all the TPUs to one another directly. I talked about throughput and goodput earlier. One of the things that we can do is, if we have a TPU rack that fails, we can replace it with another TPU rack without moving any fiber. Again, I'd have to wave my hands or draw on the whiteboard, but essentially we can say we have a spare rack available at all times, and when a rack fails, we redirect the light to that new rack. And that can be done in milliseconds.
Sonya Huang: Why do you have the fiber at all, then?
Amin Vahdat: As opposed to fully free-space? Yeah, it's a very good question. Attenuation loss, and bandwidth would drop dramatically. Furthermore, across the range of a very large building, aiming everything without the benefit of fiber to connect it all together in 3D would be challenging to possibly impossible. We've talked about it, actually, and there have been some really interesting discussions. But yes, it's mostly in fiber. But then when they hit the optical circuit switch, essentially the light then literally shines down on these tiny chips. So that's one big direction of networking, but there's a lot. I would say networking is exploding in terms of its capabilities and need, frankly, in the data center.
Sonya Huang: So cool. We could do an entire conversation.
Amin Vahdat: Yes, it is very, very cool.
Sonya Huang: Even going back to the picture from Ineffable, it was all the cables that—
Amin Vahdat: Exactly. Yeah. And those cables eventually, some of them, wind up at optical circuit switches, which in our data centers makes sense.
Sonya Huang: Okay. I want to flip to talking about power. You've been talking about goodput per watt and other units per watt. It makes me think per watt means power is kind of the binding constraint, or the scarce constraint, or the expensive constraint in some way.
Amin Vahdat: So I get asked this question: what is the biggest constraint that we face? And the reality is there is no single biggest constraint that we face. They're all constraints. They're all super hard, and they shift continuously. But if I had to answer fundamentally, I would say that power is the single most fundamental constraint that we face. Everything else, it seems like we know how to solve them, and it's a question of solving them over some period of time. Power—I think the way that you put it is really nice. It's a binding, long-term issue. We do not have a—I mean, nuclear, perhaps nuclear is going to—abundant clean energy would solve a lot of problems, for sure, when that happens. And when it happens at scale is still unknown.
Sonya Huang: And so how does it work in practice? You're standing up a new data center. It needs a gigawatt of power. I imagine you can't just call PG&E and say, "Hey, please send a gigawatt of power." So what does provisioning power actually look like? And are you having to vertically integrate all the way down to doing your own turbines? How do you solve the power bottleneck?
Amin Vahdat: Yeah, it's a big question, an important question. Our preferred model at Google is always to be utility-connected, to be grid-connected. So it's the equivalent of calling your—
Sonya Huang: Oh, so you can just call PG&E?
Amin Vahdat: Exactly. You call, very politely, your favorite utility wherever it is that your data center is. And of course you give them many, many years of notice. In other words, if we're talking about gigawatt scale, it's not something that you can say, "Hey, I need a gigawatt tomorrow. When can you start billing me?" It's something that we plan together.
For us, it's something that we also take very seriously, from the perspective of ensuring that when we work with these utilities, the costs of putting that infrastructure in place—and we could get into a long conversation just on this topic as well—that we cover those costs. Because with the way that billing works, it actually could be that by the act of the utility building out capacity for us, let's say, other people's rates could go up, in theory. What we ensure is that actually, let's say the transmission lines that have to be built or upgraded, additional utility base stations, et cetera—that we pay for those as well.
So there's a long planning process. It can absolutely be the case that—let's say that we need a gigawatt in, I'll make up a date, 2028, and the utility can get us a gigawatt in 2029. They can get us, let's say, 700 megawatts in 2028. So now we might be left with a question of how we cover those 300 megawatts. Well, one answer is just wait. Another answer is to say, okay, well, would we figure out how we generate some subset of that power ourselves? Maybe that would be with solar cells, or with batteries as backup. Would it be from other sources? So then we again work with the utility. It might also be a combination, where we might maintain some power generation local, and have the capability, even when the utility comes online fully—let's say at the gigawatt scale—where we could actually provide power back to the grid, so having that when they need it as well. One way to look at it is there might be the two weeks of the year—maybe it's the hottest two weeks of the year—where there's a huge amount of residential demand. If we have some local generation, we can then provide that back to the grid as well. So it really is working over multiple years with the utilities, for us.
Sonya Huang: Why is your preference to work with the utilities, as opposed to vertically integrating yourselves?
Amin Vahdat: The main reason is flexibility and uplift on both sides. One way to look at it is statistical multiplexing, or, if you want, the law of large numbers. If we need to have a gigawatt of power—let's say we wanted to have that with 99.99%-plus reliability—that means we have to build two gigawatts of power. At that level, 99.99 or 99.999, you have to have one-plus-one redundancy. That gets expensive. And then making sure that that's going to be ideally a clean energy source, probably right next to our data center—that can also be challenging.
Now, if we partner with [a utility], again, maybe we have some amount of that that we can bring ourselves, that we can give to the grid when they need it, and when we need it less, we can take their power. In other words, by leveraging that statistical multiplexing over a much larger base, actually everybody wins. We win, the grid wins, residences win. We prefer that. In rare, rare cases, we will do it behind the meter. But even then, we're doing it under the assumption that we're going to work with the utility, where it might be a year later that we want to be connected to the grid. Again, it's an uplift for us. It's an uplift for the grid.
Sonya Huang: Makes sense. How do you decide how big to make a data center?
Amin Vahdat: Yeah, that's an art, and it's a source of significant debate. At some point—I remember even 10-plus years ago, 15 years ago—there was a big debate at Google: should we just put everything into one data center? Back then, it was going to be a gigawatt, which was—
Sonya Huang: Simpler times.
Amin Vahdat: Yeah, different times. 10, 15 years ago, we were going to have a gigawatt data center. So one obvious concern there is single point of failure, right? If you're looking at things from a 30-year perspective, that's a huge concern. On the other hand, if you're talking about training workloads today, bigger is better. In other words, from a networking perspective, you want to have things concentrated in as small a distance as possible. But then again, two issues: single point of failure, but then also now power availability. While a gigawatt was huge 10 or 15 years ago, it was imaginable. There's no way anywhere in the country or the world we're going to get the total requirements of Google built in one place. It's just not going to be possible.
So now, what's the optimum size? Again, we have models for this, simulators, et cetera. But it also depends. Look, in some places we're going to be at the edge of the network, or we're going to be in-country—that might be tens of megawatts. We actually partner with ISPs; that might be a rack, literally. It could be a rack. For a training cluster, maybe that's going to be closer to a gigawatt. Some other sites might be hundreds of megawatts.
Sonya Huang: I'd love to understand how you think about portfolio lifecycle management. My guess would be that the biggest, newest clusters are used for training the latest frontier model, and then you recycle the older stuff and run inference on it. Is that a fair framework? Are you also standing up inference-specific clusters? How does that all work?
Amin Vahdat: It's a very reasonable framework that you're thinking about, and I think it makes a lot of sense. But I would say that the demand for inference is such that we can't just rely on whatever older training clusters that are no longer being fully used for training as the basis for inference. And also, if you think about it, we might centralize, again, in a particular year, into a small number of large sites, our training clusters, and there is benefit to keeping the network distance between them small. So wherever they're located in the world, they're probably going to be, let's say, on the same continent, or they might even be in the same portion of a continent in a particular year. So now you might be left with not enough serving capacity on the other continents. So then, okay, actually we have to go build the specialized inference elsewhere across the world. I think your intuition is spot-on, but not wholly. In other words, yes, the training clusters probably are going to be used for serving in some number of years. But then we're also having to build the serving clusters as well.
Sonya Huang: And your serving clusters, are they different than the training clusters? Are they smaller? Are they cheaper per megawatt?
Amin Vahdat: Not necessarily cheaper, actually, because for serving there is more need to co-locate compute, networking and storage. And so then we get to the inability to specialize. For training, you actually have this big, dense, uniform deployment. Whereas for serving, you're going to have to have the mix of storage, compute and accelerators. Furthermore—and it's a very interesting aspect of this that we could go into more detail on as well—for serving, you actually don't want to have too much in one place. You want to be serving your workloads from all over the planet. But now we wind up having individual model endpoints, and that actually can be variations of models. So now we have to spread these models out across the world, also accounting for locality. So they will be smaller. They'll actually be less vertically integrated, probably, on the inference side.
Sonya Huang: Super interesting. So if you made a data center five years ago—the state-of-the-art accelerator five years ago was very different, vastly less efficient than the ones that are being made today—do you actually go back and swap out the chips in those old data centers? And this relates to an ongoing debate, I think, of what is the actual useful life of a chip?
Amin Vahdat: Yeah. So I've been on record as saying this—I was surprised by how much pickup this had—but I said that our seven- and eight-year-old TPUs are still at 100% utilization.
Sonya Huang: Wow.
Amin Vahdat: And—
Sonya Huang: I think I saw this. I didn't realize this.
Amin Vahdat: Yeah. I didn't mean for it to be a meaningful statement, but it was a meaningful statement, apparently. So our older TPUs—GPUs too, but our older TPUs—are seeing significant utilization. We do replace them in the end. There's a question of, okay, the depreciation lifetime is approximately six years. So once they're fully depreciated, and given the power efficiency of newer generations, it does make sense to replace them and upgrade them.
We do replace the chips, but it's really the systems. In other words, we think in terms of pods. A TPU pod might be 9,600 chips, and it might be 140-something, 152 racks. So then we would say, okay, we're going to pull out that pod and we're going to replace it with a, whatever it is, TPU 12 or 13 or 14 pod. It might not be a perfect fit. The new pod's footprint might not be a perfect fit for the old pod's vacancy. We then just have to account for that and figure out how we would retrofit. We can't plan that, because we don't know what that many generations of TPUs are going to be. So then it's again a bit of an art to figure out how we would—and it's actually hard work—figure out how we would decom and then replace, as quickly as possible, with new TPUs.
Sonya Huang: You've written publicly about open standards. Can you say a word on that?
Amin Vahdat: Yeah. So we talked about this interoperability question. For us, while we support vertical integration, and we make it possible to extract just as much performance as you want, it's really important to not force lock-in, and not force a sort of walled garden in terms of how the system works.
I'll give one example. At Google, we've developed a model development framework called JAX. We like it a lot. We think it's really, really good, and we use it extensively internally. Many of our customers like PyTorch. So one thing we could say is, "Hey, if you want to run on TPUs, you've got to use JAX. It's the best." I'm being facetious a little bit—but maybe it is, maybe it isn't. We think it's the best, and that's your only choice because it's so great. Or we could say, look, if you like JAX, we like JAX. If you like JAX, you can use it. But if you like PyTorch, we have TorchTPU, where your unmodified models can run.
And we've seen this throughout history with many, many examples. I've used the example of IP. Why did the Internet Protocol win? There were actually many competing protocols to IP. This is ancient history, in the '70s and early '80s. Why did IP win? It was because it was an open standard, and it was this narrow waist of the hourglass, where anything could run on top in terms of software, and anything could run underneath it in terms of hardware. Open, standard, interoperable. Anyone who brought IP could plug into a router port and become part of the internet. It was beautiful. And that's what allowed the internet to explode in its growth and reach across the world.
So we really do believe in those open standards in terms of plug-in points. Now, if you want to plug in something highly specialized, you can. If you think that you've got a better thing to plug into our framework, you absolutely can. But we want to really support open standards, ideally open source around it as well.
Sonya Huang: It's too big and too important a buildout to be any one vendor's closed, proprietary stack.
Amin Vahdat: We really believe that. And it's got to be one of choice. That's also why, for example, we fully support and have CPUs, GPUs, other accelerators, et cetera.
Sonya Huang: How has day-to-day life for your team changed with AI? Where is it changing your function the most?
Amin Vahdat: So I think the easiest answer is on the software engineering side, where it's been documented externally quite a bit. We at Google, and my team, are using it to great effect in terms of our software development capabilities, but also, frankly, test rollouts, even helping with design. But on the hardware side, which has maybe gotten less external coverage, there's also been significant change. My hardware engineers—I was just looking at the numbers earlier today, in fact—are using AI. Using whatever the token counts, if you want—not the best metric, but it's still a metric—our hardware engineers are using AI as much as the software engineers are. And productivity has gone up. Time from design kickoff to tape-out is shrinking. Time for bring-up is shrinking. So productivity is going up significantly.
But then other, maybe less expected, changes: how we do data center design has changed significantly. How we evaluate—you mentioned, hey, do we do a gigawatt building or a 100-megawatt campus or 200-megawatt campus? In the past, this would be a very detailed, very spreadsheet-driven, human-driven process. And it still is to some extent, but there's a lot of AI now involved that really streamlines the process in terms of planning and development as well.
Sonya Huang: Interesting. So, like, reasoning models, or—
Amin Vahdat: I wouldn't say quite yet reasoning models. It's not replacing human judgment, but it's making it much easier to bring the necessary information together in one place and basically put the right information in front of the humans making the decisions.
Sonya Huang: Okay, I'm going to bring us home with two kind of out-there, fun questions. Question number one: orbital data centers. I've seen that Google is thinking about this quite seriously, and you don't need to comment on that. But does the fact that people are seriously running the numbers on orbital compute mean that there are serious binding constraints on Earth? And what do you think of orbital data centers?
Amin Vahdat: Yeah, it's an exciting direction. We are actually pursuing it, and we've referred to it, with no tongue in cheek, as a moonshot—as one of our big efforts that we're excited about investing in. It goes back to the earlier part of the conversation, where you raised, I think correctly, that in terms of fundamental constraints, energy and energy production is a key, key challenge.
So the bottom line is, from a fundamentals perspective, in space you have something like 40% more power available, because of lack of attenuation in the atmosphere. In other words, it's just more solar capacity: 1.4x. So that's one part of it. But in a sun-synchronous orbit, you have 98 to 100% coverage of sunlight on your solar cells, relative to 28, 30, maybe 35% on land. So just this enormous amount of power. You take the 1.4x and you take the 3 to 4x in terms of number of hours per day, you remove batteries more or less from the equation, and now you have the potential for something that can really—and of course, carbon-free—lots of benefits.
Lots of challenges. Lots and lots of challenges. So, okay, now what about cooling? Naively, you might think cooling in space is easier. It's actually harder. Reliability—we talked about how these things fail sometimes. So repairs become harder. Not impossible, but harder in space.
Sonya Huang: You’ll send a little robot up with the cluster?
Amin Vahdat: Exactly. So perhaps there'll be—this model is a promising direction. Redundancy will probably be your friend. Your earlier question on free-space optics is now going to become reality. We're probably not going to be stringing fiber between these components. So it will actually be free space, and we're going to have the lasers pointing at the receivers and calibrating in real time. There are no showstoppers here, no fundamental showstoppers.
Sonya Huang: Okay. Last question. I've had this image of the Ineffable supercomputer that you built them in my head this whole time. So here's the question: 10 years from now, what will the most frontier supercomputer look like?
Amin Vahdat: Oh my gosh. Yeah. So 10 years is right at that edge where—I would say, in all sincerity, probably some people are thinking about what the 2036 computer looks like. We might have a few thinking that far ahead, but wow, the cone of uncertainty there is way, way too wide.
If you look at the trends, the level of integration is going to be tremendous. We saw the beautiful fiber in the VR200 racks that we put together for Ineffable. My guess is it's going to be more integrated, less fiber. I won't say no fiber in 2036, but I think that from a rack perspective, the rack will look much more tightly integrated, and the amount of fiber—there'll be maybe a small fiber bundle coming out of the rack. One view you might have, from a modular manufacturing perspective, is these racks in 2036 are likely going to be manufactured centrally. So whether they have 72 or 144 or 288 or probably more—576, 1,152, pick your favorite multiple—of GPUs, since we're talking about the Ineffable case, but it'd be the same for the TPUs: deeply integrated into a rack. What would a rack be? Multiple megawatts. We don't know exactly how, but in 2036, we get to imagine things could be multiple megawatts in a single rack. And now you bring water and you bring power and you bring fiber, right? You wheel that rack in, you plug those three things in, and you're off to the races.
Sonya Huang: Okay, so it's not going to be a big alien orb off in space, then?
Amin Vahdat: By 2036, I wouldn't say it's not possible for it to be directly launched into space, and then maybe taken by a robotic arm at some space station and plugged into the right module. Yeah, maybe that for 2036. Fun, fun, fun to think about.
Sonya Huang: I really enjoyed this conversation. This is the most intense, biggest CapEx buildout in history. But I also think it's a technology revolution that is a thing of beauty. And I think you have such a grasp over the beauty of the technology and the binding constraints and how to balance it all. Google is in good hands. Thank you for taking the time to share what you're doing.
Amin Vahdat: A ton of fun. We get to live through this, and we get to define it. So thank you very much. This was a great conversation.
Sonya Huang: Thank you.
Amin Vahdat: Thanks.

