Anthropic’s Katelyn Lesse & Angela Jiang: Building an Ecosystem, not a Walled Garden
Katelyn Lesse and Angela Jiang lead the team building Anthropic’s developer platform. Angela frames the platform as a three-layer stack: knowledge, execution, and coordination. She argues the real leverage is what’s at the top: “strategies,” or meta-harnesses that give each token a different job, from advising to executing to reflecting to memory. On the question of open ecosystem vs. walled garden, they say they aren’t precious about owning the stack. Their deeper bet is standards: they hand skills and MCP to the whole industry, build connectors on the MCP spec, and help agents (Claude and non-Claude) work together. The one place they stay closed is model routing: they argue harnesses should be tuned to a model family vs. routing across models. Their frame for the ecosystem is electricity: transformative because everyone could plug in, and no company wired it alone.
Watch Now
Transcript
Chapters
Intro
Angela Jiang: The last layer of abstraction on top of this is probably the coordination layer. You have knowledge, then you have execution, then you have coordination. And at the coordination layer, we’re beginning to think of these things called strategies, where basically it’s almost like a meta harness. The true low-level harness is designed for execution, but the next one is about okay, if tokens aren’t really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, you want to start composing these orchestrated strategies that go together. And they should sit on top of all these things because at the end of the day, you still need to execute and the execution still needs to know what to do. So everything in theory should ladder together. And so I think if you were to look at our roadmap and maybe project forward a little bit where you expect us to go, we’ll move more and more from the knowledge layer to the execution layer and from the execution layer to the coordination layer in terms of the abstractions that you can see us put out.
Main conversation
Sonya Huang: Katelyn and Angela, thank you so much for joining us today. Lauren and I are thrilled to have you here. You are responsible for building Anthropic’s Platform, and so you are responsible for building what I think is one of the most important, if not the most important, developer platform in the world. And we are really excited to interview you today to understand more about what’s ahead. And so maybe just to get started, can you give us the context of what is Anthropic Platform and where do you sit within Anthropic?
Katelyn Lesse: Yeah, so Platform is both our externally facing APIs, our developer platform that people build on top of when they want to build applications and systems that access Claude’s intelligence, as well as internally we run our product infrastructure, and basically we’re the layer that our apps build on top of internally as well.
Lauren Reeder: Awesome.
Sonya Huang: What’s your north star as a team?
Angela Jiang: It’s a good question. We actually—because we have both internal and external, we actually have two north stars, which is probably like, you’d be like, why? There should only be one north star. But no, we …
Sonya Huang: [laughs] With different planetary systems.
Angela Jiang: Yes, exactly. They’re separate solar systems, so it’s fine. But on the internal side, we really want to provide literally as much leverage as possible for our internal teams to be able to ship AGI-pilled products. And we want them to be able to move fast, be able to have a reliable, great platform to be able to build on top of. But I think that key bit about speed is really intentional for us, and we really, really care about that internally.
Externally, we actually have a lot more complicated set of things, but one of the true norths that we have there is to be able to basically give any builder the tools to be able to work with Claude to build whatever they want to build. And so it’s a bit of a broad statement, but as a result, that boils itself down into being wherever that business is. We really care about bringing our platform really, really close to that business. This is why we spend a lot of time with the hyperscalers, integrating really closely directly with them, like AWS, Google, and so on.
And it is a lot of primitives that we end up creating. We want people to be able to express what they think their product should be. We want them to be able to almost do custom software in their own way. In this new world with AI, what used to be probably economically impossible was that last mile of custom software, but now in theory should be very, very achievable. And we want to give them all the tools and all the capabilities to go and do that. And so sometimes that comes in the form of primitives and APIs and higher order abstractions. And sometimes that comes in the form of just standards. So for example, skills and MCP, those are things just like Claude needs them to be useful. And we can just give them out to the rest of the ecosystem, work with everyone to help you create those things and get the best out of Claude.
So I would say externally, we really are oriented around just helping you be able to build. But internally, that orientation, while still existing, is probably more specified towards speed and being able to move really quickly.
Lauren Reeder: How do you decide what goes into the platform, and what gets externalized and what doesn’t, to decide what products should be available?
Angela Jiang: Yeah. I mean, we generally try to have a philosophy that we try to be consistent across the board. It’s actually one of the reasons why we do internal and external. There’s plenty of other platform businesses and constructs where you actually bifurcate these two things.
For us, we kind of try to intentionally keep it equal. And then as a result, we try to hold this philosophy as much as we can around, for any builder, internal or external, even though for internal builders might have some slightly different requirements in the same way any user would have slightly different requirements, we want to have the same primitives that are available to everyone.
And maybe the overarching thesis for that is that we’ve just seen the capabilities of these models grow in such exponential fashion. And it’s really hard to figure out a long-lasting form factor. I think two years ago, we were all like, “Everything’s chat.” And now everyone’s like, “Forget chat, it’s just like, agents.” And there’s going to be another form factor, another form factor. And we kind of imagine that constantly evolving. And so the best way for us to kind of enable that for everyone, and also ourselves, is to actually build a really robust platform that gives people those kinds of tools to figure out what those form factors are.
And I don’t think we, by any means, feel like we’re the only ones capable of figuring out that form factor, not at all. In fact, the more democratization we can do on that and help people and allow people to experiment, I think the more those form factors will actually kind of naturally come out of the market.
Katelyn Lesse: Yeah, and I think within our team, we’ve had moments where we’re experimenting even with just packaging up our primitives in a different sort of higher order way. And we’ve thought about okay, cool, we’ve solved this exact type of problem with this product that we’ve built into the world. And so we can go and dogfood it for ourselves, but we’d never want to fall into this trap of, like, we’re over-indexed on the problem as it needs to be solved for an internal user, because exactly what Angela said: internal users have very specific requirements, external users have very specific requirements. And so if you over-index on one or the other, you fall into a trap. So a lot of the time what we’ll do is dogfood something internally at the same time that we open up early access of some sort with external customers so that we can kind of get a range of feedback and bring those things back into the platform.
Sonya Huang: I’d love to talk about the higher levels of abstraction that you discussed. So I guess at the base level, this is just raw access to Claude, Opus, whatever tokens. How do you think about, I guess, the layer cake of abstractions above that?
Katelyn Lesse: Yeah, if you look back—so when I joined Anthropic around a year ago, the platform was basically just the Messages API. You know, we had come out with standards like MCP. We obviously have developer tooling around our SDKs and our docs and our console and things like this. But for the most part, it was a stateless API.
And what’s interesting, to Angela’s point on form factors evolving over time, is we found a lot of our customers solving the same problems over and over again that we also were solving over and over again around, as the models got better at running for longer and working with more context at a given time, you want to build agents that can succeed in a long-running context and even a remote context that doesn’t necessarily have a human in the loop.
And so we found that we could piece together our primitives and stand up all the same infrastructure that we’re finding ourselves standing up internally to power our own products and arrive at some higher-order abstractions that let you do more agentic work out of the box. And the problems that we’re solving for you are infrastructure being a hard thing to deal with. Like, how do you figure out spawning sandboxes that are going to have the right governance, and security and spin them up and spin them down when you need to? Or the storage around transcript sessions so that you can resume a session if you stop it and pick it back up later. So that infrastructure is a big thing that we wanted to be able to provide more of out of the box, and we do more of that today.
And then the second thing just being harnesses and harness engineering. There’s a lot of thought and energy going into how do I do my prompt caching and how do I manage my context window, as well as how do I actually just get more intelligence out of the model, and how do I manage my costs and things like that. So kind of we’ve packaged up our primitives a bit more in tune with the problems that we found ourselves solving to provide more of these things out of the box for people so that they can—if they’re building systems for themselves internally, if they’re building products, they can just be more focused on the problems that they want to be solving. And if they want to offload some aspects of those problems to us, they can. And that’s kind of the ethos.
Sonya Huang: And are your customers generally choosing to opt from the grab bag of stuff that you offer, or are they opting into the managed agents offering, like, just take care of it all for me?
Angela Jiang: It varies by the user group. So for, I would say, really AI-native startups, like the ones who are tinkering and experimenting at a really low layer, they’re just going to go for the primitives. And then for everyone else, these are classic, more enterprises or areas where it’s like the purpose of the startup or the philosophy behind the startup isn’t necessarily to optimize on some kind of hill-climbing piece. It’s more like stringing together a bunch of workflows and providing unique user value to that user. For those people, it’s just not their core competency. It’s not where they want to focus their time and resources, and they reach much more for these higher-order package offerings.
Lauren Reeder: What are some examples of the primitives you’ve released at different layers in the last few months? We’ve seen a few of them. We’d love to hear.
Angela Jiang: Yeah, I think maybe one framing I would give for some of the constructs that Katelyn was talking about is—and this is a bit of an oversimplification, but effectively there’s approximately, like, three layers of this cake. At the very bottom is just knowledge, and so at this layer, in many ways, it’s knowledge about the model, it’s knowledge about the things that the model needs, and it’s the ability to know how to actually do something with Claude, is maybe the way I’d phrase that. And so there, the primitives that we have spent more and more time on have been actually things of the past because we still evolve them, but they tend to be a little bit more baked.
Like for example, there’s very specific shapes and parameters we put on the Messages API, and it’s more like trying to expressly showcase Claude’s design—Claude the model’s actual design, the way it thinks, the way it respects certain parameters, the way it will do tool calls, all of those different pieces. And then we started standardizing tools, and then we started standardizing bits and pieces of context that you could put in at different moments in time, which is concretely skills and memory. And so those are the knowledge layer type of abstractions that we’ve put out over the past, I guess, year plus a bit.
The next layer of abstraction that we’ve actually started to spend more and more of our time on is once you know stuff, you then need to execute. And so at the execution layer, that level of abstraction is the part that Katelyn was talking about, around how we’re doing these higher order pieces, but what are we putting higher order there? It really is because you’re now getting Claude to execute work. It’s not just to know something, right? I can give it a question, it can give me an answer. You can string a lot of that stuff together. But now if you need to execute, do work, give me the output, edit files in a bunch of different systems, that becomes a lot more complicated and requires infrastructure to handle. And so that layer is basically, I would say, a low-level harness plus managed infrastructure as the set of abstractions.
Today, our high-level product for that is called Claude Managed Agents. And so that’s a piece, but we started to wrap more and more pieces in that. I think there’s going to be a layer on top of that. We have some inklings of it that we started to build towards, but the last layer of abstraction on top of this is probably the coordination layer. So you have knowledge and you have execution, then you have coordination. And at the coordination layer, we’ve started to expose some of these in ways that aren’t very obvious, but we’re beginning to think of these things called strategies, where basically it’s almost like a meta harness. The harness, the true low-level harness is designed for execution, but the next one is about okay, if tokens aren’t really fungible and you need to give them different jobs, like maybe this token is advising versus this token is executing, this token is dreaming versus this token’s executing, and so on and so forth, you want to start composing these orchestrated strategies that go together, and they should sit on top of all these things, because at the end of the day, you still need to execute and the execution still needs to know what to do. So everything in theory should kind of ladder together.
And so I think if you were to look at our roadmap and maybe project forward a little bit where you expect us to go, we’ll move more and more from the knowledge layer to the execution layer and from the execution layer to the coordination layer in terms of the abstractions that you can see us put out.
Sonya Huang: That’s a really cool framing.
Lauren Reeder: Very cool. How do you think this all comes together into a broader ecosystem beyond just the things that you guys are building? How do you help support people building products on top of it, and how do you help them get the most out of all these pieces?
Angela Jiang: Yeah, I think this is super top of mind for us. We really want to find a way to support as many people in doing this as we can. I think we’re still learning. A lot of the industry has evolved. We’ve seen a lot of different pieces get spun up and spun down. And I think the operative part for Katelyn and I has been in the category of making sure, at least at the base layer, that we provide as many primitives across the board as possible. So this knowledge execution coordination layer, we want to give all of that out to everyone so that people can start to compose and create on top of that. And that’s just from a pure builder point of view, I think.
Then there’s a point of view around how do you plug in with us, right? We’re also building first-party products of our own. We’ve also created some ways to embed natively with us, like for example, connectors, which are built on top of the MCP spec. And we try to be more open about those types of things. And we’re starting to figure out what are the right bits and pieces, but what we’re really trying to do is get to a place where a company is able to get created and built on, they can build whatever products that they want, they can build agents if they need to. And then those agents and those products could be things that could plug into other agents. Some of those agents could be Claude agents, some of those agents could be other people’s agents. But we want to be able to enable that transactability across the board.
And then I think in order for all of that to kind of ultimately be true, there is a bit around standard setting. And I think there’s the traditional standard setting, which is around how do systems interoperate? And that’s things that you’ve seen us do with skills and MCP. But they’re at, again, the builder layer. I think at a higher order layer, there’s also a bit around interoperability and standard setting around how we all treat safety together.
And we’ve talked to a lot of these companies, and this is less from—you know, like, philosophies aside, just more like no one really wants to have technology that’s, for example, doing negative things on their service, right? So cyber, I think, is a great example of this. You want to protect your own systems from negative actors or bad actors.
And so these kinds of standard settings are about how we can find ways to partner with more and more people to be like, yeah, we all kind of want to make sure our critical infrastructure is good. We all want to prevent fraud or any of those things from happening. And how can we work better with each of these members? I think on the last layer, we’re still evolving. And I think we’re still very much trying to find ways that we can be better and work with the rest of the industry to bring people along and work with them. But those are the higher-order primitives or pieces that we wish to have in place so they can work with folks to ultimately solve this.
I think if I were to take a step back at the end of the day on all of these things, this technology is so transformative. And it’s a little bit like electricity in a sense. Like, before electricity, you had to have a candle, and you could only do so many things. But with electricity, the reason why it’s such a transformative technology for all of us and so great of a utility is because you can actually wire it into everything. Everyone is able to actually access it. We also have standards and ways to plug in and do all the pieces that we need. And that’s not something that anybody can do by themselves. They always have to work with the ecosystem and work with partners to figure out a path forward.
Sonya Huang: How do you think about the philosophy of building an open ecosystem versus a walled garden? And how do you think about what products are really important for you to own first party versus where you’re perfectly happy to plug into other components of the ecosystem?
Katelyn Lesse: Yeah, there’s—so maybe in using Angela’s layered cake that we talked about a little bit earlier, you’ll see that on some pieces of this, like execution, for example, what we’ve done within something like Claude Managed Agents—and I think over time you’ll see us try to make this a little bit more modular—we actually aren’t precious about whether you run these things on our infrastructure. Like, it should be sandboxes that we control or it should be a storage layer that we control. Well, we actually, for example, launched self-hosted sandboxes and we partnered with Modal, Vercel, Cloudflare, and a bunch of other folks—even Amazon’s new micro VMs—to have a first-class offering where you can go plug any of those things in. We launched MCP tunnels so that you can call out to your MCP servers that are behind your firewall, and be able to punch through there.
And so for some of these things, whether it runs on our infrastructure versus somebody else’s infrastructure is actually not important to us, because the thing that’s important to us is more that the architecture of how you put together these agents in a way that will be powerful, in a way that will be reliable and scalable, we have strong opinions on that, and you can just conform to the interfaces that we put out there and plug those things in. And we think that that generally is a thing that works really well.
Angela Jiang: Yeah, I think on the verticals where we might build products, I think we have two frames here. The first one is we are always trying to figure out a form factor, like an evolving form factor. We, by the way, don’t think form factors are static. It’s like a dynamic thing. So what might be awesome for one year’s worth of AI development will probably not be awesome for the next year’s worth. And we just kind of try to have that mentality. We tell the team just overall around Anthropic, everyone’s always trying to be like, “Is this AGI-pilled enough?” And then we also have this mentality of we build something, it works, it was cool for a year, and maybe it’s not the right next thing. And so throw it away, try again.
And we tell platform users the same thing. I think that’s probably just attached to the technology. But so yeah, one principle is trying to always constantly find this new form factor. So sometimes we’ll launch products in certain areas to try to showcase a new type of form factor. It’s not necessarily because we think it’s the biggest TAM or the most important thing to go after, but sometimes we’re like, okay, this has always been a really difficult thing, and people have always communicated this way or tried some things this way. And can we show that maybe there’s a slightly different way? And because the model capabilities are so advanced now, can we try to express it a bit differently?
Lauren Reeder: What’s an example of that?
Angela Jiang: Yeah, you know, like, Claude Design is a little bit of that way. I think depending on how you squint, you might see it as a way that we’re going into design as one of the verticals. But more often than not, it’s like if you take a look at what we’re trying to do with that product, there’s a couple of decisions that were made in there. The first one is that you can actually try to offload more and more and more to Claude. And so it tries to be opinionated on just talk to it and let it really try to figure out. And yes, you can still edit it and do these kinds of things, but discourage a little of that and more just talk to Claude to go figure it out.
The second thing was it was really trying to express that actually code is a way to solve for things that you wouldn’t normally think would be the way. So a lot of people who have built generative slide decks or designs or whatever will pick the way of, like, they have some kind of design system, you integrate against the design system. It’s almost in the traditional, like classic WYSIWYG style of designing something. And with Claude Design, it was, okay, can we try to just use code purely, have Claude generate that code, and would it do a good job? And we found through some experiments early on that actually, it looks like it can kind of do that. And how can we showcase that to the world?
So that’s an example. We have a lot of other internal projects, and this kind of falls in the category of expressing form factor. We’ll all try it out internally. It’ll be super cool for two weeks, and then we move on to the next thing. We never even ship the thing, frankly. But yeah, we actually do a lot of product experimentation in that area, and that’s our labs team.
And then there’s the second category, which is that we actually do look at TAM. We’re a business. We do look at TAM. We do look at areas that we think there’d be reasonable agentic operations that would happen. In those areas, we do tend to have an orientation towards things that are more token heavy. And by token heavy—or token hungry maybe is the way I would say that—is what we mean is, like, once you spend, call it one turn, you look at the end of that turn and you say, “Am I done, or am I actually so glad that I did that thing I want to do more of that thing?” We like industries where it’s like, the answer to that question, you say, “I want to do more of that thing.”
So coding is obviously the one that we all know. And the great thing about coding is what it’s actually doing is that once you’ve finished a turn, you look at that and you’re like, “That was incredible. I’m unlocked. I’m going to do more. I’m going to build more. I can do more.” And there’s other services where it’s like, actually, when you finish that turn, you completed the job and you just move on. You know what I mean? And so we tend to go into the ones that are a bit more like there’s this kind of iterative flow. You’re going to build more, generate more together.
And then the last angle that we take a look at is that there are going to be certain business functions where they’re the buyer that we like to go to. We want to help them optimize their workflows, help them create better products there. And I think we’ve been pretty transparent with some of the verticalization. Like, we’ve done finance, we’ve done legal, and we’ve tried to narrow into specific areas where we feel like by having the right context and the right tools and putting it together in a good form factor is probably useful for us to be able to do.
Katelyn Lesse: And in each of those areas, we’re trying to do a bit of showing the art of the possible across all the different ways that you would accomplish those outcomes. And so for finance, for example, is a good one, you know, you could be a company that solves problems in finance, and you could build directly on the Messages API and you can just get some tokens and you can build everything else on top. Or you could be someone who builds on Claude Managed Agents. You can get a lot more out of the box. Or you could say, “I’m going to build a plugin—or a connector—that’s going to sit within one of our products and within those form factors.”
What we did recently, we launched Claude for Financial Services, which is like, okay, cool, we’ve got packages of skills and things like this that you could choose to use within our product, within other people’s products. We even launched cookbooks on here’s how you would use Claude Managed Agents to go and do these things. And so I think for us it’s all kind of an experimentation around we provide people all these different pieces and see where they run with it.
And then sometimes we put together products that are just packaging of all of these things. Like Claude Tag, I think, is a really good example. We had been seeing people in the industry go and—like, Shopify did this with River, Square/Block recently did this with BuilderBot. There’s, like, a few of these examples where people said, “I’m going to build like an agentic platform internal to my company and I’m going to try to give it all the right context and I’m going to make it accessible from Slack or from various other platforms that you’d want it to be accessible at.” And I think Claude Tag was very much a packaging of all those same things that anybody could choose to build something similar. But this is how we’re doing it internally. And if you would like to just plug in and go, here’s what that looks like.
Sonya Huang: What do you think people misunderstood about Claude Tag? Because there was all this ruckus about, oh my gosh, it’s just a Slack bot. Tell us what the magic of Tag is.
Angela Jiang: No, I think it’s a great question, and I do think it actually showcases a little bit of where maybe the future could be going. Yeah, I think if you look at products in the past, people are like, oh, you really attach to the form or the UI, almost, right? It looks like this, so it’s super cool. And I think when you look at Tag, yeah, the way you interact with it is that you literally tag it in Slack. And so yeah, that is the interface.
But that’s not really the important part. The important part is all the context engineering and architecture that we put underneath the hood so that Tag just works. It really should just feel like a coworker. If you go to a company and you onboard, the coworker comes into your channel and then you can chat with it. It’s proactive, it figured out what’s useful, and it just gets stuff done for you.
And so if you think about, especially non-technical audiences, this is a huge unlock. You literally create a channel and then you @Claude, or sometimes you don’t even @Claude, and you’re like, “Hey, I want to be able to do this and do that, and I can’t figure out this, and how do I actually submit an expense report again?” And traditionally, if you think about how to solve that workflow, you are going all over the place, and you’re talking to your manager and you’re talking to your spin-up buddy, and it’s really, really complicated. And today, you just go talk to Claude Tag.
And we do a lot of the hard work on the context engineering, the proactivity, a lot of the harness pieces. I think Andrej Karpathy said it really well. It’s like an org-level harness. There’s a lot of complexity baked into that. Like Katelyn mentioned, you can use our APIs to go and construct that. You’d have to do a lot of the experimentation yourself, obviously. But this is an opinionated take from Anthropic on how you can have this really awesome, always-on agent for your entire company.
And the bit that’s futuristic, I guess, is a lot of that complexity is actually like an iceberg. It’s all the stuff underneath it, that’s actually becoming the harder and harder and more useful part that we’re trying to push through. And I think we’ll see more and more that kind of tidbit that’s outside in the water. It’s just like the interface can actually constantly swap.
Today, Slack is a place where a lot of people collaborate, a lot of businesses collaborate, but also a lot of people collaborate in Teams and some people collaborate by a WhatsApp group. Or they text each other, or some people still email each other. And those could be the form factors—you can imagine agents just going there, almost taking up the same form factors as humans have taken up. It was almost a boring take, but I feel like it’s actually the most forward one, because you want the agent and you want AI to basically be like another person that’s helping you. But it’s very intelligent, can figure out all the context, and you can always have it be a really helpful assistant.
Sonya Huang: Totally. You talked about context and then harnesses quite a bit. And so your team has such an opinionated point of view on what it takes to build an exceptional agent. I imagine a lot of that comes down to the context engineering and the harnesses.
Angela Jiang: Totally.
Sonya Huang: Maybe what best practices or advice would you share with people about what you need to get right on the harness and what you need to get right on the context?
Katelyn Lesse: Yeah, I think it’s interesting because we’ve talked about we launched Claude Managed Agents as this, like, very generic but high-performing harness, because we’ve done all the nitty-gritty work that’s actually really boring and not super interesting around how do you deal with prompt caching, how do you deal with context management. You clear old stuff out of the window. Sometimes you call tools programmatically so you don’t pull everything into the context window and you can keep it clean. There’s a lot of those sort of details on the lower-level harness layer. And I think honestly, like, best practices are just stuff like prompt caching, do it. You’ll save a lot of money and token costs. Obviously, try to keep your context window clear, and then putting those things together in a harness that will be performant is sometimes specific to the task that you’re trying to accomplish. And then of course, evals. I’m surprised we got this far into this thing before one of us said the word “evals.” We’re like, you need evals to make sure that what you’re trying to accomplish is performant.
But I think where we’re starting to go—and Angela mentioned this a little bit earlier—is more of a concept of strategies or meta-harnesses. Because I do think that yes, you can, again, make this lower-level harness, it’s going to be performant, and maybe that’s interesting for you to do yourself, or maybe not, and you offload it to us. But this concept that you can take any given token and spend that token on just executing, or you could take that same token and choose to actually reflect on your past agentic sessions and write learnings to memory so that the next agent does a good job. Or you could take that token and advise with a bigger model so that a smaller model can execute and do a better job. Or you can say, “Execute, execute,” and then, like, a grader comes in and is like, “Did you do a good job? No, you didn’t. Try again.” Right? And so I think the interesting innovation is going to come more at that higher level, at the meta level. And I think optimizing within those strategies is something that our team is really excited about, and we’re starting to do a lot of work there. And I think a lot of other people are starting to feel really excited about this concept of strategies and the jobs you give to tokens. Because again, yes, there’s best practices on stuff like your prompt caching and exactly how you clear stuff out of your context window and how you write your evals and a lot of things like this. But I don’t know that there’s necessarily so much juice to squeeze, in a lot of cases, out of that layer as compared to a layer higher than that.
Angela Jiang: Yeah. And one of the reasons for that, I think, is it has to do with the generations of the models. If you look two years ago, a lot of the harness was like a scaffold to kind of tell the model to go from point A to point B. And you really had to build in a lot. You practically built one wall here and one wall here so that the thing would go in a straight line. And now the models are actually very, very steerable. And so a lot of that steering, you could just put in the prompt. You’re like, “Go from point A to point B,” and the model will go from point A to point B.
So if you have harnesses that are designed to do that kind of steering, you can delete that part. That part we actually frequently encourage, where you can delete part of those harnesses. I think various people have said things along those lines. And that’s, I think, what people oftentimes mean when they say the model will kind of consume some of the scaffolding. And in that sense, for sure, if your scaffolding is telling it to go in a direction that it can just intelligently figure out, that I think will increasingly continue to be so. But as a result of this, what the harness needs to start doing is more allow it to run longer. And so that’s where that execution bit tends to be—I think it sounds like a somewhat silly point, but I do think it results in a lot of differences, because you can go in the direction that you tell it to go, you obviously don’t want it to stop at B. You’re going to be like, “Okay, now go from B to C and then go to F and then go to Z and then come back to me on A.” Something funky like that. And in order to be able to do a lot of those things, the kinds of harnesses that you do are less the steering harness, and it’s more like these kind of strategy harnesses that Katelyn’s mentioning, which allows you to operate at a slightly higher level of thinking, which matches, I think, a lot of the intelligence gains that we’re starting to see with the model.
Sonya Huang: Do you think task-specific harnesses make sense, or vertical-specific or task-specific harnesses?
Angela Jiang: I think people have different opinions on this. Our opinion is yes. I don’t think there’s a general harness. I think there are some capabilities that are obviously very general, and they tend to be very useful. Like, coding is a capability that is very useful because you can use it across so many things, and software has just eaten so much of what is capable. So our ability to write software is therefore useful. I think when you think about very, very specific types of domains, they were going to require a couple of pieces of the harness to be customized. One of those, I do think, is how you choose to handle errors between when you do something and you hand something off to the model. So in domains where you require an extreme level of verification, that logic of how you handle that, again, I think it sounds small, but I totally understand why some people feel like they really want to own the harness, because tweaking that last bit will give you a ton of juice. And especially domains like legal and finance, where there’s a lot of consequences to not getting it perfectly correct, it’s really going to matter. And that’s going to be the difference between your product and someone else’s product being the thing that the user ultimately uses.
And then there are other domains for which I would say it’s not going to matter as much, because you’re able to compress it into a general model capability. So the tweaks where we feel like the domain specificity is really going to matter is the specific verification logic between the model and your execution. And then I think it’s going to be about some of these higher-order strategies on how well you’re able to actually allocate your token budget. I think the context bit is actually a little overdone. Like, yes, you’re going to throw in context, but any harness can actually handle a lot of context. And so that’s just more like you have the data, and if you have the data, then obviously you’re uniquely qualified to do something useful.
Katelyn Lesse: Yeah. And I think when people say “harnesses,” they often mean a lot of different things. And I think this is why, in part, there’s so many different opinions on this. You can think of a harness as literally just a loop that’s like okay, cool, like, user-model-user-model-tool, that sort of thing. Then you could think of a harness as also all of the tools that are packaged up with the harness, right? And there’s just a lot of different definitions of these things. And I think the stuff that can be pretty generic and less interesting to own and deal with is what I was saying earlier, like, getting your prompt caching right. Maybe that is not the world’s most interesting thing. Choosing to clear out old tool calls from the context window and things like that are maybe a little bit less interesting. And you go a layer higher into some of the stuff Angela’s talking about, and then you get into okay, yeah, these are things that I might want to own and control.
And so it’s interesting with Claude Managed Agents, the thing that we built today, we call it higher order, but it’s not really that high order in the sense that you can choose to define all of the tools that you want to bring in as custom tools with the harness, and we give you a lot of knobs to control. You can define skills, you can do your system prompts, you can do a whole bunch of different things—MCP servers and things like this. And I think where we want to get to is a point where you can literally just tell an agent, “Here’s the outcome I want and here’s the budget that I want to spend.” Like, ready, set, go. And you don’t think about any of those things underneath. And so I think there’s just a few different layers of this, that for certain things, you might want to sit at a different layer of what you actually go and control. And you can probably get better outcomes within some of those layers by doing a little bit more optimization work.
Lauren Reeder: Very cool. One of the things I’m curious about, and one that I love about infrastructure and platform teams, is that you get to see what the most advanced users in the world are using and learn from them. I’m curious, what are some things that you’re seeing and learning from the people building on your platform?
Angela Jiang: There’s some people that have been doing some really funky ways of handling context. We ourselves explore this a lot. That’s actually one of the reasons why Tag is such a great product is like there’s a lot of really awesome context engineering that’s happening. We’ve seen some teams be really clever about how they do that, and they are able to kind of think through, like, okay, if I have all these contexts in a bunch of different places, how can I proactively go reach out to them? How can I try to generate enough permissions across each of them and then feed that all into an agent?
And it’s interesting that I guess this is kind of a level of innovation that we’re actually very excited by. It doesn’t express itself as a completely different product form factor, but what it actually does express itself as is maximally useful to users. And we’ve been seeing this more and more with actual internal use cases instead of external ones. So companies who are becoming more AI native, basically, they’re the ones we’re seeing increasingly more and more innovation out of. And so we’ve had customers try to do this for—like, they’ve built their own custom SDLC setup in very, very innovative ways. We’ve had ones who do that for their entire back office. And just the nuances of how they stream in context, I think has been actually really interesting in terms of how they’ve been putting together the pieces. So that’s been one category that’s been really, really fascinating.
Another category that’s been really interesting has actually been with companies that are dealing with, like, really old school software. And so there’s a lot of healthcare companies that we engage with and they’re like, the systems I’m working with, they don’t even have APIs. That’s a dream. And so how can they use computer use and things like this to be able to start to automate and create more connectivity with our systems?
And that area of innovation, I think, has been really exciting. It’s been really interesting to see people try all sorts of crazy stuff, from taking a laptop and trying to run a bunch of things on it to auto-generate a bunch of things that then their agents can go and use. And this has actually probably been an area of a lot of innovation coming from a lot of our customers that we want to find ways to support better and see, like, okay, how can we make this easier for you? How can we help you with some standardization? How can we get it so that you can just have a spec and then Claude can then respect it? And so it’s much easier for you to organically connect a lot of these things. But maybe the general theme I’d give you is interestingly, a lot of the innovation that’s most exciting out there right now has been this context and connectivity layer, which has been really fascinating.
Katelyn Lesse: Yeah, a good example is we’re working with a customer who built some agents on Claude Managed Agents. They also have some agents that they built on other models and other platforms, and they’d optimized each of these agents to be good at the things that they want. They want these agents to all be able to work well together. And they were like, wow, galaxy brain, what if I expose an MCP server on top of this agent so that it can then have this other agent call a tool on that agent? And have these things just be more modular and be able to work together. And we were like, yeah, totally. And we sat down with them and worked through it, and it worked perfectly, and it was pretty cool. And so we’re seeing a lot of, again, that connectivity layer that I think is one of the cooler areas where people are innovating.
But outside of that, one thing that has been cool is just seeing the shift in industry trends of where we’re seeing a lot of our usage come from. We talked a lot about coding as a category, like, of course it absolutely exploded, and there’s so much going on there. And we’re starting to see some of these emerging trends. Like, more recently, we’re starting to see manufacturing really pick up as a category where people are building with AI, and one of our PMs getting on a flight to Detroit to go figure out what these customers need and what’s going on. And so I think we’re going to start to see a lot more outside-the-box of what people think about today sort of use cases, which we’re really excited about.
Sonya Huang: It seems like there’s now—we went through a token maxing moment of history, and now there’s the token rationalization moment of history.
Katelyn Lesse: [laughs]
Sonya Huang: What are your thoughts on that? And what should companies be doing, and then how does the platform team think about enabling that?
Angela Jiang: Yeah. I mean, it makes sense. It makes sense from the high—you start to rationalize. I really like that framing. And I think there are a couple of things that are top of mind for us on this front. Again, it makes sense, and as these models get more and more capable, you’re going to hit levels of intelligence maxing that are there, that then you want to do the next dimension. And the next dimension after intelligence will either be cost or it will be speed. And you just go through that across all possible task complexities in the distribution.
And as we see that happen, something that’s really top of mind for us that we try to spend some time with users on is what you don’t want to do is stop AI usage, right? Like, that’s kind of the wrong move. And we do actually see some of our customers do that. So oftentimes the way that AI spend has erupted inside their company has been through some kind of shadow IT. Their employees just want to use it, they find a way, they end up procuring it themselves, and before you know it, half your org has found some way to install Claude Code. And in that world, it is hard to manage, because these things are, again, very token hungry, ultimately.
And so what we try to encourage our customers is okay, you don’t want to stop the innovation. If you are getting returns on top of this, you are shipping faster than ever before, you can run more operationally efficient, then those are gains. And so the area that we actually try to encourage people is, like, if there is a way for you to construct, again, a strategy that allows you to design an architecture that says, given a task, assess its level of complexity—I mean, I’m effectively describing a router, but there are ways to do this that are, I think, a bit better now. And so this task comes in, has a certain level of complexity. For that level of complexity, you can define some rules, but for the most part, if it’s a hard task, you should probably route that to a big, super smart model. And if it’s not a hard task, you can route that to cheaper models. Designing that, I think, has a little bit—there’s a lot of technical complexity in it, but it’s very, very doable. And we actually encourage people to try those kinds of things. I think ultimately …
Sonya Huang: Do you think you’ll offer a RAG?
Angela Jiang: I think within the Claude space, it will make sense. It’s actually one of the strategies we imagine designing, because the way we’re thinking about a lot of these things is it almost feels like every month there was a new era of something. And if we just take a step back and, like, okay, it seems to be really fast. And so what are the different ways that are recomposable so we can redesign very quickly for any new whatever the cool thing is that month? And so this is in that category of things where we feel like we can actually just recompose a lot of our primitives and then design it.
I think the bit that we do feel really strongly about on the model routing front is we are designing our platform for Claude, and we want to make sure that Claude is great at solving all these things. So we’ll restrict to that space rather than, you know, I don’t think we’re that interested in saying you should route to a different model or whatever.
Sonya Huang: Makes sense.
Katelyn Lesse: Yeah. And well, some of that too is I think we have a strong belief that harnesses and the agentic layer should be tuned to the model family you use it with. And so I think there was a period where people were kind of like, “Yeah, cool. I can build a harness and build an agent and then just plug in a different model underneath.” And they were excited about routers from that perspective. And I think we’ve started to see—like, Vercel just did this with Harness Agent, for example—some of these players in the space come up a layer of abstraction and say, actually, plug in the whole harness and the whole agent that’s tied to a model family, which makes a lot of sense. And so what we could provide is a little bit better, smarter, like, how do you mix and match the right models within the model family underneath that thing, if that makes sense.
But yeah, on the general question of token maxing costs and these sorts of things, I think we’re just kind of going through what feels like a normal, natural cycle for companies, and figuring out how to make the best use of this technology and run their businesses really well and really effectively.
And it’s interesting, like, before working at Anthropic, I was at Stripe, and we were in the very reasonable era of, like, we paid a lot of attention to our AWS bill. And so if someone were to have built some background job and didn’t quite configure it correctly, and this thing’s burning through CPU or whatever it is at any given moment, and causing a big increase in spend that’s not actually worth it, we have put in place the guardrails to find that and then go ask that engineer very nicely to please turn off their background job that’s not within the bounds of what they should be spending for the thing they’re trying to accomplish. I think those are the things with AI that people are going to start to figure out.
And I think to Angela’s point, I think that gets dangerous when you’re just like, “Here’s a cap and you’re stuck within your cap, ready, set, go.” But I do think that encouraging innovation, encouraging people to create really excellent outcomes with this stuff, and then coming in from the side and looking and saying, like, okay, well, there are a few different ways we probably could have accomplished that outcome, right? And one is you take Opus and run it all night and do something crazy. And another is maybe to get a little bit smarter with the strategies that you put together in order to create that same outcome within a lower cost. And I think that’s the next layer of thinking that everyone’s going to start to do.
Lauren Reeder: Very cool. Is there anything that you guys are excited about building over the next few months that you can share a hint at what might come next?
Angela Jiang: Yeah. I mean, I know we’ve said this word, like, 20 million times, so I apologize, but we really are trying to build ways for you to compose strategies. And so that is an area that we’re trying to move into that kind of coordination layer of the abstraction. And we want to start at this front because the types of problems that we see people building, they’re at a layer where it’s like, in order to get the most return on this, you have to be a little clever about the nature of the problem you’re solving. So to give you something concrete, when you try to solve for—let’s say you want to build an agent that’s trying to do bug hunting, and you could just send one off to go and do that, and it’s going to give you a certain return, a level of return of possibility. And then people kind of get stuck at that and they’re like, okay, my next options are I can just swap the model for a different, probably bigger, model, or I could let it run longer. And that’s pretty much the only two levers you have to try to make this bug-hunting agent.
From a lot of experimentation, when we do these kinds of things, those two things are still true, but you actually have a third lever, and it tends to actually do a lot more than you think it does, which is that if you were to best-of-N the thing, it would give you a lot more returns. But just saying those words is fine, and there’s plenty of papers people have published on it. But to actually build that thing and put it into production so you can actually test it on users and see the results for yourself, that’s really, really freaking hard. And you end up building all these custom harnesses, and so forth. But we’re seeing this is where the alpha is. And it’s hard. And so in the same very simple philosophy that we talked about at the beginning, if it gets you the return that you want and it’s hard, we’re going to try to make it easy for you so then you can use it to then run the experiments you actually need to run.
Lauren Reeder: It reminds me of when people were talking about agent swarms a year ago. It’s some version of that.
Katelyn Lesse: Yeah. Has it been a whole year?
Lauren Reeder: I know. We’re finally there.
Angela Jiang: Yeah. No, I think that that’s a type of strategy. Exactly. In the same way that you have one big one that separates a bunch, that’s another type of strategy. And I think people have thought about this maybe in the way of human organization. I guess it could be similar, but if you take it to its kind of end state, it’s actually more that the token has a job. And I think it’s this job piece that we’re really indexed on and we see a lot of returns to. And that’s the thing that we want to spend time with users and the rest of the ecosystem on, on, like, how we can just make that easier for folks to then experiment. Like, we can give you five jobs off the top of our head, and that’s probably what we have internally. And if we give this out to the rest of the ecosystem, it’s probably going to be like 100,000, 200,000. Who knows what other combinations people could put together?
Katelyn Lesse: Yeah, we want to be able to keep doing this hill climbing on, like, how do you get the most value, the most intelligence per dollar, and just put that power in people’s hands. But around the edges of that, we have these personas that have things they have to work through in order to be able to really deploy AI, either within their companies or within their products. And that’s like the sort of enterprise-ready security and compliance controls and things like this.
But really, even just making the platform more modular in the right ways, being able to plug in different pieces of the solutions we’re building. Like, I want to use memory for this thing over here, right? Or whatever else it is. And having a truly excellent developer experience around that, because we spend a lot of time with enterprises who are like, “Okay, I have this walled garden. I need to figure out exactly how I can plug these solutions in.” And so we’ve got a part of our team that’s innovating on things like strategies and jobs, and trying to help you maximize intelligence. And they’re like, “That’s really cool, but I can’t actually use that any of that for XYZ reasons.” So I think solving those problems is really, really important to us.
But then the other persona is the weekend developer who’s like, “I want to go and build something useful for myself.” And they’re often doing that on top of our platform and on top of many other pieces of developer platforms in the community. And I think for some of those folks, there’s more that we can do to provide solutions that are maybe more open or more hackable, or whatever it might be, for those folks to just go wild with what we can offer them and have this really excellent developer experience. And so I think there’s a lot of stuff that maybe I would put in the category of table stakes that I’m really excited about, because I think those are the things that then unlock getting people to say, okay, yes, this thing works for me. And now I can plug in on some of the stuff that you guys are doing that’s really innovative and hill-climby to get more intelligence and save costs and things like that.
Sonya Huang: Wonderful. Katelyn and Angela, I feel—I mean, you are building one of the most important developer platforms in the world, and talking to the two of you over time, I just feel really optimistic that that platform is in very thoughtful hands that care about the ecosystem. So thank you for taking the time today to share what you’re up to and we look forward to what’s ahead.
Katelyn Lesse: Thanks for having us.
Lauren Reeder: Thank you guys.