Factory’s Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself
Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Co-founder and CEO Matan Grinberg explains how the company survived its “journey in the desert,” including the decision to hand nearly all of its revenue back to customers when the product wasn’t making developers obsessed. Matan makes the case that model-agnostic harnesses beat model-and-harness co-design, because they don’t overfit to any single one. He argues open-weight models will capture the majority of tokens by staying one generation behind the frontier at a fraction of the cost, and that CIOs will soon justify every incremental token the way they justify headcount. Ultimately, he predicts 90% of coding tokens will run asynchronously—the “dark factory” where software builds itself.
Watch Now
Transcript
Chapters
Intro
Matan Grinberg: Bezos at Amazon, it’s customer obsession. But in our mind, that’s an input metric. You don’t want to measure input metrics. It doesn’t matter if you’re customer obsessed. You could be customer obsessed and they file a restraining order against you because they don’t like what it is that you’re doing. Our job is to build something so good that our customers themselves become obsessed with us. That is our job. The analogy is if you’re a coach of a basketball team, you don’t want to tell your players before they come out there, “Hey guys, make sure to sweat.” It’s like, what? No, score points. We need to score points. And in doing so yeah, you’re probably going to sweat. And I think similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters.
Main conversation
Pat Grady: We’re here in the studio with Matan from Factory. This is our second time with Matan.
Matan Grinberg: Yes, indeed. Thanks for having me.
Pat Grady: You’re in the small and elite group of second-time Training Data attendees, so thank you.
Matan Grinberg: Oh yeah.
Pat Grady: Matan is the co-founder and CEO of Factory, which makes droids, which are autonomous agents for the art of software development.
Matan Grinberg: Yes, indeed.
Pat Grady: And Matan, we’re going to jump right in because I think you guys are a little bit of a dark horse candidate in this world of software development. It is a market that has absolutely taken off. There are folks like Claude Code and Cognition and others who have a lead, but you guys are coming up strong. Talk about the competitive dynamics and what makes Factory special.
Matan Grinberg: It’s been a wild ride. We started Factory three and a half years ago now. So in April of 2023, when the world—and the enterprise in particular—was barely ready for GitHub Copilot, let alone fully autonomous agents.
And so I think the first two years, it was kind of our journey in the desert is how I like to refer to it, because we were focused on fully autonomous agents, but engineers weren’t ready, procurement teams at the enterprise weren’t ready. And so I think retrospectively, we really honed our craft and learned a lot about how to build for developers in the enterprise. But it took a lot of time to actually come around to when they were ready to receive it.
And so we’re kind of now emerging much more, and some of these other players like Anthropic or OpenAI who have a ton of distribution are going in and bringing their incredible tools like Claude Code or Codex. The thing that enterprises are really caring about that we have learned through those two years is they do not want anyone to kind of be their single point of failure. They do not want anyone to kind of control their fate. And so something that really matters is model independence.
Everyone learned from cloud where, back in the cloud days, it was like AWS or Azure being like, “Hey, come on in, sign this three-year contract. It’s going to be so cheap. We’re going to subsidize it. It’ll be great.” And then a couple of years later, when it came time to renewal, they would 10x the contract.
Pat Grady: [laughs] Data gravity, we got you now.
Matan Grinberg: Yeah, we got you. What are you going to do, a two-year migration to go to someone else? No way. Everyone has scars from that now. And so everyone knows, look, Claude Code is fantastic. Codex from OpenAI is fantastic. We cannot put our fate in any one of these model providers’ hands. Also, if you just look at the risk profiles of the model labs versus the cloud providers, what’s the last piece of drama that came out of one of the cloud providers? Versus the model labs, it seems like there’s always some sort of chaos of internal fighting or getting in spats with the government or any other entities.
And so if you’re going to build this very important part of your business, you want to make sure that you’re robust to any of these changes. And that’s something that we’ve learned over those kind of initial two years is, developers really care about things being modular. They want to know that they can customize it to what they want. They want to know that if there’s a new model that comes out that’s faster or cheaper or more performant, they can kind of hot swap it in. And that’s, I think, one of the biggest reasons why a lot of the largest enterprises are taking the momentum that they’ve had from a Codex or a Claude Code and then are carrying that into Factory, because they get that performance from these fantastic models, but they do it without the vendor lock-in that the model labs directly would provide.
Pat Grady: And if I’m the enterprise, I’m going to be like, “Wait a minute, am I now just getting locked into Factory?” What’s the answer to that?
Matan Grinberg: It’s a really good question, because that is something that you might think of, like, okay, wait, so we’re just switching the lock-in point. All of the modularity that we build is such that if at some point you wanted to say, hey, you know what, Factory’s not staying at the frontier anymore, whether it’s the automations that you build or the skills registry that we help you create, the work that we’ve done stays in your codebase, and any of the automations that we’ve created, the artifacts also live in your codebase. In other words, there aren’t really things that we’re saying are tribal knowledge about your org that we’re keeping on our side and not giving to you.
And that’s part of the relationship that we have with customers is that we similarly want to make sure we’re providing the best experience possible. If we help you arbitrage between different models to get cost optimization, we’re giving you that optimization. We’re not taking that away from you. And I think that’s a really important part of the trust that we’re building with these enterprises.
Pat Grady: You and I were talking probably a couple months ago at this point, and I was trying to give you credit for having the right vision for this market two, three years ago. And you responded with something along the lines of, “Thank you, but being two or three years early is the same as being wrong.”
Matan Grinberg: Yes.
Pat Grady: Which I thought was a wonderful response in so many ways. Can you talk about, like, those two years in the desert, how did it feel to have this vision that turned out to be right, that nobody appreciated for a year or two? Can you just talk about that journey and what it has done to the DNA of your company?
Matan Grinberg: Yeah. I mean, in the moment, it’s really, really difficult because I hadn’t had a job before. I dropped out of my PhD to start this company, and over the course of those two years, convinced 20 of the smartest people that I’ve ever met to quit what it was that they were doing and join Factory and join us on this mission. And these are people with families, these are people with kids who are dedicating years of their lives to this problem and going customer after customer. And they weren’t ready for agents. They didn’t get it. Also, the models weren’t as performant, but I think a lot of it was behavioral.
And I mean, even just a fun anecdote of, like, giving developers an NPS survey. If you ever are giving a developer an NPS survey, they do not like whatever it is that you’re giving it to them, because developers, they vote with their feet. They are very clear what they like and what they don’t like. And if you’re like, hmm, I wonder if they like it, they definitely don’t.
Pat Grady: [laughs]
Matan Grinberg: But during that time, I think there was a lot that we were learning. There was a lot that I myself was like—I’d never had a job before. Enterprise sales is not something that comes obvious to a physicist. But at the end of the day, it doesn’t matter. There’s no—you don’t get any bonus points for being early, because who cares? There’s no consolation prize. It’s either you do the thing or you don’t do the thing. And that’s all that matters.
And for the team, it was really tough. There were points where we ended up getting good at enterprise sales, but the product still wasn’t good. And that’s a very tricky position to be in, because we ended up getting to a point where we were, like, just under $2 million in revenue and the product was not good. And there was a point in time where we realized this. Because if you’re really good at sales, you can sign contracts. You can definitely do that. But if you’re doing that and the developers don’t like your product, it’s like a ticking time bomb, because eventually they’re going to churn and it’s going to be really, really bad. We realized this, and we proactively gave all of those customers their money back.
And I remember that was one of the most difficult decisions to make, because not only is there a group of 20 people who are getting ridiculous offers from all the labs, they have these huge financial incentives to go elsewhere. There are all these other companies that are doing well, and they decided to do this. And then we’re going to say, “Oh yeah. Hey, by the way, that little bit of revenue we managed to get, we’re actually going to give it back because we don’t think product is making their developers happy.”
Pat Grady: Why did you make that decision?
Matan Grinberg: We sold them on a good vision and convinced them that this is the right team to work with and that we were going to deliver the solution for them. But we realized that the way that we had sold them on it and the product that we were delivering was not up to snuff in a way that I don’t think it would hold true to one of our operating principles. And one of our operating principles that I really like is create obsessed customers.
Pat Grady: Yeah.
Matan Grinberg: This kind of flips over Bezos’s thing, where Bezos at Amazon, it’s customer obsession. But in our mind, that’s an input metric. And input metrics? You don’t want to measure input metrics. It doesn’t matter if you’re customer obsessed. You could be customer obsessed and they file a restraining order against you because they don’t like what it is that you’re doing. Our job is to build something so good that our customers themselves become obsessed with us. That is our job. The analogy is if you’re a coach of a basketball team, you don’t want to tell your players before they come out there, “Hey guys, make sure to sweat.” It’s like, what? No, score points. We need to score points. And in doing so, yeah, you’re probably going to sweat. And I think similarly to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters.
And I think, coming back to this, the product that we were delivering was not creating obsessed customers. And we wanted to make sure—like, this was a group of the smartest people I’ve ever met. We were getting there. We were getting a lot of intuition. Things were starting to come together internally. We could see internally we were starting to become a lot more agent-native in how we were doing things, and the product was kind of scratching that itch. But we were kind of ahead of our customers, and we wanted to maintain trust with our customers so that when it does hit, we can come back to them and say, “Hey guys, this is the real deal, I promise.” And to build that credibility, we had to say, “Hey, look, even though you were maybe happy to continue, we’re going to give you this back and say, three months from now, I think it’ll be ready. Give us some time and I promise we will knock your socks off.”
Pat Grady: How did your customers react when you had that conversation?
Matan Grinberg: Some of them were like, “Oh great, sounds good,” because I think it wasn’t something that they were obsessed with. Some of them were a little bit confused, but I think generally, especially enterprises, they’re not used to these things. A lot of times, an enterprise’s budget, once it’s gone, it’s gone, and no one really cares.
Pat Grady: Yeah.
Matan Grinberg: And so some of them didn’t even know if they had a mechanism by which to take back the money.
Pat Grady: [laughs]
Matan Grinberg: But it’s a difficult thing to tell investors who believe in you, too. I remember having the conversation with Shaun. I think Shaun, obviously, he’s stayed really close with the company, so he was very much on the same page. But it’s kind of a scary thing to be like, “Hey, by the way, remember all those updates where you’re saying, ‘Hey, look, the revenue is going up?’ It’s about to go down to zero.” It was a scary thing. And I think it was kind of a leap of faith of, like, we see the signal internally early that this is the direction we need to go. We need to kind of pivot the approach on the product.
But I remember that all-hands where we told the whole team, that was like one of the worst months of my life. Not everyone was going to say, like, “What the hell is this? What’s going on?” But it’s kind of the looks on their faces where they kind of go a little bit pale and they’re like, oh boy, is this just the early signs and we’re about to sink completely?
Pat Grady: How’d you keep the team together through that?
Matan Grinberg: I think honestly, the only reason the team stayed together is we were so ruthless about hiring early on where it was like, people that are genuinely really, really obsessed with the mission, which our mission is to bring autonomy to software engineering. And, like, really, really caring about that, making sure everyone was also—like, very clear feedback loops as to the fate is in our hands. It’s not like this is like, oh, something that I go do. It’s like we all have a part to play in making this work. And I think embracing how much it sucked was also, I think, something that was very valuable.
Pat Grady: Just being honest about it.
Matan Grinberg: Being super honest about, like, yeah, this sucks. Look at those competitors, their revenue is going up like crazy. This is not good. We are in a very bad position. We just had to give back all of our revenue. We need to really get our shit together. And in the moment, I think, retrospectively, those are the moments where really the deepest bonds are made. If you talk to people who are athletes or even academics or whatever, whenever you’re in the stressful period, whether it’s cramming before finals or—we have some rowers on our team, and I think that’s an example we always go to.
Pat Grady: That’s a pure pain sport.
Matan Grinberg: It’s pain. It’s literally just—there is one number that quantifies your performance. It’s just, what is your time on your 2K, or your time in—but embracing that is what creates those enduring bonds, such that afterwards we know what it’s like to be at rock bottom. We know what it’s like to lose. I mean, when we first started the company, our valuation was $5 million. A lot of our competitors, a lot of the companies out there these days, they don’t know what it’s like to not be a unicorn. That’s manifestly what they are day one. Whereas we have been there kind of in those dark moments, and not a single person left.
Pat Grady: Yeah.
Matan Grinberg: That makes us so resilient and so strong that going forward, things are going a lot better now, but there are going to be really bad times. But we have that resiliency in our DNA that I’m not sure some of these other companies do.
Sonya Huang: I love that. So talk to us about what changed. And I’m curious about your comment from earlier that the models getting better is not the most important thing that happens, because at least in my mind, the models getting better is the most important thing that happens. So just help me understand.
Matan Grinberg: Yeah, so a couple things. So one is the interaction pattern that we were building for before was too ambitious. To your point, we were right in that what we were building for was fully autonomous agents, but it was two years too early, which makes it wrong. And fully autonomous agents require a complete change in behavior from the developer. And we were trying to do that out of the box before they were even using tools like Copilot. It was just too much of a leap. It was too much of a step-function jump.
So it’s an important day—September 26, 2025 was when we first put out basically the Droid CLI. And the Droid CLI met developers where they were in a manner that previously these fully autonomous agents did not. And also its performance was completely state of the art, and it was model-agnostic, so it could use every model that was out there. September 26 was also two years after we initially started. So the world had gotten much more used to using things like autocomplete. By late 2025, most engineers were using an autocomplete tool, and many were starting to, at the time, use a chat interface to ask an agent to go do changes wholesale, so the more agentic interaction.
However, what we see is that if you go back now and use this agentic interaction, some of these older models are still good. So the biggest thing that changed was developers, and in particular in the enterprise, being open-minded to this new way of working. In particular, developers, they’ve established their workflows over the last 30 years. They can be stubborn. A lot of them were like, “No, no, no. My craft could never be done by an AI tool.” So a lot of it was just understanding how to work with these tools and having the willingness to go in and try, and also the intuition about what are the guardrails that you need to provide in order for it to succeed.
Sonya Huang: Yeah.
Matan Grinberg: And so I think it was a combination of both of these things. The model’s getting better, so you need to do less in the way of providing guardrails, but also developers lowering their guard and being like, okay, you know what? Let me go try and do these things. It’s going to go do things I don’t like.
And then also, there’s a certain degree to which when Andrej Karpathy tweets about something, then every engineer suddenly is like, okay, maybe this is true. And Andrej started to tweet about this agentic work—early on, he wasn’t as open to it. And then him being more open to it genuinely just changed some people’s minds. Which is funny, but that’s some of the things that go into behavior changes. You hear it from people you trust. You start seeing it from people within your organization who are maybe a little bit more agent-native. But these things together is kind of what changed that.
Sonya Huang: And now we’re all going to be on Slack. [laughs]
Matan Grinberg: We might, we might be on Slack. We might be pushing the limits of Slack, which I think is going to be another interesting thing.
Sonya Huang: That’s cool. That’s cool. Okay, so September 2025, you launched the Droid CLI. You said “frontier performance soda.” What does that mean for you?
Matan Grinberg: There’s the benchmarks, which have a very short half-life. Anytime there’s a good benchmark, it gets benchmaxed within, like, three to six months.
Sonya Huang: Yeah.
Matan Grinberg: At the time, I think the one that we kind of championed when we launched, and it ended up becoming a pretty good benchmark, was Terminal-Bench. Prior to that, the one that was kind of leading was SWE-bench, which kind of took some open-source projects and some examples of issues that were then solved. The problem with that was it was very focused on Python and scripting or individual file changes, whereas Terminal-Bench was more—one, it was in the terminal setting, so it was things like scheduling runs and things that were not just changing the code file, but general software development tasks. And that was something that we ended up having really frontier performance on. Now it’s benchmaxed to the extreme, to where I think models that come out now are like 90 percent on it. And I think there’s a very short time horizon from putting out a good benchmark to then it being kind of in the training data.
Sonya Huang: What goes into building a great harness? It seems like there’s almost a lot of FUD in the ecosystem of my harness is better than your harness, and you need to own the model to have a good harness, or actually you have a better harness if you don’t own the model. What’s your mental model for—benchmark maxing aside, what keeps you at the frontier?
Matan Grinberg: Yeah, so some general things that matter are the way you do caching. So cached tokens end up being like a tenth as expensive, and so one big piece of performance for a given harness is what is your rate of token caching?
Another example would be how do you perform while in compression or compaction? So typically, when you’re dealing with a long session, you’re going to exceed the context limit of the model itself. And so the harness will do some sort of summarization, compression, compaction, whatever you want to call it. And the way that you perform during that compaction is a big determining factor of how good your harness is. And tests that they do for that, they call it needle in the haystack, where you have some long thread and maybe there’s one piece of information that’s really important. How often will your harness preserve that through compaction?
Other examples are tool use, or how does it use the environment to validate whatever work that it’s doing? These are things that you can kind of have individual metrics on and that we kind of have our own internal benchmarks to measure how do the out-of-the-box agents do versus how does Factory perform. I think one thing that naively everyone believed initially was if you train the model and you build the harness, you’re going to make them better together.
Sonya Huang: Yeah.
Matan Grinberg: And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better.
Sonya Huang: My intuition would be model harness co-design makes you better.
Matan Grinberg: Yes.
Sonya Huang: What’s the intuition for why it’s actually not?
Matan Grinberg: It’s very analogous to the idea maybe, I don’t know, 10 years ago, back in ML days before GPT-3, “I want to train my personal AI, I’m going to give it all of my data, because I want it to know me.” Turns out the answer was: train it on the whole internet, and it’ll be so much better for you than if it were just trained on your data. So there’s a sort of analog that emerges where it’s what data is to a model, models are to a harness. Where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular. And there are certain intricacies about different models that you can learn from and then improve different models’ performance in your own harness.
And this was why, for example, we kind of stopped doing it, because Terminal-Bench got so benchmaxed. But initially, when every new Opus or GPT model would come out, it would perform better on Terminal-Bench in Droid than it would in Claude Code or Codex. And this is something that I think was somewhat frustrating, because from a lab perspective, you ideally want it so that it’s better together, because then that means you have to use their harness and you can’t use a different one. But I think the reality is having that multimodal harness ends up getting kind of frontier on all of those aspects.
Pat Grady: Is there a good example or illustration of that? Conceptually, it makes sense. Is there an easy way to illustrate it?
Matan Grinberg: Maybe a good example of it is like, if you’re familiar with the different behaviors of Opus and GPT-5.6 right now.
Sonya Huang: I am. He’s not. [laughs]
Pat Grady: [laughs]
Matan Grinberg: I mean, loosely, loosely—I mean, to be fair, honestly, these days I’m not doing it as much either, but I will say loosely, Opus is kind of like that super friendly colleague where you’re like, “Hey, I want to go do these 20 tasks,” and they’re like, “Okay, cool. Hey, by the way, five of those tasks, I realized we didn’t need to do. Don’t worry about it. I got other of these done, did it this way.”
Sonya Huang: They’ll be like, “Tonight’s not a good time. Let’s pick it up in the morning.”
Matan Grinberg: Yeah. And let’s go get a beer afterwards and hang out, whatever. Meanwhile, GPT-5.6 is like, “Absolutely, I will do every single one of those and nothing will stop me. I’m not going to sleep until there’s—” it’s kind of very OCD and meticulous. But sometimes you want one where it actually realizes hey, that list of 20 that you gave me, actually, here’s a better way of doing it anyway. 5.6 is more methodical.
If you build a harness for each of those, there are actually different things that that harness will then be good or bad at. So for example, one thing that typically agents will do is they’ll have a to-do list of, like, if you have a task, it’ll go and generate a to-do list. And the Claude Code harness can, in some cases—and this is maybe less relevant now, but I think earlier, this is just a more illustrative example—earlier it was really strict to make sure it would stick to the to-do list, because the model itself would typically wander. Meanwhile, Codex wouldn’t do that because the model itself was really, really OCD about that.
Sonya Huang: That’s right.
Matan Grinberg: But if you’re a user, you want to have the same experience regardless. You want to make sure if you switch to a different model, you’re not going to suddenly lose track of whatever things that you are working on. And so there are certain things where, like, maybe in some cases you really want robust tool use—and there are tools that you use to do these to-do lists. You want really robust tool use, and you want to make sure that no matter what, if I’m a user, I want to see my to-do list there. There were some cases where it would just not have the to-do list. And so these are things that kind of improve the general performance. And the to-do list matters because if you’re doing some crazy migration and you don’t have the to-do list, and then you’re in this long session where there’s compaction, that might get lost in the summarization. And then now you forgot what your seventh step was, and that could be one of the failure modes. That’s kind of an example of how …
Pat Grady: That’s a good example. Yeah. Yeah.
Sonya Huang: That’s a great example. Okay, so we talked about one type of maxing, benchmark maxing. Let’s talk about token maxing.
Matan Grinberg: Yes.
Sonya Huang: Because it feels like the world has changed a lot. We’ve gone from token maxing to now cost rationalization. What does that mean for Factory?
Matan Grinberg: Yeah, so maybe I’ll lay this out just so we’re all on the same page of, like, the way that we see what’s led us to this token maxing. So loosely, there was this phase one where—maybe phase zero was no one believed in AI. Then phase one, everyone believes in AI. And then boards were like, “Mr. CEO, what are you doing about AI? What’s your AI strategy?” And Mr. CEO was like, “Shit, I don’t know. What’s our AI strategy? CTO, make sure everyone goes and uses AI.”
And so then phase two is, CTO is like, “Okay, shit, we gotta make sure everyone uses AI. Let’s start putting it in performance reviews. Let’s make public rankings of who’s using tokens the most, because everyone’s stubborn. No one wants to use this stuff. They’re all skeptical.”
And then we enter phase three, which is everyone sees these ratings, they see that it’s part of their perf reviews, and they’re like, “Okay, I’m going to use AI for everything.” And that’s kind of phase three. It’s this token maxing, where people are using Opus for literally everything. “What’s the weather in SF? Opus, tell me. I don’t know.”
Pat Grady: [laughs]
Matan Grinberg: There are banks that we are working with where they are spending literally hundreds of thousands of dollars a month on people asking things like literally, what is the weather? Or, like, tell me about Python. Trivial questions that you could Google, people are asking Opus. And the reality is this happened because we were so worried about adoption that we overcorrected and we’re like, adoption by any means necessary.
And I think that’s actually a decent approach. It’s probably faster to do that and then curb usage or make usage more responsible than it is to start limited and be like, you can only use it for this thing, because when you have people that are stubborn, first you want to just prove that it works and then you can get kind of more mature about it. Where Factory fits in, I think one of the most important things that we do is that we have the Factory Router, which allows you to dynamically route to different models based on the task that you’re doing. So if you’re asking what the weather is, you probably don’t need the very frontier of human intelligence to answer that for you.
Sonya Huang: Or you really do.
Matan Grinberg: I mean, it depends. I don’t know, it depends on what kind of answer you’re looking for—giving you a full, down-to-the-molecular-level answer of what’s happening. Allowing that, but also more importantly, for every enterprise, something that no one’s dealing with yet, but 12 months from now is going to be the case, is not everyone needs the same tokens. Having a blanket kind of token cap for every individual in some large bank, let’s say, makes no sense. So every CIO is going to need to answer for every incremental token, where do we put it? And right now, it is super not obvious how you would do that. Right now we’re saying, oh, the PMs who are vibe-coding dashboards get the same token limits as the engineers who are building critical infrastructure. That’s probably not the best thing to do. Or similarly, you might be dealing with COBOL codebases where Opus is not the best model to use, but instead maybe some fine-tuned model on that codebase in particular.
The point of the router is that we can kind of accommodate these different constraints where maybe you say, you know what, this part of the org, they’re just vibe-coding. They can use Gemini Flash. This part of the org, they’re doing COBOL. We fine-tuned this great model to work on COBOL. Let’s route to that when we’re working on that part of the codebase. Maybe this other part, we really care about reliability, so let’s generate the code with OpenAI, test it with Anthropic, review it with Gemini, things like that.
And we can actually take in your routing procedure instructions in natural language, so you could even say things—like, it’s not purely deterministic, it can even be, “Pat, I don’t know what he’s doing.”
Sonya Huang: Give Pat Gemini Flash.
Pat Grady: Come on!
Matan Grinberg: I think we really need to avoid having them use open models because, whatever the reason, we don’t like the way open models perform here. And we’ll do internal benchmarking to know which models are better at which of these tasks.
Sonya Huang: How close are the open models at this point? Which one’s the best?
Matan Grinberg: GLM 5.2 is incredible. It’s at the point where internally we have no token limits for our engineers, and half of our tokens are open to open models.
Pat Grady: Wow!
Matan Grinberg: Yeah, because they’re just faster and they’re cheaper. They’re just as performant. And I think the thing that everyone gets wrong is everyone is comparing GLM 5.2 to the latest model, like Opus 4.8 or GPT-5.6, but really they should be compared to Opus 4.7 or GPT-5.5.
Sonya Huang: Why?
Matan Grinberg: Because generally the open models come later, and they’re kind of a generation behind. And that’s kind of—like, the frontier models will be frontier. The question is: Are the open models getting as good as, like, frontier minus one? And the answer is unequivocally yes, which I think is a really, really interesting outcome. It’s great for consumers—and by consumers, I don’t mean individuals, I mean the consumers of the APIs, because if you’re a business that is doing software engineering, your job is, at a very high level, to solve problems. And if we can allow you to solve those problems faster and with cheaper models that are just as performant, that means you can solve more problems. That is a good thing. And it is a very good world where there is not a monopoly on intelligence, but instead kind of a garden of intelligence that you can pick and choose when you’d like.
Something that we joke about is, like, on this intelligence allocation thing, if you’re trying to get a tutor for your daughter in algebra, you can probably find someone cheaper than Albert Einstein to be that tutor. Now it might be that she eventually goes and becomes a leading physicist or something, in which case, yeah, maybe let’s get Albert Einstein in there. But most likely, you can get a high school student or something like that. And it’s probably much more cost-effective for you as well to do so.
Pat Grady: Since you guys do the model routing, if there’s a pie chart that shows the complexion of models being used by your customer base today, what did it look like a few months ago? What does it look like today? What do you think it’ll look like in a year?
Matan Grinberg: Yeah, I will caveat this with saying that right now enterprises haven’t gone too opinionated yet into the routing procedures.
Pat Grady: Okay.
Matan Grinberg: This is something that will happen over the next six to twelve months. But right now they’re just going from no router to router. That’s kind of the first change. Then it’s going to be the exact nature of the routing. At the beginning of the year, it was less than one percent of tokens that went to open models. In the first quarter, it became a single-digit percent. It has now crossed into being a double-digit percent of tokens. Now percent of tokens is not always the same as percent of cost, because the open tokens are cheaper, but it is pretty crazy to see the growth there.
Sonya Huang: What’s your forecast?
Matan Grinberg: My sense is that we will asymptote towards the vast majority being open, just because it provides you more optionality and it’s cheaper. But that’s of token share, not necessarily of leverage share, because maybe there’s one percent of tokens that are incredibly, incredibly valuable and are very key decision-making. And then the rest are more like implementation tokens, or kind of lower stakes, if you will. I don’t think there’s going to be a world in which it’s ever going to be 100 percent.
Sonya Huang: Yeah.
Matan Grinberg: I think the frontier of intelligence will inherently always be valuable for every business, just because the stakes are going to get higher and the kind of intricacy with which you think is going to be more important, but we’ll be better at offloading certain tasks. And you can loosely think of this already with the way orgs are structured, where in general, engineering leaders are more tenured engineers who in theory have more wisdom, and each minute of their brainpower is higher leverage, in theory. And you can also imagine, like, consider a human engineer, and try mapping over the course of their day how much brainpower they’re using. It’s probably going to be really low for a lot of it, but then there are going to be some moments where they’re going pretty high. They’re deeply concentrating and thinking about some systems design problem or whatever. All of those low-leverage moments, we want to automate away. And those very high-leverage moments—sometimes we’re referring to them as the eureka moments, or the moments where they’re doing something that’s very high leverage, what if those aren’t just moments, but what if those are hours at a time? Because you don’t have to deal with all the other stuff.
And I think that’s kind of the way to think about intelligence allocation is if you’re an engineer and you’re writing docs, that is such a low-leverage use of your time. You’ve become an expert in your craft, and you spend hours writing docs. I remember it was actually valuable. I remember Stripe had so much alpha for just having incredible docs. But imagine all the other stuff those incredible engineers could do if it wasn’t writing documentation. We should live in a world where everyone can have docs as good as Stripe, and that is strictly beneficial for everyone. And then the question is okay, what do those really smart engineers do with their time once they don’t have to do that?
Sonya Huang: Maybe this is a good time to talk about business model, given that, especially with the rise of open-weight models, the cost differential—I imagine that means very different things for your cost structure, but very similar value delivered to customers. How do you think about business model and pricing?
Matan Grinberg: Yeah, this is more what our customers want and need, as opposed to what we want and need. So for example, I think right now usage-based is clearly the way to go. We want to be aligned with what they are doing and what we are doing. I think seat-based doesn’t make sense, at least for what we are doing. My sense is that eventually we will change to outcome-based. Now I don’t think the enterprise is ready for that—and we’ve learned our lesson from those first two years, we are not going to impose things, right? But my suspicion is that in the 2030s, things will probably look more outcome-based.
Sonya Huang: Yeah.
Pat Grady: What does outcome-based mean for your market? What would be the definition of an outcome?
Matan Grinberg: So maybe here’s a way to put it. Right now we are usage-based: the more tokens you use, the more you pay, the more we get. Now since we are model-independent, with our router, we are kind of pointing a token cannon at either OpenAI, Anthropic, AWS, GCP, you know, any one of these people. To a certain degree, this is like a really dumbed-down version of a marketplace, where right now the buy side is an engineer who wants a task done. And then you have the model providers who are saying, either in benchmarks right now, they’re like, we perform at this cost and this performance, and then we determine who we go to for that given task. There’s a world in which, if it’s so important to get these tokens, they might kind of bid in a certain way, saying, look, here is our cost for this task. We will get this task done at this cost, no matter what. But they’re pricing it such that they hope that they can make a margin there.
Pat Grady: Yeah.
Matan Grinberg: If they price it wrong, they’re at a negative margin. If they price it right and win the bid, then they get the positive margin. And the way you determine if the task was successful is by some validation loops, because no one is using these tools anymore where it’s just like, “Write me code. Great, thank you.” It’s generally, “Write me code, and here’s how I know it was done well.”
And similarly, if you are a model lab and you are given, “Here’s a task, here’s the validation criteria,” you’ll be able to say roughly how much you think you would be willing to pay to get those tokens. And you want to have some margin on that. And then in that world, that’s basically—that’s a way that you kind of dynamically shift from usage-based to outcome-based. I think that there are so many questions with this—and this is very much forward-looking, but I think there’s a lot of questions about how do you subdivide tasks. Divvying that up, I think, is something that’s not obvious.
Pat Grady: Yeah.
Matan Grinberg: But as these tools get better, doing things like that actually become way easier.
Pat Grady: Yeah, that’s fascinating. Yeah, if you can scope a task and then create a competitive marketplace, that’d be a fascinating version of the future.
Matan Grinberg: Yes. And as a user, it then creates an incentive to be very thorough in your validation criteria.
Pat Grady: Yeah.
Matan Grinberg: Because there are stories where you ask an agent to fix your code and it deletes your code.
Pat Grady: Yeah.
Matan Grinberg: It’s like the solution is just get rid of it all.
Sonya Huang: It’s like that Silicon Valley episode. You can see how prescient it was.
Matan Grinberg: Yeah, so you need to make sure your tests are very thorough, because technically it could hit all of your validation criteria.
Sonya Huang: Or the son of Anton will go rogue.
Matan Grinberg: Yeah, exactly.
Pat Grady: Yeah.
Matan Grinberg: Yeah. Good reference.
Sonya Huang: Maybe zooming out a little bit. You named the company Factory—actually, you named it Droid before Factory.
Matan Grinberg: That’s right.
Sonya Huang: But you named it Factory before this concept took off, and now it feels like everybody wants to build a software factory. Where do you think we are today in terms of the building of software factories, and how close are we to the ultimate vision of a software factory?
Matan Grinberg: Yeah, everyone has a software factory, whether they know it or not. It’s just a very inefficient one. It feels like pre-industrialization, where people were manually sewing things together, or woodworking, or whatever it might be. And these things are very inefficient. Right now, if you go to an organization that has more than 10,000 people and you were to ask about the process by which they decide and release a feature, there is hundreds or maybe thousands of people in that process, and most likely they couldn’t even draw it for you. There’s a very low likelihood that they would know what that process looks like. That is not because they think that is the right way of doing things. That is just kind of the nature of building large software as it is today.
But with these systems, so much tribal knowledge can be codified. So much of this stuff that typically would require, oh, we need to ask this guru who’s been here for 30 years who has the wisdom. Oh, we then need this approval and that approval. Oh, and I forgot there was some doc that said we always have to do this checklist. And it relies so much on human behavior and redundancy. So much of that can be automated and refocused on what actually moves the needle for our business. And I think this move towards software factories is a move towards figuring out what are the actual inputs that determine what features we need to build. And that might be inputs from the customers, inputs from the market, inputs from product leaders at the company.
And let’s be very clear, these are the signals, the inputs that we are taking in here. Okay, great, we have those signals. Then what is the process by which we build this? And really mapping out the assembly lines of how you are building software is really important, because then you get to close the loop and say, did this actually deliver an outcome for our business?
Talking before about the tokenomics, if you’re that CIO and you’re faced with that question of where do you put every incremental token, really, the question two years from now is going to become, where do you put every incremental dollar? And so you’re going to have to be asked: do you put that incremental dollar towards headcount or towards tokens? And if tokens, to where in the org?
And these are things that you can only really know when you have these kind of feedback loops that give you examples of, like, hey, by the way, we made those decisions based on this data, and it did not matter at all. We added these new features and no one cared. It didn’t create more retention, it didn’t create more usage or whatever metrics that business is looking to optimize. And the only way to do this is you need kind of more rigor and more process. It almost feels like 10 years from now we’re going to look back at this previous era of software and it’s going to feel like businesses in ancient times where they didn’t do accounting.
Pat Grady: It’s going to be like marketing in the day of Mad Men, right?
Matan Grinberg: Yes.
Pat Grady: Where it’s all creative, and you have no idea what’s actually working.
Matan Grinberg: It makes no sense. It’s like, “Oh yeah, let’s ship that feature. Oh, I think it went well. Yeah, I got some metrics on that.” It’s like, no, if you guys read the blog post that Jack Dorsey put out about how every company is like an AGI, there’s also this degree to which if your company is an AGI, you want to optimize the weights.
Pat Grady: Yeah.
Matan Grinberg: You want to figure out what nodes are doing what things, which are load-bearing, which are not, which need more tokens, where do you need more nodes. And in order to do that, you don’t train a model by vibes—I mean, okay, actually you kind of do.
Pat Grady: [laughs]
Matan Grinberg: But I guess more importantly, you don’t do backprop in a model by vibes. Like, you are running those actual calculations, and you are seeing when we change this node, what happens. Now you might be making bets on how to change the model by vibes, but it’s pretty mathematical in what you were doing. Meanwhile, at companies, people are determining token budgets just by shooting from the hip. People are laying people off by shooting from the hip and just being like, oh yeah, like, 20,000. There is no way there is science to laying off 20,000 people. That is just like, here’s a chunk, let’s just see what happens.
Instead, I think in these organizations, the way they can do things is much more mathematical of, like, this part of the business matters a lot and does better if we give it more tokens. It doesn’t actually matter if we give it more humans. So let’s give them more tokens. There might be other parts of the business where actually giving them more tokens doesn’t matter, but more people matter, because if we build more relationships with our customers and deeper relationships with our customers, that matters. But these are things that we’re going to need quantitative insight on. And you need a software factory to do that. Otherwise, you’re just, like, shooting from the hip and just guessing, which won’t work as well.
Sonya Huang: In the limit, how much do you think people will spend on tokens versus on engineering headcount?
Matan Grinberg: It’ll depend on the business. I think every business will have a balance. And it just depends on—look, an easy example is generally salespeople, they probably don’t need that many tokens if they’re good salespeople, because generally where they provide the most alpha is when they’re in the seat face-to-face with their customers, talking about the customer’s problems, understanding how they build software in our case, and how we can make that more efficient, more productive. They can use tokens a little bit of, like, generate them an AI debrief, take some notes, help them with the follow-up. But that’s so minimal, the number of tokens, it basically doesn’t matter. Like, if you add more tokens to the sales team, it probably won’t change their output. If you add more humans to the sales team, it probably will.
Meanwhile, engineering teams are pretty different, where engineering teams generally, it seems like, you want people to own an outcome end to end, but then if you give them more tokens, they can produce a lot more. And then there’s a lot of places in between of, like, operations, finance, marketing. These are places that are neither here nor there, where I think they’re somewhere in between, and it kind of depends on your business. But I think every business is going to have to ask, what is our core competency?
Something that we see a lot in the market—or we used to see, and now they finally kind of hit reality, but what we used to see is, oh, we’re going to build our own software development agents. And we’re like, okay, you’re a consumer logistics company, are you sure you want to do that? They’re like, yeah, yeah, yeah, we have to do this. And it’s like, okay. And then six months later, it’s like, wait, actually, this is not a core competency for our business. We don’t want to hire AI engineers to be doing this. Our core competency is consumer logistics. That’s what we want to focus on. And I think this is an opportunity for every business to double down on their core competency and what matters for them, and then procure externally whatever it is that doesn’t matter for them.
Sonya Huang: Yeah.
Matan Grinberg: A trivial example of this is like, I don’t know, in the days of the early internet, you probably had to be a programmer to build a website. And websites generally help if you’re a pizza shop, because people come to your pizza shop, they want to be able to order, whatever. At that time, would you say it was a core competency of a pizza shop to have engineers? Certainly not. That is kind of a byproduct of a brief moment in time. But then there were companies out there that help you build a website, you don’t need to be technical.
And then this is why we live in a world where most pizza shops don’t have an engineering department, which I think is probably a good thing. And I think similarly, a lot of businesses have dealt with the reality of if you want to do XYZ other thing, you have to bring in people of this type of role. But I think that’s been something you had to do, not because it’s a core competency of the business. And allowing businesses to focus and double down on the things that they are best at, I think, is going to be good for the consumers of their business. And so I think we’re just going to see a lot of ruthless refocusing on what actually matters, which is going to be cool to see.
Pat Grady: Well, on that, so every company kind of has to go through this process of reinvention. You know, 10 or 20 years ago, people talked about digital transformation, and I don’t know if anybody’s given it a buzzword now, but AI transformation, something of that sort. A couple years ago, you ran into a bunch of organizations that just weren’t ready to deal with autonomous agents. You’ve seen your customers start to change, and so the question is, when you look at your customers as they kind of go up this maturity curve and sort of reinvent themselves for the future, any good tricks or techniques that you’ve seen them use to repot themselves a bit?
Matan Grinberg: Yeah. I mean, I think surprisingly, the companies that have been doing company-wide hackathons really end up doing well. It seems relatively trivial, but just setting aside a day where everyone in the workforce is just like, build shit with AI, it really sets the tone and sets the pace.
Pat Grady: Sonya’s giving me a look.
Sonya Huang: I tried to force him to build stuff with coding agents. It didn’t go so well.
Matan Grinberg: We’ll work on it. We’ll do it after this one.
Pat Grady: We gave it a great effort.
Matan Grinberg: But that’s it. It’s literally just setting aside the time to do it. And even if it fails miserably, it’s fine. And also, the orgs that are okay with failing.
Pat Grady: Yeah.
Matan Grinberg: It feels like there are some who are like, we need to do it exactly right. We need to make the right decision from day one. No error. You’re going to make mistakes, everyone is going to. And the orgs who are kind of leaning into it and embracing it to a certain degree, I think, are succeeding.
Like, one of our largest customers is EY. EY is not necessarily known to be at the absolute frontier of AI, but I think for them, they were just like, look, this matters. There have been other transformations that we were late to. We’re not going to be late to this. We’re just going to go in. We might mess up, but—obviously, respecting the things that you’re not allowed to mess up.
Pat Grady: Sure.
Matan Grinberg: Put those aside. But, like, let’s go and get our engineers to mess around and build this stuff and see where it breaks and understand what they like and what they don’t like. I think that really matters a lot in the ones that we’re seeing succeed. And also the ones who are pretty bold in reinventing the processes that they’ve put in place and just saying, hey, there are no sacred cows. Let’s put this aside, try something out. If it doesn’t work, we’ll put that sacred cow right back. And I think that’s been kind of a determining factor there. And when it comes from within—if it comes from the board, probably not going to go well.
Pat Grady: Yeah.
Matan Grinberg: If it comes from within, like the tech team or the ICs or the leadership, that’s when we see it go better.
Sonya Huang: Do you have any predictions for the most important changes that are going to happen in your space over the next, call it, 12 months?
Matan Grinberg: A lot of AI consumption is going up like crazy, and everyone’s super, super excited because the revenue’s going wild. A lot of this is synchronous usage. In other words, if everyone woke up sick tomorrow, a lot of Claude Code usage would be zero because it’s all just, “Hey, Claude Code,” or “Hey, Codex,” or “Hey, Droid.” I think in 12 to 24 months, 90 percent of tokens will be asynchronous tokens.
Pat Grady: Hmm.
Matan Grinberg: So these are going to be droids on their own, autonomously being like, “Hey, here’s some signal that I found from a customer. Let’s go fix it. Or let’s go create a first-pass solution to this.” And I think that is going to be where the real agent-native stuff begins, because right now we’re still kind of in copilot mode. If you’re going to an agent and say, “Hey, go do this for me,” it is more agentic because it’s not going to come back and ask you a ton. But it’s still like you are kicking it off.
If you guys have ever been to Tesla’s factories, which is one of the sources of inspiration for the name, it’s just robotic arms everywhere going and doing stuff. It’s not like there are people there going and attaching the widget to the thing. And this idea of a dark factory, where the lights are off and things are just happening, that is where software development is going. That’s kind of where the name came from is like, Elon was always talking about how the factory is the machine that builds the machine.
Pat Grady: Yeah.
Matan Grinberg: And that’s been something that we took to heart. And I guess, also, that combined with his whole thing about how you’re destined to become the opposite of your name. And in our case, Factory becomes artisanal, which is kind of a good flip there.
Pat Grady: What’s your most optimistic version of the future, both for Factory and for the world at large?
Matan Grinberg: So I think short term, there’s going to be a lot of turbulence, because I think a lot of companies have misallocated resources pretty poorly. There’s been a lot of bloat. And I think the correction that’s going to happen there is going to be really painful for a lot of people. And I think that’s something that I think every AI CEO should really bear much more responsibility for than they currently are. And also figuring out ways to address and kind of ameliorate it in some way, because this is something that’s going to be very painful for a lot of people.
Now, I have optimism that we can actually address that faster than we think. We just need to start now in terms of addressing that. Now the longer term, and why I think this is a good thing, is—and why I don’t believe at all the BS that people are saying of oh, engineers are going away. Generally, there is a huge number of problems in the world. A large subset of those problems can be solved with software. A small subset of those problems are currently being solved with software. And so in the short term, this means that okay, first there’s a given problem that was overallocated engineering resources. So okay, we need to reallocate those. Reallocating those is a very kind of cold way of saying some people are going to lose their jobs.
But I think the thing that’s going to happen in the longer term is we need engineers. Engineers are some of the best systems thinkers and the best problem solvers. And there are so many problems that can be solved with software that are not being solved with software. And so that means that we are going to take those engineers and have them go and solve problems that previously were not being solved. That is such a net good for the world, because again, there are so many of these problems that we are not solving. And also, there’s so many problems that we are maybe solving, but with really shitty software. And this is going to enable people to solve it with incredible software.
And the vision for Factory is that we are kind of the factory that allows them to go and build this incredible software to solve these different problems. And these problems range from things that are trivial like, you know, government software typically is not very good, whether it’s DMV or IRS web, all that stuff is generally a pretty poor experience. We don’t need to live like that. We can live in a world where all software is really fantastic.
But also things like pharmaceutical research. So much that goes into solving diseases is not just a biology problem. A lot of it requires the best software engineers in the world. And previously, those problems haven’t allocated the right dollars to attract the best engineers. But now, because of what’s happening, I think we will be much more closely allocated to, like, these are the biggest problems, let’s get the best minds and the best problem solvers to solve that. I think it’s kind of our job as an industry to do that reallocation as quickly as possible, so it’s not ten years, but maybe like six months or a year.
Sonya Huang: Wonderful. Matan, I think the clarity and consistency of your vision over time has just always been very inspiring. And then just seeing how much you’ve grown as a leader and how much Factory has grown as a company, even since the last time we did this Training Data episode, it’s truly awe-inspiring. So thank you for joining us again to share what you’re up to.
Matan Grinberg: I appreciate it a lot. Thank you.
Pat Grady: Thank you.