EPISODE 2026-10-06

AI’s Memory Wall and AI Engineering Meets Wall Street

Thomas Sohmers on Positron AI’s commodity-memory inference chips and using frontier models for chip verification; swyx on AI Engineer New York, finance as the next vertical, and managing coding agents. Nathan and Prakash discuss The Curve, AI safety spending, and a new 3SUM result.

▶ Full show on YouTube𝕏 Live broadcast

Is memory, not compute, the real constraint on AI inference? Thomas Sohmers of Positron AI joins Nathan Labenz and Prakash Narayanan to explain the company’s bet on commodity LPDDR memory instead of HBM, why he thinks looped models save memory capacity rather than bandwidth, and how GPT-6 Astra and Opus 5.5 now run chip verification loops on their own. swyx then previews AI Engineer New York, argues that finance is the next vertical after code, and describes what separates thoughtful AI-assisted work from slop.

The hosts open with Nathan’s takeaways from The Curve, including frontier competition, distillation, and labs spending more compute on safety, and close with a new 3SUM result, the NanoGPT speedrun, and how much AI multiplies their own output. Performance, utilization and spending figures are Sohmers’s and swyx’s own claims, not independent measurements.

The rundown

  1. 1:22Opening10 min
    Notes from The Curve: power, competition, and safety computeNathan reports broad agreement at The Curve that AI will get very powerful, debate over whether self-improvement levels off, how few labs are at the frontier, and labs spending more compute on safety.
    Open segment on YouTube ↗

    Prakash Narayanan opened the show on Tuesday, October 6, and asked Nathan Labenz to share his takeaways from The Curve, the invite-only conference held over the weekend at Lighthaven in Berkeley that Prakash also attended. Labenz described it as a rare event built to bring contrasting views together: frontier-lab insiders and short-timelines believers alongside "AI as Normal Technology" authors and initially skeptical DC policy people. He said the notable shift this year was broad agreement that AIs will get very powerful, with the live debate now whether recursive self-improvement will "foom" or level off.

    Speaking under Chatham House rules, Labenz relayed what he called newsworthy remarks from a senior frontier-lab executive: pretraining and synthetic-data pipelines keep delivering, the latest models may already have the research taste for paradigm-level breakthroughs (an elicitation problem, in the executive's view), and there is likely a level of intelligence that should not be exceeded. Pressed on whether he would back a FLOP cap on the next pretraining run, say 10^27, the executive reportedly called that reasonable. Prakash questioned what such a cap implies for the compute build-out; Labenz argued compute would still have many uses, including safety monitoring.

    On competition, Labenz reported a consensus that only two or three labs are truly pushing the frontier, with Gemini 4 possibly making it three, and that xAI and Meta are keeping pace partly through distillation by way of the RL-environment industry, which has Claude building environments that other labs then train on. Prakash added party chatter that Musk has trouble retaining talent and is betting on infrastructure with a 2029 to 2031 horizon, while Anthropic and OpenAI plan for AGI and RSI around 2028 to 2029. Labenz said attendees leaned even shorter, with many talking about next year as the critical period.

    Labenz also described labs spending more compute on safety, citing Jensen Huang's 20/80 design-versus-verification analogy, activation monitoring, and using frontier models to red-team and repair hackable RL environments, which he called a proto-scaling law for cheating. Insiders expect to stay bottlenecked on alignment but to find enough low-hanging fruit for the next generation or two, beyond which all bets are off. Prakash agreed that it is normal not to see past two generations on an exponential.

    Prakash raised Anthropic's rush to IPO, reportedly slipping from late October to early November, and wondered whether it reflects cash-flow needs or a desire to spread shares before the next revenue ramp. Labenz said the IPO never came up, but that audience members pressed Anthropic leaders on whether they would press their commercial advantage, given the safety community's impression at founding that the company would not push the frontier; he said they got no real answer. The hosts then segued to the guest.

    Timestamp links open the original source recording.

    Definitely the strongest statement I've heard about imposing kind of a hard cap on capabilities even for a time.

    The more you clean up the RL environments, you do, in fact, get lower rates of cheating downstream.

    Lightly edited · timestamps jump to YouTube
    1:29

    Prakash Narayanan: Good morning. It is Tuesday, October 6, 9 AM. Nathan, good morning to you.

    1:33

    Nathan Labenz: Good morning, Prakash. How are you today?

    1:37

    Prakash Narayanan: I'm very good. We met over the weekend at Lighthaven, which has now become an infamous venue, that den of infamy. Lighthaven had a conference called The Curve over the last few days, and you attended. A good number of people from the frontier labs attended too. I saw Jack Clark there, and a number of our former guests, [unclear] and Zvi Mowshowitz. I was walking around thinking, that guy looks really familiar, and then realizing, oh, that's [unclear], the field specialist. So a good number of leading lights in the field. What was your takeaway from The Curve?

    2:28

    Nathan Labenz: It was a jam-packed weekend, and it was great to see you in person for a change. We've logged a lot of hours here in the studio, but not too many in the flesh. The Curve is a really interesting and pretty unique event in the AI space. It was started by a couple of people who had the idea that we need to bring more contrasting opinions together. My own meta-strategy in AI is to observe that everybody is so deep down their rabbit hole, advancing their frontier, usually having success, but the pace is so breathless that nobody has time to come up for air and see what's going on. I'm kind of doing the zag, trying to cultivate the broadest view possible.

    3:13

    In the event world, we're going to talk to swyx today about his upcoming AI Engineer event, which is a great example of a leading event where lots of people go and do a lot of learning. I've been to one of those too, and it was a great experience. But it's very focused on a particular point of view: how are we going to make these AIs help us become better engineers, build better products, power enterprises. It's not really a place where people come together to hash things out. The organizers of The Curve were trying to create a space where people with quite different perspectives could come together and at least attempt a meeting of the minds.

    3:58

    That's what it's been about for each of the three years it has run. It's been people from frontier companies, short-timelines people, fast-takeoff people on the one hand, but also folks who have published articles like "AI as Normal Technology," and people from DC who were initially very skeptical, thinking this was more Silicon Valley hype, but are coming around and trying to create a space for sense-making and shared deliberation. It's been really successful in that regard. People have come to it in a positive spirit, and certainly not everybody leaves agreeing on all the key issues, but it has earned a reputation as a place where people have a real exchange of ideas across paradigms, ideological divides, backgrounds, coasts, you name it.

    4:54

    Prakash Narayanan: As you say, there are some divides there. Can you speak to some of the divides that you saw over the weekend?

    5:05

    Nathan Labenz: One of the big observations in AI in general is that events have come pretty fast. If you were looking ahead a couple of years ago, you would think that with two years of additional information, actually seeing how strong the AIs get and who's been right or wrong, people would start to agree on the state of things. That has happened a lot less in the broader world than I would have expected, though I did feel some of it at The Curve. There are all these different splits: pro open source, because it's going to save us from concentration-of-power issues; anti open source, because it's going to create bio risk. There are tons of these dualities, and except on whether AI is going to get really powerful or not, I think both sides generally have pretty good points, and there are pretty serious trade-offs involved.

    5:50

    But the big thing people have come together on, at least to a significant degree, is that it does look like the AIs are going to get really powerful, and this is something we're really going to have to reckon with. We can't just hope it goes away, or feel that if we wait, the bubble will burst or the fever will break. This time around, people were definitely like, okay, it's getting pretty serious. There were very few perspectives that... as Helen Toner once famously put it, long timelines have gotten super short. Another corollary is that what is called normal in terms of expectations is getting pretty incredible. Even the people there with the most modest expectations for what AIs will be able to do were like, yeah, it's going to do an awful lot. The question on the capabilities front has become: is RSI going to foom? Is it going to be fast for a while and then level off? We've got these super powerful things, but is it really going to run away from us? Those are still the capabilities questions we're getting asked, but there has been some healthy updating across the community based on what we've seen.

    7:34

    Prakash Narayanan: Were there any limits or constraints that the frontier labs said should be enforced on players in the space, including themselves and others?

    7:57

    Nathan Labenz: That was a huge topic. Probably the biggest divide in perspective is inside the frontier companies versus outside them. Everybody knows they're living in the future and have visibility that the rest of us don't. I had a chance to talk to and hear from them, in some cases in sessions, in some cases in passing conversations, all under Chatham House rules, so I'll abstract away. We're talking about founders, executives, top researchers at frontier companies. I asked a small set of them, why come here? Before doing a couple of sessions where I was interviewing or moderating, I asked the same question: what makes this next hour a good use of your time?

    8:42

    They repeatedly expressed that it's important to them to share as much as they can with the broader community. They didn't want this to turn directly into a media product, but they felt this is an event where people come together in good faith. There are a lot of people they've known for a long time, and some new people too, in a pretty high-trust, high-goodwill environment where they could share their honest perspectives. Definitely not trade secrets, of course, but a candid take on what they're seeing internally, what they believe internally, what their outlook is, what the plan is, and how we can come together and potentially impose some of these constraints.

    9:27

    There were a number of things I thought were newsworthy. First, a senior executive at a frontier lab said that pretraining continues to deliver: the models are, at the base level, getting stronger and stronger, and there doesn't seem to be an obvious stopping point. Scaling laws are holding, or maybe even bending favorably, because the data quality is getting better, they have increasingly different architecture tricks, and there's lots of work on synthetic data. Synthetic data is definitely delivering. You need more pretraining data, so what do you do? You give the model a bunch of data and ask, what could I transform this data into that would make the model smarter if I pretrained on it? You're converting reasoning tokens, test-time compute, back into pretraining data, which goes into the next generation of models. That's one way in which recursive self-improvement is happening.

    10:58

    Of course there are opportunities for architectural improvements as well, but a big emphasis was on data quality: we can now spend tokens on the existing data to augment it, to create better, cleaner, different lenses on it, feed all that back in, and the models get better. So this executive was really emphasizing that pretraining continues to deliver and the models are getting super powerful. RL is important too, but his take went as far as saying he believes the latest frontier models do have the ineffable research taste needed to start making paradigm-level breakthroughs. Excuse me, I lost my voice a little talking so much.

    11:44

    That's a key statement. He believes the pretrained models have the capability to make paradigm-level breakthroughs, and that it's currently an elicitation problem: the RL isn't bringing it out well enough, and they don't know quite how, so it's rare. This is why you need 10,000 agents to do a Millennium Prize problem, which is verifiable too, so RL is working better there. Maybe you need a million agents to stumble onto something eventually that would give you paradigmatic change at the ML research level. But he believes that capability is there, and they'll figure out how to elicit it. Then he said a couple of things that I thought were legitimately newsworthy. One was that there likely is a level of intelligence that we just shouldn't go past. "Is" might be a little strong, but there likely is.

    12:45

    Prakash Narayanan: What?

    12:47

    Nathan Labenz: Yeah. It was the first time I'd ever heard that from a frontier lab leader, and it was not minced words. He wasn't saying there's definitely a level, or that he can articulate what it is. But this is a person whom everybody would recognize by name and position, saying, I think it is likely that there's a level we shouldn't go past. Whether that means not ever or just not right now was a little ambiguous, but it was a pretty firm statement. Definitely the strongest statement I've heard about imposing a hard cap on capabilities, even for a time.

    13:30

    Prakash Narayanan: What confuses me about that statement is that some of the increases in intelligence we're seeing, number one, are coming as a result of increases in compute, which has traditionally been what has happened over the last decade and a half. Number two, they're coming as a result of refining the input data. The input data is being refined by greater intelligences, so you have much better input data. When we say we shouldn't exceed a point of intelligence, that automatically means you have to look at the inputs going in, and the inputs are compute and this refinement of data. So if you say we shouldn't go beyond that point, that also means you shouldn't build compute beyond that point, or refine data further beyond that point, which means a lot, really. It means the end of the compute build-out, which is pretty significant, right?

    14:33

    Nathan Labenz: Well, I think there are a lot of uses for compute. I had a chance to follow up on this. This combination of pretraining really delivering and scaling laws, if anything, maybe bending more favorably, although it's tricky to know what you count. Do you count all those FLOPs you use to create the synthetic data? Maybe you should. How to count these things is always vexing. But given there might be a cap on how far we should go, I asked a follow-up: okay, how do we operationalize that? Would you be prepared to sign on to a limit on the number of FLOPs that go into the next pretraining run? I don't know what the number would be, and again, what exactly would count, but should we say no more than 10^27 FLOPs in the next pretrain or something like that?

    15:18

    I anticipated that being a bit of a straw man for him to react to, or tell me what was bad about it and ideally give me a better idea. But what actually happened was he basically said, yeah, I think that could be reasonable. The end. Not much of a fight there at all. Again, the details will be incredibly important if we go down that path, because the more you put into enriching the data, you can definitely play a shell game of hide-the-compute these days. But the sentiment was, yeah, we might have to do something like that. And then, of course, I can anticipate your question. I know you well enough to predict the next Prakash token at this point: what about Elon? What about

    16:17

    Prakash Narayanan: Zuck?

    16:18

    Nathan Labenz: Mhmm. There were a couple of really interesting things people said along the way. This is not from one individual, but a takeaway from the event in general: two to three companies are really pushing the frontier. With Gemini 4 being announced now, we maybe have a three-horse race again; we'll see as that comes online. And even the other two we would usually think of in the top five, Elon and Zuck, are actually distilling a lot, whether they know it or not. The way they're distilling these days isn't calling the Claude API, getting traces, and feeding them in. Rather, the cottage industry of RL environment makers and sellers is in fact functioning as a distillation conduit.

    17:03

    What they're all doing is having Claude create these RL environments, selling the RL environments to other frontier model makers, and then those companies do their internal RL on an environment that only exists because Claude was smart enough to make it. That's how they are constantly capturing the capability advances that the true frontier companies have made. So there wasn't actually a ton of talk about competition beyond the top three. One person, another founder-executive type at one of these companies, went as far as to say that yes, they have kept up. Obviously there's a gap, and it's been persistent, but they've maintained a consistent following distance.

    17:49

    This is in contrast to Anthropic's old prediction from their fundraising deck, which I always think about, that maybe in 2026-ish the companies that train the best models would be so far ahead that nobody would be able to catch up. Obviously that hasn't happened. But the synthesis of all this is that if they hadn't been releasing the models, and if people weren't able to use these various direct and indirect distillation techniques, they actually think that probably would have happened: the other companies would not really be keeping up but for the leak of intelligence in all these different ways, including the RL industry that is gradually finding ways to transfer capability from frontier models to their competitors.

    19:01

    Prakash Narayanan: I did speak to some people at the party, and one of the prevailing pieces of conventional wisdom there was that Elon doesn't really have talent because no one wants to work with him anymore, so Elon was not regarded as serious in that sense, because he's already chased away everyone who started working with him. On Elon's side, he is very focused on the infrastructure build-out. He's making a gamble that infrastructure matters more than the models and that he'll be able to catch up using his build-out. He's structuring things for 2029, 2030, 2031: TeraFab, things like that will come out much later, and the SpaceX satellites are a 2030s thing. He's structuring for a much longer period of time. And I think Anthropic and OpenAI are both structuring for more like the two-to-three-year mark. They're looking at a 2028, 2029 AGI, RSI. And we

    20:19

    Nathan Labenz: Yeah. Honestly, even sooner, I would say, was the vibe. This is probably somewhat biased toward a short-timelines crowd, but again, that crowd includes the executives, top researchers, and founders of these companies. They were talking more about just next year.

    20:38

    Prakash Narayanan: Yeah.

    20:39

    Nathan Labenz: It was really a sense that the critical period is beginning now. RSI is kind of at hand. Definitely not full agreement on how fast it goes, or if or when it tops out, but a shared sense that the decisions made over the coming months, and the way compute is used over the next year, could be really critical. There wasn't even much 2028 talk, to be honest, which is pretty crazy. On the compute build-out, we'll see where value accrues, but I think Elon's bet is not necessarily a bad one even if he can't really compete at the model layer. I don't count him out, but I'm just reporting the news from the front. In terms of whether a cap on pretraining scale, or another compute input cap, that might constitute a pacing mechanism, means the end of the compute build-out, I would say no.

    21:24

    Another question I had the chance to ask relates to what everybody has heard from Jensen [unclear]. His comment was that these companies are going to have to mature and change the ratio of where they spend their resources. He said at NVIDIA, they spend about 20% of their effort on designing the chip and 80% on verifying, validating, and testing all the edge cases and making sure it's going to last a long time and be reliable. His point was that they've been working really hard until now, putting everything they can into making their models capable enough to be useful. Congratulations, you've succeeded. Now you're going to enter the era where you have to make them safe, reliable, and trustworthy. So I asked one of these people what they thought of that take from Jensen, and the response was basically, yeah, that's kind of reasonable, again without knowing exactly where the numbers will shake out.

    22:54

    The possibility is that the majority of compute in the future could be going into all sorts of safety measures, which could be monitoring, could be chain-of-thought monitoring. People are getting kind of bearish on chain-of-thought monitoring.

    23:13

    Prakash Narayanan: Yeah.

    23:14

    Nathan Labenz: But then there's internal activation monitoring. They're already spending a significant percent of compute on monitoring, and it sounds like that might go up a lot. Another thing they're spending compute on in a big way right now is fixing the RL environments. Everybody understands at this point that if you have sloppy RL environments that reward cheating, you're going to get a lot of cheating. So what to do about it? I wouldn't say we're quite at the point where they have a scaling law for cheating, but it sounds like they might be getting kind of close, because they're looking at all these RL environments and going nuts with them. They're taking their best models, and this is another

    24:08

    Prakash Narayanan: way of kind

    24:12

    Nathan Labenz: of another form of recursive self-improvement. A model was trained on a bunch of hackable RL environments where cheating was rewarded, so it's a bit of a cheater. That's a big problem, but you can still get lots of valuable work out of it, as we do all the time. They are now applying these models to the RL environments themselves and saying, hack this environment. That is now the direct task. They're finding all these places and ways in which the environments can be hacked, they're fixing those, and I'm sure they're throwing out some environments that are fundamentally flawed. Gradually they're driving down the rate of flaws, which means they're also driving down the rate at which cheating is rewarded, which in turn means you get less cheating from the models as they come online.

    24:57

    I got the sense, although I think this is still, as they say, an open research question, that they are seeing a pretty clear relationship: the more you clean up the RL environments, the lower the rates of cheating downstream. Obviously they'd love to have zero hackable environments. That's going to be tough. It doesn't seem like we're all the way to a scaling law, but we're seeing a proto-scaling law, in the sense that you can spend a lot of compute driving this down and get better behavior on the other end. I heard repeatedly that they expect to be bottlenecked on safety and alignment, but it seems that for at least the next generation or

    26:07

    Prakash Narayanan: two,

    26:08

    Nathan Labenz: there is still enough low-hanging fruit in the safety and alignment domain, in part just by fixing RL environments, that we should expect things to continue to move pretty fast. We should expect more powerful models that they will probably feel confident enough to release. I would expect maybe two orders of magnitude less cheating by doing this. I don't know, and this is where the exact shape of the scaling law gets important. If you have one environment that rewards cheating, is that enough to teach your model to cheat? Probably not, I would guess, but maybe it could be.

    26:55

    But how low can they drive that? My guess is they'll take the rate at which the environments are hackable super low, and cheating will go much lower too. We'll have to see what that ratio looks like. It seems they're probably going to be able to get comfortable releasing the next couple of things even as they consider themselves bottlenecked on alignment. But past that, they're like, yeah, all bets are off. It's really hard to say what happens past the next one to two generations.

    27:26

    Prakash Narayanan: I will say that, being on the exponential, it's not unusual that you can't see beyond two generations. That's par for the course. You would have to be a little bit insane, or a cultist, to believe that in five or ten years we're going to have [unclear]. That's a tough one. So it's not unusual that you wouldn't be able to see past that period. I wonder what comes next. Last year we had this end-of-year process where both companies dropped models around Thanksgiving, and everyone kind of checked out for a month and a half and had a chance to use these things at home. When we came back in January, we started to get, okay, coding is kind of solved, and by March it was this enormous climb, all the way up.

    28:11

    Right now, Anthropic is rushing toward the IPO. It was supposed to be late October, and now it's pushed back to the first week of November. Last call is just for Thanksgiving. So I think they're trying to rush out the IPO, and the question for me is why, which I'm not too clear about. I don't know if they have a cash-flow necessity, because they have more than $100 billion of commitments for next year. Or do they want to get the shares into the hands of the public before the next ramp-up in revenue, and spread out the wealth before things happen? Or is it something else? I don't know, but there seems to be intense pressure to get the company to IPO this year, which I think is rather unusual.

    29:33

    Nathan Labenz: I believe that never really came up in my conversations at all. The assumption seemed to be that these two companies will be able to raise the capital they need one way or another, whether public, private, whatever the case may be. There was definitely some interesting discussion around how hard Anthropic will press its advantage, to the degree that it has one right now. OpenAI folks seemed kind of unsure what to make of Anthropic, and there were a number of questions from the audience to Anthropic people over time, along the lines of, okay, we're trying to get serious over here.

    30:18

    Going way back to Anthropic's founding, and there are a lot of people at this event who were in the mix at and before that founding, the sense the AI safety community had was that Anthropic was committed to not advancing the frontier. They wanted to be at the frontier so they could be relevant and do all the research, but they weren't going to push it, because they didn't want to make the race condition worse. People challenged some Anthropic leadership on that, and their response was, well, what we said was that we would never publish methods that would advance the frontier. And people were like, I don't think that's really what you were saying at the time. I don't know exactly what was said in every instance, but people had the impression that they were not going to be pressing commercial advantage.

    31:03

    And now they've been asked: right now you probably have a little lead if anything, but it's pretty competitive. OpenAI is trying to get its house in order, you definitely have your work to do too, but you seem to have maybe a little more in order than they do. Are you going to chill and let them get their house in order, rather than try to run away with the competition, or not? We didn't really get an answer on that, so I don't know what to expect. There actually wasn't much discussion at all of when the next model will launch, that kind of thing. So I didn't come away with any major updates. I would hope they would show a little friendly restraint, but I'm not sure what to expect.

    31:58

    Prakash Narayanan: Yeah. The story keeps changing, right? And it's not unfair, because as things on the ground change, you do have to change what you're doing.

    32:12

    Nathan Labenz: Yeah.

    32:13

    Prakash Narayanan: So

    32:14

    Nathan Labenz: As President Xi said, a man of wisdom adapts to circumstances.

    32:17

    Prakash Narayanan: Indeed. And speaking of adapting to circumstances, let me kick off a great segue to our guest for today.

  2. 11:28Interview20 min
    Thomas Sohmers: Commodity memory, looped models, and AI chip designThomas SohmersPositron AI’s cofounder explains its Series C, why it uses LPDDR instead of HBM, why looping saves capacity rather than bandwidth, and how frontier models now handle verification work on its Asimov chip.
    Open segment on YouTube ↗

    Thomas Sohmers, cofounder and chairman of Positron AI, joined to discuss the inference-chip startup's new funding round and its bet on commodity memory. Sohmers clarified that the round was not led by Oracle, which is a customer, but by Valor Equity Partners, Atreides Management, NEA and others, with Netscape and SGI cofounder Jim Clark participating. He said Positron went from cold start to a shipping server in about 15 months by starting with FPGAs, and is now using AI to compress the loop on RTL, verification and place and route for its custom Asimov silicon.

    Asked by Nathan how he taxonomizes chip strategies, Sohmers said Positron aims to win on performance per dollar and TCO by building a fairly general linear algebra processor, and dismissed analog and photonic approaches as still science experiments. He said Positron uses commodity LPDDR memory rather than HBM, claimed 93% sustained memory bandwidth utilization against roughly 30 to 40% for NVIDIA GPUs in decode, and described a Credo-partnered memory chip reaching 72 LPDDR5X channels and up to 2.3 TB per chip for Asimov. On a roughly 4.5x memory price increase, he argued that all vendors bear the cost and that the value delivered per gigabyte has risen even faster.

    Nathan pressed on looped transformers and the hardware angle. Sohmers traced the idea from duplicated-layer experiments on open models to rumors about GPT-6 Astra, and suggested early exit and expert re-routing within the forward pass as extensions. He argued that looping chiefly saves memory capacity, including for the KV cache, rather than memory bandwidth, since the same bytes must be fetched on each pass. Prakash asked what Positron cut before design freeze; Sohmers said features were cut for speed, including optimizations for sparse attention and linear transformers that he thinks US labs are not yet using at scale.

    On AI in chip design, Sohmers said Asimov's effort splits roughly 60% design to 40% verification. He said GPT-6 Astra and then Opus 5.5, handed full access to Cadence Palladium emulation infrastructure, closed the verification loop on their own, reading documentation and building test harnesses, a step change over GPT-5.6. Asked by Prakash about genuine innovation, he said he has not yet seen the models out-design his team, but would not be surprised if brute-force iteration soon does.

    Nathan asked about token spend and manufacturing access. Sohmers said tokens are now Positron's largest non-manufacturing line item, peaking above $100,000 a day after Astra's release before tempering when Opus 5.5 matched it on many tests at a quarter of the price, and he insisted on using only the best model for real development work. On supply, he credited his COO's long TSMC relationships, said Positron is a direct TSMC customer with TSMC-backed Venture Tech Alliance as an investor, sticks to N3P, and has sidestepped the memory bottleneck by using commodity memory. The segment closed after a brief connection drop.

    Timestamp links open the original source recording.

    We want to be the TCO win, and TCO wins require you to be fairly general.

    Looping, really, is more memory capacity savings than memory bandwidth savings.

    I really don't want anyone to be running a real software or hardware development task on anything that is less than the best model.

    35:25How did the Series C come together?
    Sohmers said the round was not from Oracle, which is a customer, but was led by Valor Equity Partners, Atreides Management, NEA and others, with Jim Clark participating, on the heels of the Oracle deployments and ahead of next-gen tape-out.
    36:17How did Positron speed up chip design and deployment?
    Sohmers said it took about 15 months from cold start to shipping a first server, helped by starting with reconfigurable FPGAs and now by using AI across RTL, verification and place and route for the custom second-gen chip.
    38:18How do you taxonomize chip strategies, and where does Positron fit?
    Sohmers said Positron targets TCO via a fairly general linear algebra processor that is generally manufacturable, using commodity memory instead of HBM, while he sees analog and photonic approaches as still science experiments.
    41:43How have rising memory prices affected your TCO projections?
    Sohmers said a memory quote is up 4.5x in a year and might double again, but all vendors bear the increase, and the value delivered per gigabyte has risen even faster.
    44:03What trade-offs and engineering challenges come with commodity memory?
    Sohmers said commodity memory is the only thing that scales cost-effectively as context and agent counts grow, and Positron compensates with 93% bandwidth utilization and a Credo-partnered memory chip that reaches 72 LPDDR5X channels.
    48:17What parts of networking and coordination disappear when a model fits locally?
    Sohmers said Asimov's up to 2.3 TB per chip is eight times a B300, so models needing eight GPUs fit on one device, removing all-gather and all-reduce overheads and software complexity.
    50:17Why do models loop within the forward pass, and how does that relate to hardware?
    Sohmers traced looping from layer-duplication experiments to rumors about GPT-6 Astra and suggested early exit and expert re-routing as extensions. He said the hardware benefit is capacity, not bandwidth, since the same bytes are fetched each pass.
    54:51Is looping about keeping weights on-chip to cut memory bandwidth?
    Sohmers said no: on-chip SRAM is far too small for reuse, and the main gain is halving capacity, apart from a narrow MoE benefit if the same experts are reused.
    57:43What design decisions do you defer or cut before freezing a chip?
    Sohmers said Positron cuts features to ship faster, including sparse-attention and linear-transformer optimizations, and does not regret it given its focus on big US model companies.
    1:01:00How does looping affect the KV cache?
    Sohmers said the KV cache is per layer, so halving layers halves capacity, but the same total bytes must still be fetched on each loop.
    1:02:27What is your design-to-verification ratio, and where is AI really helping?
    Sohmers said roughly 60% design to 40% verification, with GPT-6 Astra and Opus 5.5 closing the loop on Palladium gate-level emulation by reading docs and building harnesses, a step change over GPT-5.6.
    1:07:51Have you seen genuinely new innovation rather than rule following?
    Sohmers said not yet: the models question unconventional choices and default to traditional designs, but he would not be surprised if fast iteration soon matches or beats human design.
    1:11:56How does token spend compare to salary spend, and how do you secure manufacturing capacity?
    Sohmers said tokens are the largest non-manufacturing line item, peaking above $100,000 a day. He credited a direct TSMC relationship through his COO, TSMC investment, and a conservative N3P node for capacity.
    Lightly edited · timestamps jump to YouTube
    32:38

    Prakash Narayanan: Our guest for today is Thomas Sohmers. Sohmers is cofounder and chairman of Positron AI, a company that builds hardware for inference, running a trained AI model to answer questions, write code, or perform other tasks. Positron concentrates on the memory systems that feed data to a processor, aiming to make those tasks cheaper and more energy efficient. Thomas previously founded Rex Computing, a processor startup, and was a 2013 Thiel fellow. He later became principal hardware architect at Lambda, which provides access to graphics processors for AI developers, and worked on its early cloud infrastructure. He then served as director of technology strategy at Groq, another AI chip company.

    33:23

    His experience includes both designing processors and putting other companies' processors to work in data centers. Positron was founded in April 2023 and shipped its first Atlas server in August 2024. Atlas uses programmable chips. The company's next generation, Asimov, is custom silicon intended for Titan servers. Positron reports that more than 50 Atlas racks have been deployed at Oracle. Alongside that work, Thomas has questioned a familiar assumption about computing equipment, that advancing software necessarily makes yesterday's machines less useful. In February, he pointed to the first V100 systems he deployed at Lambda and said they were still in service. Positron has recently raised a Series C, I believe, from Oracle.

    34:08

    Let's get Thomas up here. I'll just wait a moment while he comes up. But it's interesting, because I think we now have a series of very successful chip startups. At this point, we have Cerebras. Cerebras was, for a moment last week, in trouble, as SemiAnalysis put out a note saying that Astra Ultrafast is not using Cerebras chips. SemiAnalysis then put out a note saying that's because they're fully sold out. But in the time period between those two notes, Cerebras' stock had taken a beating. It seems that Cerebras is currently being used by our friends at Jane Street. So let me bring up Thomas here. Hi, Thomas.

    35:15

    Thomas Sohmers: Hey, guys. Sorry for that technical difficulty.

    35:18

    Prakash Narayanan: No worries. It's a pleasure to have you on the show.

    35:22

    Thomas Sohmers: Thank you. Thanks for having me.

    35:25

    Prakash Narayanan: The big news is you guys have just concluded a Series C fundraising, I believe, from Oracle. Tell me a bit about how that whole deal came together.

    35:35

    Thomas Sohmers: The funding was not from Oracle. We have Oracle as a customer; they're the first large-scale deployment of our first-generation product. This new fundraising was led by Valor Equity Partners, Atreides Management, NEA, Andra Capital, and Jim Clark, the cofounder of Netscape and SGI. It is very much on the heels of the Oracle customer deployments, and in anticipation of our next-generation silicon taping out at the end of the [unclear].

    36:17

    Prakash Narayanan: One of the things that you've done is speed up the process of designing and deploying the chips. Can you tell us a little bit about how you did that?

    36:26

    Thomas Sohmers: Positron was started about three and a half years ago. A core founding tenet was that we want to get real hardware into customers' hands as quickly as possible. It was about 15 months from cold start of the company to shipping our first customer server, and about three years total to win Oracle as a customer. That's pretty unheard of in the new AI hardware space. Obviously it's a great team, but it also came down to leveraging AI and going pedal to the metal on iterating on the design.

    37:11

    That came down to us starting off with FPGAs as the basis for our first-generation products. These are field-programmable gate arrays, hardware that can be reconfigured, which gave us a very rapid iteration loop and spearheaded the rapid development of our first-gen product. For our second gen, which is full custom silicon, the advancements in AI over the past year have been amazing, specifically in our ability to close the iteration loop on everything from RTL development and all the front-end chip design work, to verification, which historically is the biggest time limiter for chip designs, all the way to full back-end place and route and all the complex physics involved. AI has drastically accelerated our own development of chips to make AI cheaper and more energy efficient.

    38:14

    Nathan Labenz: Sounds like recursive self-improvement to me.

    38:17

    Prakash Narayanan: Indeed.

    38:18

    Nathan Labenz: Several things I would love to dig into there. For starters, how do you taxonomize the different strategies that chip companies have at a high level? GPUs are highly flexible, and we all know about the CUDA moat. Some people are going as far as burning particular models into the silicon itself. How do you see that constellation of big-picture strategies, and where do you guys fit?

    38:52

    Thomas Sohmers: My core focus from the beginning, as I said, is rapid iteration. Number two is performance per dollar and TCO. A lot of other chip companies have tried to carve out niches, be it ultrafast, or focusing on one particular model and getting efficiency there. We want to be the TCO win, and TCO wins require you to be fairly general. So we set out to build a fairly general linear algebra processor. We think about what the unifying factor is, not just of all transformer models today, but if I were to ask whether an architecture designed 50 years ago could still be used today, can the underlying hardware we design today still be used for whatever computational applications exist 50 years from now?

    39:37

    Thankfully, the fundamentals of linear algebra have not changed. What we're building won't be fast in 50 years, but it will still fundamentally function. Coupled with that general philosophy, what we build also needs to be generally manufacturable and something that can scale. There's a whole other set of categories, which can be lumped into a bucket, of people trying new analog circuit designs or light-based photonic computing. A lot of those things are very cool as an engineer, but they are still much more in the science-experiment stage than something that could be mass-market commercialized.

    40:22

    With our philosophical underpinnings, and in how we've executed and are pushing our next-gen product, it's all about how we can scale to gigawatts of compute as quickly as possible. So we're leveraging commodity memory technology as a core part of our supply chain, rather than the very expensive and very supply-chain-constrained HBM that everyone else in the high-performance AI accelerator space is using. It's a pretty different approach, and the aim is both to provide the best TCO for customers and to scale out to massive deployment sizes.

    41:43

    Prakash Narayanan: There have been increases in prices of all kinds of memory modules in the last year or so, pretty significant increases. Has that affected your projections on the TCO going forward?

    42:03

    Thomas Sohmers: The sad reality is that I was looking at a quote we had a year and a week ago for memory, and it's gone up 4.5 times since then. I really hope this isn't the case, but I wouldn't be surprised if it ends up being another doubling over the next year. On TCO, with all else being equal, everyone else is in the same boat in having to increase their COGS and then their prices, so that doesn't affect us that much from a competitive standpoint. The fear would be that if the cost goes up and you're delivering the same value, you have a problem. But a big reason everyone bears these cost increases is that you're delivering more than five times the value compared to last year.

    42:48

    From a primitive capitalist perspective, the whole point of pricing is to let the things of higher value receive the goods that are limited. The system is working in that sense, because memory is getting to the applications that have the greatest value, even if that value is generating AI videos that we also share as memes. But compared to a year ago, the capability of a model using x gigabytes of memory is way more than five times what it was this time last year.

    44:03

    Nathan Labenz: We've got an abundance of memes already, if not an abundance of everything else we might want. Help us understand a little more. I know there are different kinds of memory, and we know the very basics, but what kind of trade-offs does it involve, and what engineering challenges do you take on to get the advantage of using commodity memory versus the very high-end memory? And how does that flow all the way through to the customer experience?

    44:39

    Thomas Sohmers: I definitely did not expect this level of price increase, but at the start of Positron I did think memory was going to be the limiting factor going forward, even though people can debate power and other parts of the infrastructure stack. My fundamental belief was that there isn't going to be a stop anytime in the near future to scaling laws and getting increased value from making models larger, and that eats into memory on one side. On the other side, I think the main limiter today of AI being applied to most applications is context length, being able to hold more context per user and then scale that to drastically more users.

    45:24

    Today the driver of, quote-unquote, users or individual sessions is really just having more agents. Three or four months ago I had on average two to four agents running constantly in the background; now I'm running 15 to 20. Multiply by how many people are using AI and how many completely separate concurrent contexts there are, and it adds up very quickly. That memory driver had us say we have to use commodity memory, because that's the only thing that's going to scale and be cost effective. When we set that as the constraint in architectural design, we had to come up with very clever, innovative solutions.

    46:10

    The two main pieces, and what we showed with our first-generation product, start with achieving extremely high memory bandwidth utilization from a compute architecture perspective. NVIDIA GPUs, on average, get between 30 and 40% memory bandwidth utilization in the decode forward pass of a transformer model. Even though they advertise, say, 8 terabytes per second of theoretical memory bandwidth with a B300, you only see something in the ballpark of 2 terabytes per second of realized bandwidth.

    46:56

    That comes down to a whole bunch of GPU architecture details, and to reuse patterns that don't really exist in transformers, which they designed their hardware architecture around for training and other workloads. With our first-gen product, we were able to hit and sustain 93% of theoretical memory bandwidth, so theoretical and realized basically become one. That's a massive 3x improvement right there. But to really take advantage of the belief that models are just going to get larger, we have to scale to much more capacity. So we're partnered with Credo Semiconductor and have developed a memory-chip solution that lets us go from the maximum number of LPDDR channels found in any other product, on the order of 12 to 16, up to 72 channels of LPDDR5X with this decoupled memory chip. LPDDR is the same type of memory that is in phones and laptops.

    48:17

    Prakash Narayanan: So walk me through a model that fits locally. Which parts of the networking or coordination start to disappear?

    48:26

    Thomas Sohmers: A very big part of it is that if you have all this memory, with Asimov, our upcoming chip, we have up to 2.3 terabytes of memory capacity per chip. The B300 shipping today tops out at 288 gigabytes. A lot of the original deployments actually used 180 gigabytes as the base SKU for the B200s. And as reported by analysts, NVIDIA for next generation is actually cutting down the memory capacity because of cost and everything involved, so they're only going to be at 192 gigabytes with Rubin Ultra.

    49:11

    Even if you use that 288 as the current benchmark, we've got eight times more memory capacity per chip. That means what would have needed eight GPUs from a memory perspective, we can do on a single device. And it's not just the silicon cost savings. Whenever you have to scale to more than one device, there are overheads: all-gathers and all-reduces for each layer, really each multiply, of sharding across those devices. By having it on a single device, you remove a lot of the complexity, both software and from a deployment and build-out standpoint, and it also comes down to real performance in the end.

    50:17

    Nathan Labenz: The nature of the hardware people can access determines in part what kind of models they can deploy. A huge topic recently has been looping within the forward pass and what impact that might have on our ability to monitor models effectively. Chain of thought is becoming less reliable as a monitoring technique. My best understanding from what I've heard is that it's not so much because of the looping, and more because pretraining is getting more and more powerful and able to juggle more advanced representations internally. I'd love your perspective and breakdown on how these hardware considerations relate to looping. I think I know the answers to some of these questions, but maybe start with, why do we loop? Why is that even a tempting thing to do? And then branch out across different hardware options and how that might shape the future.

    51:24

    Thomas Sohmers: It's been very interesting, from the looped transformers paper to its being the rumor that this was one of the big advancements with GPT-6 Astra: that simply repeating the forward pass gets you that improvement. It's crazy, because there was work that I think was only being done on the fringes of the open-source transformer community two or three years ago, of taking something like a Llama 70B and duplicating layers within it. You'd take a 70B model, strategically repeat sets of matmuls in sequence, and turn it into something like a 100-billion-parameter model. And you got better results from it, even though it was repeating the exact same matmuls. So there is an

    52:15

    Nathan Labenz: Imagine what you can do when you train it to work that way.

    52:18

    Thomas Sohmers: Exactly. There were also a couple of interesting things years ago that preceded the multi-token predictor. Today, for a speedup, you can add a layer at the end that tries to guess what the next token will be after the one the rest of the model has predicted, and you can do very cheap prefill on that to see if it's correct, so you basically get two tokens for the price of 1.1. There was a precursor to that idea called Medusa heads.

    53:04

    The reason I'm bringing this up is that I think this concept can be expanded in many ways: training sidecars, or things within the layers themselves, to better predict what's going to happen and take some insight of where you are in the layers. If you already have very high confidence about the next token and multiple layers have agreed, decide to exit early. Or if you're really uncertain, find which layers in the network would increase that probability. These are the same sort of functions the MoD router performs in determining which experts should be selected on a per-token basis.

    53:49

    If you extend that concept: if you're really unsure, with a very large set of equally weighted probabilities for the next token, and you're 75% of the way through the layers, do you just route back and determine that you need a different set of experts? I think there's a ton that can happen within the layer boundaries we have today. Doing this as part of pre- and post-training is optimal, but if I had more free time, and thankfully I can get some of it by sending off agents, it's nice to be able to run experiments on open-source models today. I wouldn't be surprised if these things are already being done at scale in the major model labs.

    54:51

    Nathan Labenz: Can you go a little further into the hardware connection of that? I'm coming in with assumptions in very fundamental terms about why loop at all. I think it has to do with memory bandwidth: now I can keep the same weights on the chip, and I don't have to shuffle things in and out as much. No?

    55:15

    Thomas Sohmers: Not really. The assumption in inference is that you have all of your weights locally in DRAM, sharded over some number of devices. You do get some amortization, but with the size of experts and the sizes of caches on chips, there's not much reuse opportunity; by the time you're done with the layer, you've gone through many times the amount of memory that's in the on-chip SRAMs. There are two ways a loop could occur. In an MoE, you may reuse the same experts, which would be an advantage if you don't have to go through an expert-parallel fan-out routing decision again.

    56:01

    But I think that's fairly narrow, and I don't know what's being done in the big labs. I would guess that if you are looping, you probably want the opportunity to select different experts than you did the first time, to get improvements in quality. That's a guess. From a hardware perspective, looping is really more memory capacity savings than memory bandwidth savings. Say you trained two models from the same base set, and trained one to loop, so it has only 10 layers instead of 20. Model B is 20 layers, doing very simple work, and is double the size of the 10-layer version. If the looped version gets 95% of the same quality of results at half the size, you're probably going to deploy that one.

    56:47

    So you're saving more capacity than bandwidth, because in doing the second loop you still have to do the same memory fetches and the same number of matmuls. There's the exact same number of bytes that need to be moved and the exact same number of FLOPs in the model A and model B scenarios. It's just that model B, the 20 layers, has unique, different weights for the second group of 10 layers, and that means you require twice the memory capacity.

    57:43

    Prakash Narayanan: When you design a chip, there is a certain point where you freeze the design and hand it over to start the rest of the process. What are the decisions that you delay or defer until the very last moment before you hand over?

    58:10

    Thomas Sohmers: There are a lot of features, at different stages of chip development, that we left on the cutting room floor. Most are things we still want to do in future chips, but it's about what you can actually accomplish, especially with our goal of iterating quickly. It's better to get a product out quickly, even if it doesn't have all the features you want, so you get all the learnings, product feedback and revenue from actually having shipped something. So we are on the side of cutting features to save time, primarily. It's not just design time, but the verification of that component and how it integrates with the rest of the system.

    58:55

    Everything I've cut, I'm going through it in my mind and thinking I really hate that that had to be cut, but I don't hate the fact that we did, because it's what is enabling us to move quickly. To pick one thing: a lot of what's happening in the sparse attention space and linear transformers. There are a bunch of optimizations we could have made differently if a year ago we'd thought sparse attention was going to take off the way it has, at least in the open-source model space. I'm still not convinced it's being done at scale in the big model labs, but it makes sense given the constraints of the Chinese ecosystem on hardware supply, which effectively force them in that direction, so they're innovating in that way.

    1:00:26

    As far as I'm aware, none of the major US model labs are doing anything like MLA or DSA or any of these other techniques the Chinese labs have used. Our primary customer focus is the big, big model companies, so I don't necessarily think we made the wrong choice in focusing on the things we did.

    1:01:00

    Nathan Labenz: Can I do one more follow-up on the relationship between memory and this looping concept? I genuinely want to understand it better. How does the KV cache relate to this? When you said that if you loop you can have half as many layers and half as many weights, does that also mean the KV cache is half as big?

    1:01:29

    Thomas Sohmers: The per-user KV cache is on a per-layer basis, so for the exact same reasons as with the weights, we're cutting it by half in that simplistic example. It would be the same for KV caches or, if you're doing something like delta nets, the state matrix as well. It's still a capacity win, because even when you do that loop again, you still need to fetch the same KV cache. It's reduced capacity, but you still have to move the same total number of bytes. If you were doing 20 layers instead, there is, quote-unquote, a second fetch of the same amount of data that you would have had if you were doing 10 layers repeated twice.

    1:02:27

    Nathan Labenz: Gotcha. You mentioned how much AI is helping you on the verification side. In a notable interview Jensen Huang recently gave with Ezra Klein, he talked about how NVIDIA spends something like 20% of its time and energy designing and 80% verifying, validating, ensuring long lifespan and reliability. What do your ratios look like, and what more color can you give about specifically where AI is really helping? In general there's the question of whether AIs can make meaningful new discoveries. Sometimes they obviously are, but not in all domains. What about in your domain? And does that matter if it can accelerate the other 80% so dramatically that it still creates a phase change in your cycle times?

    1:03:29

    Thomas Sohmers: For us it's a little different, because we're starting from scratch, so there's a lot more base-layer design capability that we have to build. For our first full custom silicon, Asimov, it'll end up being somewhere around 60% design focus to 40% verification, thinking from project start to tape-out. RTL freeze was roughly halfway through, and then we still needed to fix some things in RTL in the back half, when you find problems. The most amazing thing, which I don't think would have been possible three or six months ago and only became possible with GPT-6 Astra and now Opus 5.5, is in our verification work.

    1:04:14

    For verification, we use Cadence Palladium emulator systems. These are giant racks full of custom ASICs specifically designed to do gate-level emulation of the silicon, which is itself something like a hundredfold improvement in the speed at which you can emulate your chip design and run verification tests. The crazy thing about Astra and Opus 5.5 is that the agents themselves have been able to do completely closed-loop iteration and testing on this. People thought it was crazy, but we handed these very powerful agents full keys to our internal infrastructure and said, here's the Palladium piece. I highly doubt these models had Palladium documentation in their pretraining, especially because a lot of the documents are new as software updates.

    1:05:45

    But looking at the agent traces, the reasoning traces and tool calls, it went and looked up the documentation, read the full PDFs, effectively compressed them by generating its own markdown cheat sheets of what to do, and then built out all the testing infrastructure in its own harnesses to access this. It's mind-blowing. Regular software RTL simulation runs on the order of 10 hertz for us; if you remove a lot of the debug pieces that slow things down, we can run full-chip emulation on the order of 500 kilohertz. That's several orders of magnitude of speedup, and it gets us closer to this RSI loop.

    1:06:32

    Right now that's really focused on implementing test programs, finding cases where they fail, and writing out reports, which get reviewed by other agents and by humans in the loop. When this really started working, about six weeks ago, around the beginning of September when Astra 6 came out, it was a huge step-function improvement over GPT-5.6, which could not do that full closed loop and still required humans at different stages. It's going to be amazing to see how that same paradigm shift applies to every part of the chip development process. I don't see it not applying to every part.

    1:07:51

    Prakash Narayanan: One of the commentaries on AI in general is that it's a lot of process-following work, for example taking a manual for a particular piece of software or hardware and implementing it or following the rules behind it. The criticism has been that it is not really true new innovation. Have you seen any signs of unexpected innovation rather than just rule following?

    1:08:28

    Thomas Sohmers: I have not yet. That's the most disappointing element of all. As you can probably tell, I think all of this is insanely amazing, and we're going to continue this exponential trend. I'm somewhat proud that it has not figured out some of our cleverness. It has questioned things, like saying this seems like a bad design decision, until it gets explained why something's done or it actually runs tests and understands why you're not doing it the traditional way a systolic array is done. Models are night and day better than a year ago. But I remember using one of the early Opus 4 releases, I forget which, having a very long dialogue, and thinking how amazing it was, that we were days away from AGI or something. Now I wouldn't trust it with my dog.

    1:09:16

    I remember looking at some of our internal design specification and asking it to compare and contrast how it thought about it versus a TPU or traditional systolic array designs, and it pointed out differences. But when I isolated that and asked a regular Opus to design something with these specifications, it would take the traditional approach. That's the training data it had. So it does make me feel a little proud, even with how unbelievably amazing the models are today, that it won't come up with the ingenuity we had as human designers. Do I think that will hold for another six months? Not sure.

    1:10:40

    I think the amazing turn of events in the past few weeks, in terms of us having access to these models, is that because it can iterate so fast and do experiments by itself, I wouldn't be surprised if it came to the same conclusions we did, or maybe made better things than we would by ourselves, just because it can go through so many different design possibilities. Think of the story of OpenAI solving Navier-Stokes: the same approach of taking 10,000 agents and sending them off for millions of man-years' worth of effort, and it found a solution. It's brute force, it's expensive, but I think that's a valid strategy that has only become possible very recently.

    1:11:37

    Nathan Labenz: Not many people left standing in the "AIs can't do what I do" category, so congratulations on your place for as long as it lasts.

    1:11:47

    Thomas Sohmers: For a couple more weeks at least.

    1:11:49

    Prakash Narayanan: So...

    1:11:50

    Thomas Sohmers: Yeah. But I'm very eager to, for one, welcome our new robot overlords.

    1:11:56

    Nathan Labenz: How does this translate today to token spend versus salary spend? And, not so much a money question as an environment question: you've just raised nearly a billion dollars. What is it like trying to get capacity from manufacturers? I imagine that's very competitive, and probably where a lot of this money is going to go, but even with nearly a billion dollars you're competing against massive whales throwing around unbelievable sums. How are you spending internally on AIs versus humans, and how are you managing to win a place in, I don't know who's manufacturing for you, but TSMC's heart, so to speak?

    1:12:50

    Thomas Sohmers: On the token spend side, it's become our single largest non-manufacturing line item very, very quickly.

    1:13:01

    Nathan Labenz: Bigger than human salaries?

    1:13:03

    Thomas Sohmers: Yeah.

    1:13:04

    Prakash Narayanan: Wow. Just recently.

    1:13:06

    Thomas Sohmers: It eclipsed human salaries and then came back down; I'll explain in a moment. Six months ago, our monthly token spend across all employees was equivalent to a single employee, granted the company was half the size. So that was very easy to budget: Claude plus GPT and a little bit of Cursor and other stuff, all for the cost of a single employee. It really started in May, when GPT-5.5 and Opus 4.6 were massive, massive steps up, and

    1:14:09

    Prakash Narayanan: Dropped off for a moment.

    1:14:10

    Nathan Labenz: We lost you. Come back. Oh, you're back.

    1:14:16

    Prakash Narayanan: Oh,

    1:14:17

    Nathan Labenz: you were back for a second.

    1:14:20

    Prakash Narayanan: He's still in the room, so I think it's just

    1:14:23

    Nathan Labenz: Yeah. If he's like me, he's got a faulty cable that keeps coming loose at inopportune

    1:14:28

    Prakash Narayanan: Can you hear me? Alright. Sorry about that. Back. Difficulties. Coming up.

    1:14:36

    Thomas Sohmers: Alright. Am I actually back now? Okay.

    1:14:38

    Prakash Narayanan: Yep. I think so.

    1:14:39

    Thomas Sohmers: Sorry, I'm dialing from my phone because my computer was having issues earlier. That's what happens when you get a phone call. So, our token spend just massively increased. Six months ago it was equivalent to one employee; in June it started to be multiple employees, and now it's more than that. Our annualized expenditure: we were peaking in the days after [unclear], when we were really trying to push its boundaries, at over $100,000 a day on tokens.

    1:15:26

    That tempered down a little in the weeks after as we optimized and weren't trying to run as many parallel experiments. A saving grace was that Opus 5.5 came out, and on a lot of our tests, not everything, it was doing better than Astra at a quarter of the price. This is the right moment to be alive, because of this competition in the market. We didn't tell anyone to cut their spending even though we were seeing exponential growth in token spend. We felt we were getting good ROI on it, and we have very high confidence that costs are going to come down.

    1:16:11

    We're using the major model labs' models for everything today, but we also have faith that we'll be able to move things to local models running on our own hardware. We do a tiny bit of this with GLM 5.3 today. But personally, my philosophy is that I really don't want anyone running a real software or hardware development task on anything less than the best model. I don't care if it's a tenth of the cost per token. It's just not worth the expense.

    1:17:16

    Nathan Labenz: That's fascinating. We're at time. I'll let you get back to work, but I'd love to hear, very briefly, what it's like competing for manufacturing allocation.

    1:17:29

    Thomas Sohmers: We're in a great position, largely thanks to our COO, [unclear], for whom Positron is his lucky 13th startup. He has the honor of having been the very first customer of TSMC in North America, back in 1987, and he's now done 13 startups all working with TSMC over these decades. So we've had a very close connection with TSMC from the start, and we're a direct TSMC customer, which is very rare these days. Basically every other AI chip startup is going through a value chain aggregator like Broadcom or Marvell.

    1:18:14

    On top of that, announced with our Series C, we have Venture Tech Alliance, a VC fund whose sole LP is TSMC, so TSMC is an investor as well. We've not had really any worries. Even with our aggressive production ramp, TSMC is the best foundry in the world and is scaling capacity like crazy. We're not aggressively chasing the latest process nodes; we're sticking with N3P, so there's a lot of capacity out there for us. And we really solved the core supply chain bottleneck, which is memory, by going in the opposite direction of most people in the space.

    1:19:15

    Prakash Narayanan: Amazing.

    1:19:16

    Nathan Labenz: Incredible stuff. Congratulations on the round and the early success. I'm sure there'll be a lot more to come. I don't envy our editing agent who's going to have to try to find highlights out of this section; it's going to be a question of what it will cut. This was action-packed, so I really appreciate it.

    1:19:33

    Thomas Sohmers: And you've got a great guest coming up next.

    1:19:35

    Nathan Labenz: Indeed.

    1:19:36

    Prakash Narayanan: Indeed.

    1:19:37

    Nathan Labenz: Alright. Thanks for being here. Great meeting you.

    1:19:40

    Prakash Narayanan: Bye bye. Thanks for being here. That was incredible. Let me pull up our next guest.

    • AI token spending overtook payroll at Positron

      0:00 / 0:00
    • AI chip memory: Positron's eight-GPU capacity plan

      0:00 / 0:00
    • AI model looping: Why less memory isn't less compute

      0:00 / 0:00
    • AI chip verification: Agents learn Cadence's toolchain

      0:00 / 0:00
    • TSMC capacity: Why Positron avoids the newest nodes

      0:00 / 0:00
  3. 31:10Interview17 min
    swyx: AI Engineer New York, finance, and managing coding agentsShawn Wang (swyx)The AI Engineer cofounder argues finance is the next vertical after code, says engineers who manage agents well are in demand, and discusses token bandwidth, his Kill My SaaS bounty, and OpenAI’s Decisions API.
    Open segment on YouTube ↗

    The segment opened with Prakash Narayanan introducing swyx (Shawn Wang), cofounder and CEO of AI Engineer, cofounder and editor of Latent Space, and adviser to Cognition, and plugging the October 12th-14th AI Engineer New York conference. swyx, after a brief audio dropout he blamed on Codex hogging his machine, said the event is the third New York edition and the first focused on finance. He argued finance is the next vertical to break out after code, citing his own past as a sell-side options trader and hedge-fund quant, and described it as knowledge work in its most verifiable form.

    Nathan Labenz asked how working software engineers feel about the latest coding models. swyx said college students are nervous but that engineers who can manage coding agents are in more demand than ever, while describing employees under performance review for submitting unexamined AI output, which he called paying for someone else's "LLM psychosis." He said taste, depth of insight and understanding at the module level matter more than years of experience, except for security and backend scalability, and offered a race-condition bug from two agents' code as a cautionary example.

    Prakash asked which parts of the stack are poorly suited to agents. swyx pointed to compute, especially CPU shortages, resumable long-running execution such as Temporal and Restate, networking latency, and per-person and per-company token bandwidth. Pressed on managing token spend at scale, he said he has no special solution, but argued large companies manage by team-level goals and cost metrics, citing Amazon's business-unit GMs and Jensen Huang's 60 direct reports.

    On his Kill My SaaS bounty, swyx said it began as a way to replace an overpriced SaaS and drew so many submissions that evaluating them became the bottleneck. He said his event team, initially skeptical of vibe coding, converted after seeing the results and now uses Devin to make changes in an hour or two. Asked by Nathan how this squares with strong demand for engineers, he argued the software pie is expanding as spreadsheets, calls and emails become custom software, with total silicon supply as the true limit.

    The conversation closed on JEV and OpenAI's decisions API. swyx said he is friends with JEV's creator, that its Latent Space episode is the show's biggest ever, and that many use cases are overdone, since it is a System 1 model meant to be as unremarkable as an if statement. After swyx repeated the conference details and signed off, the hosts reflected: Prakash on a management gap between mid-size and very large companies that AI agents might fill, and Nathan, with Gemini-sourced estimates, on Cursor's very high revenue per employee and the possible effect of "log in with OpenAI" on token economics.

    Timestamp links open the original source recording.

    The phrase I've recently taken to is: I don't want to pay for someone else to go through LLM psychosis. I can pay for my own psychosis, that's fine. But when you work for me and I'm paying for your tokens, you'd better be producing thoughtful stuff.

    You are a human using these tools. You cannot let these tools think for you. You have to actually think to use these tools.

    The real limit on all this is the sum total of humanity's token bandwidth, which is literally the amount of silicon we can produce.

    1:21:32Which big trends will AI Engineer feature, and why finance for the New York event?
    The conference runs four times a year, and New York is its third there, now focused on finance. swyx argued finance is the next vertical breaking out after code (labs are prioritizing finance plugins), and he brings his own early career as a sell-side options trader and buy-side quant.
    1:25:24What is the work product in AI finance: code, spreadsheets, presentations?
    swyx said spreadsheets are very much part of it, but finance spans back, middle and front office, insurance, payments, banks and hedge funds. He framed it as knowledge work in the most verifiable form, so AI naturally proceeds there after code.
    1:26:56What's the vibe among software engineers: empowered or threatened?
    swyx said college kids are a bit worried, but engineers who are plugged in and good with AI tooling are in more demand than ever, since the cost of software fell and demand rose. The scarce skill is managing coding agents productively rather than producing slop.
    1:29:48What is the gap between slop and thoughtful work with these tools?
    swyx said it is depth of insight versus breadth of coverage: people let tools think for them and accept false precision (his YouTube analyst's 14.7% figure). People with taste notice when someone has no real relationship to the content.
    1:33:03How much does a software-engineering background matter for that taste and judgment?
    swyx said not much, except for security and backend scalability roles. He argued long experience can hurt, and that data-oriented people who can turn logs and traces into evals, and who juggle five to ten agents, are more valuable; what matters is making modules understandable, not reviewing every line.
    1:37:40Which parts of the stack are poorly designed for agents, given agents overuse primitives like databases?
    swyx said compute, especially CPUs, is the most consistently reported shortage. He pointed to resumable long-running orchestration (Temporal, Restate), latency and edge-versus-cloud tradeoffs, and per-person and per-company token bandwidth (roughly 2,000 tokens per second for him, a rough guess of 2 trillion for Cognition).
    1:41:13How do organizations manage token spend as managers can no longer evaluate everyone's output?
    swyx said he has no special solution but sees hiring capped at roughly 100 to 400 people a year at leading companies. At very large companies you manage by team goals and cost metrics, treating teams as black boxes, as Amazon's GMs and Jensen Huang's 60 direct reports illustrate.
    1:45:41What was the Kill My SaaS project and what did you learn?
    swyx described it as a bounty on mid-tier SaaS that should not exist, started when he balked at a roughly $40,000 subscription. Lessons: evaluating many submissions is hard (evals are the bottleneck), large bounties draw low-quality vibe-coded entries, and his event team switched to the winning approach because Devin lets them request changes in an hour or two.
    1:52:04How do you square strong demand for vibe coders with likely business failures and big-tech layoffs?
    swyx argued the software pie is expanding, as spreadsheets, calls and emails become custom software and people customize like they do handbags. The real limit, in his view, is humanity's total token bandwidth, i.e. silicon supply.
    1:56:50What do you think of the JEV release and the decisions API; has JEV reduced token spend?
    swyx said he knows its creator, Diogo [unclear], ran the first podcast on it (Latent Space's top episode ever), and is in OpenAI's decisions API trial. He called it a good idea but said many use cases are overdone: it is a System 1 model meant to be as unremarkable as an if statement, and many people are using it where any small classifier would do.
    2:00:56Where and when is the event and who should come?
    swyx said AI Engineer New York is October 12-14, tickets should sell out by the end of the week, Latent Space subscribers have a code, and details are at ai.engineer/nyc.
    Lightly edited · timestamps jump to YouTube
    1:19:56

    Nathan Labenz: Graphics package.

    1:20:03

    Prakash Narayanan: Our next guest is swyx, or Shawn Wang. He is the cofounder and CEO of AI Engineer, which runs conferences and workshops for people building software with AI. He's also cofounder and editor of Latent Space, a publication and podcast about AI and the people developing it, and an adviser to Cognition, the company behind the coding agent Devin. He also founded Smol AI, whose AI news project combines automated reading and summarization with human selection. The next AI Engineer New York conference is scheduled for October 12th through 14th, with a focus on applications across banking, investing, insurance, and financial technology. Let's bring up swyx.

    1:20:56

    Nathan Labenz: Heyo.

    1:20:58

    swyx: Hey, guys.

    1:20:59

    Nathan Labenz: How are you?

    1:21:00

    swyx: Can you hear me?

    1:21:01

    Prakash Narayanan: Yep. Yes, you're good.

    1:21:03

    swyx: Okay, alright. Thanks for having me back on. I was just saying I hung out with Tom for two hours at DevDay, and he's such a fountain of advice. He's very low key, very humble, but he's building a monster chip company in record time, at speeds that no other company of his size and cohort has done. So definitely want to keep close to him.

    1:21:32

    Nathan Labenz: Yeah, it helps to be a galaxy-brain genius, I guess. I was just joking: he's one of the few people who can still credibly claim to be able to do something that models might not be able to do, at least for now. But even he, it sounds like, doesn't expect to hold on to that for all that much longer. So you've got another AI Engineer coming up. What are the latest trends you're watching? I know you're the benevolent dictator of the community and always scouting the horizon for new topics to get ahead of. What have you decided are the big things that you'll be featuring at the upcoming event?

    1:22:14

    swyx: So first of all, both of you are invited. I keep trying to invite Nathan specifically, and he's always on some family trip. But at some point, AI will be more important than family. We do it four times a year. The idea is that a conference that happens once a year is horrible, because AI proceeds at superhuman speed. The next one is New York, our New York flagship. This is our third New York conference, and we're coming back with a focus on finance this time, which I'm pretty excited by, but it's also somewhat of a risk for us, because we're very much on home turf talking about code and coding agents. But

    1:22:59

    I definitely think that, as far as the other verticals that are breaking out like code, finance is the next. You can clearly see it from the labs. We have all prioritized finance plugins and finance verticals for our forward-deployed-engineer operations, all that good stuff. But I have a particular affinity for this one because I used to work in [unclear]. I see it in my opening address that I'm drafting right now: it has basically taken the last 20 years of my life to build up to this, because I spent the first seven to eight years, first of all, I went to college, at work. [unclear]

    1:23:39

    Prakash Narayanan: Oh, you cut out. Come back, swyx. Come back, swyx.

    1:23:49

    Nathan Labenz: So everything's working great today except our guest's hardware, it seems.

    1:23:59

    Prakash Narayanan: Oh,

    1:24:00

    Nathan Labenz: you're back.

    1:24:01

    swyx: I think I got kicked out by Codex, which didn't like that I was chatting. So that's a Codex alignment fail. I just have stuff running. By the way, I'm like a super token billionaire now, because I always have stuff running in the background and operating my [unclear]. That was the first time it actually shut me down during a call. Oh, and then they kicked out Prakash. Is he here? Can you still hear me? Yep. You guys have your weird [unclear] platform. Anyway, sorry. Where was I? Yeah, so finance.

    1:24:45

    I spent the last 20 years building up to this, mostly because I spent the first half of my career in finance. I was a sell-side options trader, trading interest rate and currency derivatives. Then I switched to the buy side, in hedge funds, doing basically quantitative portfolio management. And then I switched to tech. So for people who only know me for my tech career, there's a whole finance phase. In some ways I'm one of the few people in the world qualified to bring AI and finance together, and it just so happens that I think this is also genuinely the next breakout vertical that people should be watching.

    1:25:24

    Prakash Narayanan: When we talk about AI finance, with AI code the work product is obviously the code itself, the functionality behind it. When you talk about AI finance, what is the work product that people are looking at? Is it code? Is it spreadsheets? Is it presentations?

    1:25:47

    swyx: Spreadsheets, actually. I considered having an entire track that is only about Excel: automating Excel and building models. Spreadsheets are very much part of it. But in the same way that code has a development life cycle, with the CI/CD side, the app side, the testing side, finance has back office, middle office, front office. There's insurance. There's payments. There's commercial and consumer banks, investment banks, and then hedge funds. And there are information services providers that provide horizontal services for all these players.

    1:26:33

    So it's a little bit hard to pin down. Basically, this is knowledge work, but the most verifiable form of knowledge work. So it's very logical that AI started at code, which is very, very verifiable, and then you proceed on to the next most verifiable domains.

    1:26:56

    Nathan Labenz: You guys are doing so much on Latent Space right now, which I'm enjoying. You're moving into science, which has always been a passion of mine, AI for science generally, so I'm glad you've joined into the AI-for-science coverage. When it comes to just engineering: a couple of years ago, all the engineers were figuring out how to use these coding tools, and of course they've gotten better and better. With just the last couple of models, I have felt like another phase change, where it's so effortless to create so many things that I really am bottlenecked on my ideas.

    1:27:41

    I'm getting things that basically just work the first time, or if there's a little missing piece, it's one quick correction to fix it. What is the vibe you're getting from people who have been software engineers building software applications? Are they still feeling super empowered? Are they starting to feel threatened? For a while people were saying there's going to be so much demand for more software that it'll be fine, there'll be even more software engineers. What's the vibe among those who are most in the know about how this is actually playing out?

    1:28:18

    swyx: I think college kids are a bit worried. But other than that, if you're relatively plugged in and very capable with AI engineering tooling, you are in more demand than you've ever been, because your expected value is higher than it's ever been. This is one of those paradox-type things: the cost of creating software has gone down, therefore the demand has increased a lot. In particular, there's demand for people who can manage coding agents productively instead of producing a whole bunch of slop. And I have been in that situation. I am both an engineer and an employer of engineers

    1:29:03

    and of people who are not engineers who are vibe coding. The phrase I've recently taken to is: I don't want to pay for someone else to go through LLM psychosis. I can pay for my own psychosis, that's fine. But when you work for me and I'm paying for your tokens, you'd better be producing thoughtful stuff. I have two or three employees right now who are basically under performance review because they are just giving me slop. And that's really bad for them. They don't see it. They're like, what do you mean? I think this is perfectly fine. And then, well, you're not producing any value. I can just prompt Claude; I don't need you. So there's that perspective. Most software

    1:29:48

    engineers there

    1:29:49

    Nathan Labenz: Can you unpack that gap a little more? I'm probably also mostly producing Claude slop.

    1:29:54

    Prakash Narayanan: What's

    1:29:55

    Nathan Labenz: the delta?

    1:29:58

    swyx: Going for depth of insight rather than breadth of coverage. My YouTube guy will say, oh, your videos went up 14.7% week on week. And I say, why? And he doesn't know, because he's never actually watched the videos. He just used Claude to analyze them, and Claude doesn't have video analysis or any contextual understanding. And I'm like, dude, you are a human using these tools. You cannot let these tools think for you. You have to actually think to use these tools. Because the tools simulate thinking and provide so

    1:30:43

    much false precision, who gives a shit about 14.7%? What did I do right? What can I do better? And clearly Claude or GPT is not tuned for most specialist tasks like "help me optimize my YouTube video." That's not something Anthropic really cares about, and why should they? There are lots of areas for human domain expertise, but people substitute the lowest-energy, lowest-cost means of production, which is to chuck something into Claude and hope I don't notice. I had an experience with an agency. You know how these launch videos come up, and the agency says, oh yeah, we made that? I reached out to one of those agencies and said, okay, I want to work with you. They said, give me your requirements, and they came back to me with very clear Claude stuff. Clearly

    1:31:28

    maybe they have some talented humans in there who serve their high-priority clients, but the rest of the customers they're serving just get Claude slop. And I happen to have high enough standards that I thought, first of all, this really doesn't get what we're going for. To some extent, slop is fine, as long as it's useful and good, but this was clearly not the case for this particular project I was trying to hire for. They were just like, oh yeah, you caught us, we're going to move on. So they're just hunting for people with low standards, or who don't care, or who have enough money; they don't know anything about the business and are just taking a shot. Whereas the people with taste care about the end product and will have an immediate reaction: clearly you have no relationship whatsoever with this core content, and you're just doing Claude stuff. And I think that's the best. You guys have a clipper, right? It's an agent clipper, and I'm sure it takes a lot of tuning to get somewhere. But probably, if you put your full human attention on clipping, you'd do better. It's

    1:32:58

    hard to articulate what that gap is, but you know it.

    1:33:03

    Nathan Labenz: So this whole taste question: to what degree is it coupled or correlated with a background in software engineering? I do feel you're absolutely right, at least for now, and long may it continue, that there's a role for the human to exercise taste and judgment and push for depth of insight. But if you were in the market to replace a couple of people because they're just giving you Claude slop, how much would you need them to have experience building software to do that job, versus just

    1:33:48

    product sense or high standards in general? I'm not reading any code. Most people I'm talking to aren't reading any code. There's probably still some stuff around hyperscaling platforms that I don't know the agents can do super well. But do you care about the part of the resume that says I've been coding for X years, or is it not so relevant?

    1:34:16

    swyx: I don't care so much, except for, let's call it, security roles and anything around backend scalability. I really need you to know the differences between GCP and AWS and what I can do on each of those. I just recorded an interview with the Supabase founders, who are scaling Postgres to a scale we've never seen before, and good luck trying to hire somebody to scale Postgres with a completely self-managed sharding solution that works at YouTube scale. There's only

    1:35:01

    one person in the world who can do that, and they hired the guy. You'd better make sure you don't lose data, don't cause downtime, and scale economically, because basically agents consume databases at something like a 500-to-1 ratio compared with human developers. Other than that, you can mostly buy code, and then it is about taste. And it's not so much about length of experience; sometimes length of experience works against you, because you have a set way of doing things and you don't understand how to run more than one agent at a time. At this point you should be relatively comfortable juggling five to ten ongoing

    1:35:46

    things, and you're definitely much more of a manager than an individual contributor. I do think that people who look at data rather than code are more valuable these days. Being able to say, here are the logs, here are the traces, here's the schema, here's the input and output, and capturing that and turning it into an eval. All of this is the merging of AI engineering and ML engineering that people need to upskill on, and then managing runs of that at scale. I think the person with 10 or 20 years of software engineering experience really reviews every line and tries

    1:36:31

    to make every line make sense. Whereas now we just need to make the modules make sense. I can allow slop in there because it helps me go faster, as long as I contain the slop in things where I completely understand the whole system. Where you go wrong is when you have too many modules, too many black boxes you don't understand, and then the code also gets confused. As an example, I'm making my own sort of Slack competitor, and I saw a bug where the messages weren't loading. I refreshed and the messages loaded, I refreshed again and they didn't. It's the same exact code; what the hell is going on? It turns out there were two code

    1:37:16

    paths and a race condition. Why? Because two different coding agents worked on it at different times and each made their own thing. A human would never do that. An agent would potentially sometimes do it, because sometimes things fall out of the context window. But you need to have oversight of the module, and whatever is inside the module can be a black box.

    1:37:40

    Prakash Narayanan: So can you describe this: we're obviously seeing a transition from humans being the primary users of the Internet to agents being the primary users. And as you pointed out, 500-to-1 on things like databases. What are the other parts of the stack that need to be scaled at this point? It seems like many of the primitives are going to end up being overused. Which parts of the stack are really poorly designed for agents right now?

    1:38:20

    swyx: Compute. We covered this in January, February, when the GPU shortage was spilling over into a compute shortage. Everyone knows the GPU shortage, everyone knows the memory shortage, but the CPU shortage is the most consistently reported thing among all the people I interview right now. What does that mean? It means we need more fluid compute, which is the first [unclear] term for it, and different forms of serverlessness that are very ephemeral, where you can execute and resume, because agents need some time to do LLM calls or network calls, whatever.

    1:39:06

    Basically, long-running states that are resumable. I used to work on this at Temporal, by the way. There's existing technology doing it super well, which is why Temporal is doing super well; they've been in this game the longest. There are other competitors, like Restate, that also work on this orchestration problem. Beyond that, I'm really thinking about two things. One is bandwidth or networking, where latency starts to really matter. That's an optimization of which point of presence you're communicating with, or which region you're in, versus how much compute you should be doing in the cloud versus locally, at the edge, on devices.

    1:39:52

    That is definitely happening a lot with the robotics companies we're talking to. And don't treat it as just a robotics problem; robotics is the first harbinger of what you'll all eventually be doing if you consume enough inference at their scale. So there's all that. And then finally, on a very personal basis, bandwidth: how many tokens per second can I get, both on a single-call basis and in aggregate across all my calls, and then in aggregate across my company, because I have enough people managing all these calls. Right now, my tokens per second personally

    1:40:37

    as an individual, across all the agents I'm managing, is on the order of 2,000, and Cognition is roughly 2,000,000,000,000. That's a rough estimate, a [unclear] number; don't treat this as a real indication of their spend. It's in a rough ballpark. And I think for a scaled AI company, that's probably your overall total bandwidth of productive AI usage, summed across agents and humans.

    1:41:13

    Prakash Narayanan: Let's talk a little about that. You've already said you've had issues with poorly spent tokens, parts of your organization that perhaps didn't have the critical faculty to evaluate the tokens being emitted with sufficient thinking. So how does that work as you start to scale organizations and a human manager has less ability to evaluate whether their subordinates are spending tokens properly?

    1:41:56

    swyx: Wait, so, yes, you need the ability, but are you asking if I have a solution?

    1:42:04

    Prakash Narayanan: Yeah. How do you think that would work? Would organizations be capped in terms of the amount of manpower they can successfully use, because you need people to monitor the humans who are using the tokens?

    1:42:24

    swyx: I think this comes down to management philosophy in general, and I don't have any core competitive advantage there. I do talk with a lot of CEOs who are going through this problem. On one hand, yes: the leading companies, the names you've all heard of, have a practical hiring limit, because they just don't have enough managers to manage the people they're hiring. It's still pretty good, on the order of 100 to 400 people a year. But if you asked them to onboard 4,000, they would struggle.

    1:43:06

    Prakash Narayanan: Right.

    1:43:07

    swyx: I think the other perspective is this. That is normal, let's call it seed to Series D, Series E hiring. At the large, more-than-10,000-person company level, nobody cares about the individual humans inside a team anyway. You manage by numbers: what are your three top goals per team, and are you delivering those so I can plug you into the rest of my org? Inside that black box, I don't care. I'll hold you to some cost-efficiency metric, but hire the humans you want, spend the agents you want. And that

    1:43:51

    actually is much more scalable. That's the bitter lesson for organizations: you don't need to know every single person in your company. You just need to know your team leads, and I'm going to fire the entire team if the team lead is not performing. You're managing Walmart, you're managing Johnson & Johnson, Procter & Gamble; there's just no way to do it otherwise. It's not very humane, but people have spent centuries building up the disciplines for this. I don't think anyone is particularly inspired by how you manage a 200-year-old legacy

    1:44:36

    Fortune 500, but this is what people do, and I think that's perfectly fine. They have business managers, P&Ls, and other objectives they optimize for. Amazon was very like this: you have business units, and every business unit has a GM, who is basically a CEO. They cost it, they run it, they're plugged into an overall strategy, and I think it works well. I think probably more companies should try that. That's a bitter-lesson scaling of humans rather than making sure everyone is well attended to. At the very limit, take one of the richest men in the world, Jensen: he doesn't have one-on-ones with any of his direct reports. He has 60 direct reports. That's very contrary to everything I've ever done in my tech career, where the maximum you have is eight. It is just a different type of scaling, and I do think more companies should probably embrace it, but it is a very big cultural change.

    1:45:41

    Prakash Narayanan: Indeed. Let me segue a little bit. You did a Kill My SaaS project.

    1:45:49

    swyx: Mhmm. Not done yet. But yes, it's actually very successful.

    1:45:53

    Prakash Narayanan: So why don't you tell us a little about it? It came about because you were using a product for AI Engineer, and you found they were overbilling you, or you weren't happy with how much you were paying for the amount of work they were doing. And you decided, okay, let's do it over a weekend. I'm going to run this Kill My SaaS project, put up this SaaS there, and encourage people to come up with solutions. Is that a fair summary of how you started?

    1:46:24

    swyx: Yeah. Basically, you can call it a bounty on mid-tier SaaS that should not exist. And this was definitely one of them: half a salary, or a third of a salary, for a piece of software that I don't own and that nobody enjoys using. Let's just throw that out. If I'm going to spend $40,000 on this SaaS subscription, I can spend that on tokens, and that buys me a heck of a lot of tokens. That would be the ideal outcome. The only problem is, first of all, for people involved in KMS [unclear] who haven't received the results yet: sorry.

    1:47:10

    We're working on it. Second of all, this is the problem with launching without planning. I just did it on a vibe, to my entire audience. We had so many submissions that we then had to eval them, and so evals are the problem. We're no longer checking whether you fit two or three requirements. We're checking the entire UX of the entire flow from three different perspectives: organizer, attendee, and sponsor or speaker. Then you need different logins and all the different workflows, to submit different applications across the different formats that we do. It gives me an appreciation for building frontier evals. If you're going to code or manufacture entire programs from scratch, then you need the ability to eval them. I actually shipped evals to help my participants on the second day, but for really frontier things, the human just has to play-test through all of

    1:48:18

    Prakash Narayanan: it.

    1:48:19

    swyx: on top of their regular day jobs.

    1:48:20

    So yeah,

    1:48:22

    that is the tough finding. And be aware that if you offer a large bounty, like $10,000 as the prize, you get a lot of submissions, because people want to vibe code for fun, and a lot of them will be low quality.

    1:48:37

    Prakash Narayanan: Right.

    1:48:38

    swyx: They'll just say, Claude, make no mistakes, go do this. It'll make mistakes, and they'll just submit. So the verification load is very imbalanced, because they spent zero thought on this thing, they just threw it in there, and you have no idea whether it's low quality or not. That would be my general lesson. But overall, very successful. And initially, I have one of the most adversarial non-AI teams in AI. What is my business? I hire event professionals who are super old school. They do everything in spreadsheets. They work with unions. They have to worry about the physical placement of the meter boards in locations. These are not

    1:49:23

    glamorous, high-tech jobs, and they're very suspicious of anything new and vibe-coded and techy. So I hire these guys and then make them work for AI; that's my job. Initially they were like, we will never use this vibe-code thing, I want to use the tried and tested stuff that has been working for Microsoft and all these other guys. Then they saw the quality of the submissions and looked at their existing platform, and they said, yeah, okay, we're going to switch. You basically need to provide proof that this is feasible. And then the ongoing benefit, one of the benefits of my partnering with Cognition, is that I give all of them access to Devin to modify the code. Anything they don't like, they can request the change and have it in pretty much one to two hours, which they've never had before. Just to give you an idea of the way we work with these SaaS companies right now: when we request a change, they'll say, okay, that sounds pretty cool, it's on our Q3 roadmap.

    1:50:29

    Nathan Labenz: Yep.

    1:50:30

    Prakash Narayanan: Right?

    1:50:31

    Nathan Labenz: Someday.

    1:50:33

    swyx: And we don't have confidence that they'll actually be done in Q3. By the way, we requested something in Q1 and they only just landed it. So SaaS is quite cooked if you're mostly a CRUD app. I would say there are a lot of UX issues left, though. If you look at SWE-bench and think, oh wow, we're at 90 on SWE-bench, you're looking at the wrong thing, my guy. If you haven't actually tried to completely vibe code a SaaS that you use, you don't understand how models are still very bad at all this.

    1:51:09

    Prakash Narayanan: Indeed. It sounds as though you had to set up a SaaS in order to evaluate your Kill My SaaS submissions.

    1:51:17

    swyx: Yeah. We had the thing we wanted to kill, which I tried to keep anonymous for their protection. Then we evaluated each submission against a side-by-side playthrough. That's the beauty of computer use being good now: you can just point and click and screen-capture and then compare, compare, compare. We have a 200-page document that documents every flow, and we make sure each of those is represented in your clone. And that's mostly it. The extra thing we put on ourselves, because this is also a hiring challenge, is that we ask them to put in some human taste: what would you change if you could do things differently? You are allowed to go off script, and why would you not?

    1:52:04

    Nathan Labenz: So can you synthesize this experience with your first statement, that if you're decent with vibe coding, you've never been in higher demand? It seems to me we're headed for a world where, certainly, there's a lot of latent demand for software that hasn't been met yet, but I have a hard time seeing how we don't pick a lot of this low-hanging fruit and end up with a lot of companies going out of business. I'm also thinking about layoffs from, like, Meta; that's another canary in the coal mine. Are we going to see mass layoffs from big tech? How long can the party go on?

    1:52:47

    swyx: Yeah. This is the thing I had to deal with when I was considering joining Cognition, because I said, you guys are at a $2,000,000,000 valuation, you did it in two years, congrats, this is faster than I've ever seen in my career. And now they're at $48,000,000,000. The simple answer is that if you look at it as a fixed pile of software engineering, that is the conclusion you will arrive at. But if you look at it as, my competition or my TAM is spreadsheets, then all spreadsheets made by anyone with any sort of productivity

    1:53:32

    tooling can be turned into custom software, put on rails with a beautiful UX and automations, and then the demand for custom software becomes a lot larger. Beyond that, if it's phone calls, if it's emails that go back and forth, and eventually physical meetings between people, all of that can be incrementally turned into more and more custom software and hardware and models, by the way. We are just growing into the long tail. Why is Salesforce so damn big? Because people can customize Salesforce. But at some point they need to stop paying the $300K-a-year baseline subscription and spend $30K building their own personal CRM, and they're good. In fact they're more than good; they're happier, because they can now modify it to whatever they want. You should have an idea of the sheer diversity of human needs and desires, over and above just ego. There's a baseline ego of, oh, I made my own thing, you have your thing, I have my thing. There's some demand from that. But beyond that, you immediately start customizing: I like the buttons this way, I like the flow this way,

    1:55:03

    whatever. The sheer amount of customization people can do... we don't even have the same amount of customization in our software as we do in our handbags. What the hell? Clothes have been around longer than computers, and we have so much customization in clothes, so much demand for different brands and varieties. We don't really care about benchmarks for clothes; it's just, this is the style that I like. I think software engineering is only done when we have Kate Spade versus Louis Vuitton versus Coach for software. We're not even at the point of looking at features or price. We're just asking, what's the brand feel? I identify more with this brand. That's what truly commodity software is: you start differentiating on substance and you start differentiating on vibe, and there is a lot of vibe right now.

    1:55:48

    Basically, the pie of software is expanding because it's so easy to customize, and therefore it will. We just need to get the infra there and scale inference. The reason I made the point about scaling my personal token bandwidth versus my organization's token bandwidth is that the real limit on all this is the sum total of humanity's token bandwidth, which is literally the amount of

    1:56:33

    silicon we can produce. That's going to dictate pricing and availability and rate limits and all these things. The real players in the room are only focused on that, and the rest of us are just fighting for scraps within the fixed pie that we already have.

    1:56:50

    Prakash Narayanan: Let's talk a little about fixed pie and non-fixed pie. We had this release called JEV, a calibrated decision-making model, and it went viral; a lot of people were playing around with it. OpenAI, during DevDay, introduced a decisions API, which seemed like a JEV-like undertaking. What did you think of the JEV release, and have you used JEV? Has it been useful in reducing token spend within your organization?

    1:57:29

    swyx: Yeah. Not only have I used it, I was friends with Diogo [unclear] for four years before the launch, and I insisted that we be the first podcast he did about JEV, and that is now our number one podcast ever at Latent Space. Secondly, we're also in the trial for the decisions API with OpenAI, so we have a podcast on that as well. That one is not the same thing; it's basically a Luna model with a different inference layer on top of it. Generally, it's a very good idea. It's so good that I wish I'd thought of it first, but Diogo earns the credit there. I think what's also true among the people I talk to is that

    1:58:14

    a lot of the use cases you see on Twitter are probably overdone or dumb, in the sense that you could use any other small model to do exactly this thing, but because JEV is hot right now, you create a bunch of content about JEV. Or the other perspective is that this is another classifier. You could always train classifiers, just not as sexy. Now people are throwing classifiers at things, or using a classifier model where they would previously use an LLM because they didn't know anything other than LLMs. So that's great; it's additional options

    1:58:54

    Nathan Labenz: in

    1:58:55

    swyx: in the toolkit. I think people are missing the whole point of why they don't call it a decision model. They call it a System 1 model, because it's meant to be integrated into your software and go into the background as unremarkable as an if statement. People aren't really doing that. They're like, let me replace every point of structured output with JEV, and they're probably going to run into issues, because they don't care about intelligence. If you notice, most of these use cases aren't evaluating whether it plays very smart or does any planning. It's just, can you make a decision quickly? And you could always make a decision quickly. So there's some nuance there that's maybe being missed. It doesn't matter, because I think the category creation was so successful that it's an overall good thing for the industry.

    1:59:57

    Nathan Labenz: Well, maybe offline, I'll see if you can help me get off the JEV wait list so I can get into it firsthand myself.

    2:00:03

    Prakash Narayanan: Oh, I see. You send them...

    2:00:05

    swyx: Yeah, you've got to send them good memes. But yeah, they are constrained by compute, I will say that.

    2:00:10

    Nathan Labenz: I don't doubt it. I don't know the numbers on that podcast, but I can say as a listener that it's not just one of my favorite Latent Space episodes, it's one of my favorite podcasts truly ever, because his passion and difference of vision for where AI can go, how we can work with it, and what the experience can be like was so invigorating that I wanted to run through a wall when I was done listening. And that's before even using it. Somebody who has clearly put a ton of time and energy into developing his own positive

    2:00:56

    vision, and it was no surprise hearing him how much it resonated with everybody. I know you've got to go; we can let you get back to work. You've got a lot coming up before the event. Why don't you tell us one more time when, where, and who should come, and then we'll let you get on with your day?

    2:01:14

    swyx: It's next

    2:01:15

    Nathan Labenz: week

    2:01:16

    swyx: in New York, October 12th to 14th, and we're about to sell out of tickets, probably by the end of the week. So definitely get your tickets if you want. There's a code for Latent Space subscribers if you can dig in your emails. And it is at ai.engineer/nyc.

    2:01:37

    Nathan Labenz: And if you're not a Latent Space subscriber, what are you doing, folks? Come on, it should already be in your inbox. That one's on you. swyx, thanks for joining us. Always a pleasure to get your take on things. Few people are more plugged in to what's going on in the Silicon Valley development world, and it's always a pleasure to get your perspective. Good luck with the event.

    2:01:59

    swyx: Thank you.

    2:02:00

    Prakash Narayanan: Bye bye.

    2:02:01

    Nathan Labenz: Thanks, man. Alright. I pity our Claude that's going to have to try to cut this down into a highlights episode. I think this is banger content from start to finish today.

    2:02:21

    Prakash Narayanan: We learned quite a bit, actually. There's a real difference between the 100 to 400 people maximum you can add through Series C or D, and the mid-tier gap, around the 2,000-person to 10,000-person company, where you're struggling to manage and grow fast enough and trying to figure out your management techniques. You're in this gray zone where you can't just assign the team a project and expect it to be completed, or fire them, which you can do at a larger firm. And at the same time,

    2:03:06

    you don't have the bandwidth to interact individually with every team and see the internals. So there's really a management gap there, and my guess is that gap gets filled by AI agents. So you jump from 2K to 10K people rather than scaling slowly from 2K to 10K.

    2:03:28

    Nathan Labenz: I don't know. Have we seen any companies that have been super successful lately grow headcount in the way you would have expected if they were blowing up years ago?

    2:03:48

    Prakash Narayanan: I would

    2:03:49

    Nathan Labenz: say we need to put a research agent on this, but it seems like everybody is keeping the squad pretty small. How many people did Cursor have, for example, at acquisition? I don't know, but it wasn't

    2:04:02

    Prakash Narayanan: I think it was less than 1,000 people.

    2:04:05

    Nathan Labenz: Yeah, that'd be my impression too. Let's do a

    2:04:09

    Prakash Narayanan: quick, between several hundred and 1,000, I think. Maybe 300 to 500. If you just look at revenue per employee, the requirements are very high now, in the $5 to $10,000,000 per employee range. And OpenAI is very big right now, definitely thousands of people, maybe close to 10K and maybe more, but the revenue per employee is very high, and the valuation per employee is also very, very high. Anthropic is probably about half the size, or 60% of the size, of OpenAI, and also has very high revenue per employee. Also to note, both firms tend to hire startup founders and CTOs and have them self-motivate.

    2:04:54

    This is also a strategy used by Rippling, the HR and back-office software company. They hired former founders and assigned them entire business lines, because they were absorbing other HR companies business line by business line. They would assign these former founders to lead teams and build out those business lines. I think that's pretty similar to what OpenAI and Anthropic have done: they've hired people who have founded companies in the space before. Several CTOs of SaaS companies were hired to run products that would eventually be competitors to their SaaS firms. So I think that is definitely a strategy people are using.

    2:06:15

    Nathan Labenz: So Gemini, meanwhile, estimates pretty much exactly what you said: 300 to 700 employees for Cursor, and about $10,000,000 in revenue per employee at Cursor. For comparison, Meta and Google have always been extreme standouts in that regard, and they operate around $1 to $1,500,000 in revenue per employee, so almost a full order of magnitude higher at Cursor. Gemini also puts OpenAI and Anthropic in certainly elite territory,

    2:07:00

    but a little lower than Cursor, more like $3 to $5,000,000 in revenue per employee. Obviously a big part of the reason it has to be so high is that a lot of the money goes to tokens. In some cases companies are still even losing money on the tokens they're serving to customers. I know Cursor probably at times has done that; I'm sure at this point they have positive unit economics on most users. But when you're selling tokens as part of your subscription, a huge part of your top line goes to that. One thing, if we'd had another minute with swyx, I would have been very interested to get his take on how people are reacting

    2:07:45

    to "log in with OpenAI," so you can bring your tokens along with you to your SaaS app. That was not the kind of event I was at this weekend, where talk of that kind of thing was happening, but I suspect it's going to be a huge win. I would be surprised if it's not, because it's such a weird environment to be in, to have to make millions of dollars per employee because you're turning around and spending, in some cases, a majority or even going into the red just to power the product with the tokens. It's so much nicer if people can bring their tokens along with them, at least at first. So that may change what

    2:08:30

    is required, I imagine, a bit. But yeah, these are heady times, to say the least.

    2:08:43

    Prakash Narayanan: Indeed. Let's maybe

    • AI slop at work: Why swyx put staff on performance review

      0:00 / 0:00
    • AI agents face CPU shortages: Why pause-and-resume matters

      0:00 / 0:00
    • AI coding agents vs SaaS: Changes in hours, not quarters

      0:00 / 0:00
    • AI coding demand: Spreadsheets become custom software

      0:00 / 0:00
    • AI coding agents caused a race condition in his app

      0:00 / 0:00
  4. 47:59Closing12 min
    A 3SUM breakthrough, the NanoGPT speedrun, and AI for scienceThe hosts discuss Claude’s reported role in a new 3SUM result, a large NanoGPT speedrun record, how much AI multiplies their output, and the prospect of AI-driven paper retractions.
    Open segment on YouTube ↗

    Prakash Narayanan opened the closing stretch with the weekend's wave of math results, focusing on a new 3SUM result. He explained that for 20 to 30 years no one had gotten the problem below n squared operations, and that the new authors reached roughly n^1.99992 with numbers too large to matter practically. In his telling, the barrier was broken when Claude, working on an Anthropic employee's cryptography problem that assumed n squared was the limit, undermined that assumption as an aside.

    Narayanan added that Anthropic paid two leading mathematicians to reformulate and publish the solution as first and second authors, with Claude and Anthropic's role noted on page 63. He framed it as a preview of a workflow where models solve and humans verify and take credit, amid ongoing unease in the math community, and noted that 20 or 30 further results dropped on Monday after people spent the weekend using up their tokens.

    Nathan Labenz said he didn't know the problem well but pointed to the four-minute-mile effect of barrier-busting, then turned to the NanoGPT speedrun, where an accepted record cut almost half the training time, more than the last 45 improvements combined, with faster but contested claims following. He relayed that the author credited human insight more than AI, that the company reportedly holds back a better internal optimizer, and Narayanan identified it as Hyperstition and wondered why it hasn't been acquired.

    Labenz then revisited a question he heard from Ajeya in May: how many copies of himself, or how much faster, would he need to run to match his AI-assisted output? He said his answer moved from 2x to perhaps 4 to 5x, and far higher for software work, while Narayanan guessed around 10. Narayanan argued that people build permanent "sensory organs" and translation layers, like Zvi's AI News, that displace work for good and act as a capital asset, so the true multiplier is higher than token counts suggest.

    Labenz reported from The Curve that frontier-company people expect AI for science to be next, with one person saying the best science would be impossible without AI within a year. The hosts discussed the likelihood of a "mega retraction" as models flag junk or fraudulent papers, with Narayanan citing the Stanford president's resignation over image fraud and predicting strong social pushback. Labenz also relayed a secondhand account that doctors and the AMA have been receptive to AI, which Narayanan attributed to doctors being overwhelmed since Google and helped by tools like OpenEvidence, before Labenz closed the show.

    Timestamp links open the original source recording.

    For 20, 30 years now the answer has been no, and now that you have techniques that get you below n squared, you'll find other techniques.

    It's not just the number of tokens that you spend on a daily basis, but the fact that you've spent those tokens and that becomes a kind of CapEx, a capital asset you've built for yourself.

    In no more than a year, it will not be possible to do the very best science without AIs playing a big role in it.

    Lightly edited · timestamps jump to YouTube
    2:08:46

    Prakash Narayanan: Let me segue a little bit to other news. Over the weekend we've had a series of math discoveries. I still haven't caught up; it's an unending series, with major conjectures falling left and right. One that's perhaps worth discussing is 3SUM, and obviously people are making a bunch of jokes about the name. It's an interesting story. If you take three numbers and sum them up, can you get to zero or not? That's all it is, and it's basically a search operation over the space of numbers. So 2, minus 10, and plus 8 equals zero is a 3SUM solution.

    2:09:32

    Prakash Narayanan: The question was whether you can get the number of operations below n squared. For 20 or 30 years the answer has been no, and all of these people have been trying to prove otherwise without getting below n squared.

    2:10:17

    Prakash Narayanan: The authors who published yesterday or today managed to get the exponent down to about n^1.99992, by finding a combination of numbers in the number space that brings the search time under n squared. That's significant because it breaks the barrier, and once a barrier has been broken, other people will find further optimizations to reduce the number. For the moment there's no real impact, because the numbers these authors found are extraordinarily large, so it's not a big deal. But now that you have techniques that get you below n squared, you'll find other techniques, and in the end this will help matrix multiplication and other very basic mathematical operations. Summing three numbers is about as basic as it gets.

    2:11:02

    Prakash Narayanan: The way it was found: an Anthropic employee gave a cryptography problem to Claude and asked Claude to solve it. That cryptography problem relied on n squared being the limit of how quickly you could do something. While working on it, Claude broke that assumption and came back and said, in effect, you asked me to solve the [unclear] problem, and I kind of just solved it by undermining the entire basis for that set of problems. On the one hand, it's yet another case of finding a single value that disproves the entire theorem. On the other hand, it's absolutely amazing that Claude could do this at all, and did it as an aside, really.

    2:12:34

    Prakash Narayanan: A number of things will happen over the next few weeks and months. This is somewhat related to P equals NP. 3SUM is a P problem, where answers are quick to find, and NP is where answers are quick to check. Right now we assume these two regions are not connected, but this is starting to go into that territory of asking which of these problems are really unsolvable, or faster to solve than we assumed, and which are essentially solved.

    2:13:19

    Prakash Narayanan: The other interesting thing about this problem is that after Claude solved it, Anthropic approached two mathematicians who are leaders in this field, paid them to publish, and gave them the solution. They reformulated it in a readable manner and published under their own names, listed as first and second author. Claude is listed, and the fact that Anthropic gave them the solution is noted on page 63. This is being claimed as the way things are going to be in the future: models solve these things, we hand the solutions to mathematicians to verify, and the mathematicians who verify or explain them end up taking the credit. There's still a lot of unease in the math community about what the right thing to do is. This was the recommended process from the math committee that was put together a few weeks ago.

    2:14:06

    Prakash Narayanan: But it's notable that we come out of the weekend, and on Monday there are something like 20 or 30 math results that just dropped. People spent the weekend using up their tokens, and on Monday they dropped proofs of things that have been standing for 20 or 30 years. So there you go.

    2:15:15

    Nathan Labenz: Never a dull moment. I'd never heard of this problem and I'm still not entirely clear on how meaningful it is. The guidance I got from AIs was that it's not likely to be immediately super consequential, but that barrier-busting effect can at times be super important, like the classic four-minute mile. Once it's demonstrated that somebody can do it, lots of people start doing it. So it will be really interesting to see whether there's a wave of follow-on optimizations that bring the number down substantially further. Something similar seems to be going on right now in the NanoGPT speedrun department.

    2:16:00

    Nathan Labenz: My understanding of the discourse is that it's a little unclear what should count, because people are starting to take approaches that were allowed by the rules, but not necessarily what people expected. So I need to study it more. But headline-wise, it seems we've gone from something like 70 seconds on this classic benchmark, which people have been working on for quite a while, to much less. The idea is to train a small model to reach a certain loss as fast as you can in wall-clock time. This result took almost half the time off,

    2:16:56

    Nathan Labenz: which was more than the last 45 improvements combined. That one has been accepted. Since then a number of other people have come on with additional claims. One was just a few seconds off, one was maybe half the time off again, and there's one that's even under 10 seconds now, but that is still under review and not fully accepted. It again uses some techniques that may get rules-lawyered, or that maybe constitute hacking the rules to a certain degree. One thing I did find quite interesting about the accepted one, which seems to have kicked off this wave of significant drops in time, is that the author said AI didn't play a big part in it.

    2:17:42

    Nathan Labenz: The senior executive leader at the frontier company I mentioned earlier believes the models can probably come up with ideas of this quality, but that it's an elicitation problem, and the author didn't rely on AI super heavily in this case. They're obviously getting lots of help, but in terms of core insight it was mostly attributed to the human, not so much to the AI, on this particular drop. This company also does large-scale training. They made a point in their publication that the optimizer, which accounted for a significant but not majority share of the time savings, is actually not as good as the one they use internally. So they have an even better optimizer improvement in-house, but they're keeping that to themselves because

    2:19:02

    Prakash Narayanan: This company is Hyperstition. I think the NanoGPT record was a form of performance marketing, intended to show the capabilities of their optimizer and their team in general. I'm surprised they're still independent. I'm surprised they haven't been bought yet, given the severe shortage of talent. I'm surprised they haven't received offers or had their staff poached.

    2:19:37

    Nathan Labenz: Well, it could be a licensing deal in their future. Who knows? Wild times. It's been a busy weekend, an intense information firehose that I'm still processing. I took thousands of words of notes at the various sessions, and even after talking to people informally I'd think, I'm going to write down what I heard here to make sure my human brain keeps track of it. Here's a question for you. I heard it from Ajeya in May, and I had occasion to reflect on how my answer might have changed this weekend. In May she asked me and others in an audience: how many copies of yourself would you need to do as much stuff as you are doing today with AI assistance? Maybe at some point we'll have to get into the coordination problems of having lots of clones of oneself. I wasn't thinking of it in terms of stepping on one another's toes, but just another way to think about it: how much faster would you have to run yourself? If you could speed up your own

    2:21:11

    Prakash Narayanan: inference,

    2:21:12

    Nathan Labenz: how much would you have to speed yourself up to get as much output as you are getting with AI enhancement? At that time, I

    2:21:21

    Prakash Narayanan: said

    2:21:22

    Nathan Labenz: 2x: two clones, or twice the speed.

    2:21:26

    Prakash Narayanan: So, meaning, let's say you don't have AI assistance. How many copies of yourself would you need in order to perform the same amount of work?

    2:21:36

    Nathan Labenz: Right, that you are doing with AI assistance. Yeah. So I said 2x then. And this weekend, thinking about it again, I thought it might be more like 4 or 5 now.

    2:21:48

    Prakash Narayanan: Mhmm.

    2:21:49

    Nathan Labenz: But then last night I came home and was asking for all kinds of enhancements to my personal, as swyx says, made-to-measure agent console that I use. In that domain, software in particular, it's way higher. What I did last night with about 12 prompts would easily be weeks' worth of work. So that ratio is honestly hard to calculate, and it throws a bit of a wrench into the thought experiment. Of course, the other answer would be that I wouldn't do that stuff at all, but that's fighting the hypothetical. Even subtracting out the software development and just imagining the content, the sense-making, the writing, giving talks, I still think I'm right now in the 4-to-5 range. That's a finger to the wind. I haven't measured it, and I don't intend to deny myself AI for weeks to try to measure it. But what's your intuition for yourself?

    2:23:07

    Prakash Narayanan: I don't know, maybe 10. The way I look at it, a lot of what we build for ourselves tends to be a couple of different things. We build sensory organs: we want some firehose of data, there's a lot of it, and we want it compressed into the things that are important for us to see. You can find that pattern in many places. For example, Zvi created AI News. He set up pipelines over unstructured data, Discord channels on AI, arXiv, all of these different places, which got filtered and ranked until he could send out these very compressed emails. Those became sensory organs for the entire community to know what's going on.

    2:23:52

    Prakash Narayanan: I've seen so many repeats of this pattern, which I call sensory organs. I also see translation layers. People build a lot of them, translating between, say, a math thing and language. And once those things are built, when you use them they displace a portion of what you would have done, and that displacement is permanent. If you were the kind of person who woke up in the morning and looked through Y Combinator, Hacker News, and four or five or six different pages to keep up on AI, and that used to take 45 minutes of your morning, and all of a sudden AI News compresses it to five minutes, that's a permanent compression. So when you say let's take away the AI agents, you also have to say let's take away all of these affordances we've built for ourselves over the last year or so. And once you take away those affordances, you'd need to replace them with an actual person, kind of.

    2:25:23

    Prakash Narayanan: So I have permanently displaced a lot of work that I used to do, and to live without AI completely, you would need quite a lot of people, because we've replaced lots and lots of what we do. And it's permanent, and getting more and more thorough, with more and more stuff being consumed by the agents. The main thing is that we have built these tools for ourselves. That's AI and software: they allow you to build tools for yourself. Someone like Nathan is building tools for himself continuously, things that will change his organization and alter how the entire organization does things going forward. So I think the numbers are actually much higher than we think. It's not just the number of tokens you spend on a daily basis, but the fact that you've spent those tokens and that becomes a kind of CapEx, a capital asset you've built for yourself. I think that's pretty significant.

    2:26:54

    Nathan Labenz: Buckle up. One other big takeaway we can end on from the weekend at The Curve, in terms of expectations, and this was shared across the frontier companies, and again shows swyx's foresight and, to a degree, our own: AI for science is very much where they think we are headed in the not-distant future. One sentiment was, is AI going to be superhuman at just some things, like the verifiable tasks, or superhuman at everything? The middle position, which I think is very credible, is that we are seeing superhuman performance at anything we care about and are really committed to investing in. That doesn't mean full generalization to every domain. But even in domains thought of as not inherently verifiable, they feel that when they put their focus on it, license whatever data they need to license, apply lots of processing to augment it and create synthetic versions, and put some RL on it too, the whole package is broadly understood to work on essentially any problem they really choose to focus on.

    2:28:26

    Nathan Labenz: So AI for science is coming next. One person went as far as to say that in no more than a year, it will not be possible to do the very best science without AIs playing a big role. We'll see. Maybe somebody will prove they can still do it. Tom from earlier is still designing chips better than the agents can, but he certainly couldn't do it fully, at anything like the speed he's doing it, without them. And the expectation is that this comes to much of science over the next year.

    2:29:23

    Prakash Narayanan: I wonder how that plays out. Math is very nice: one theorem depends on another, so if you disprove an underlying theorem, you disprove everything on top. I wonder to what extent science like physics, biology, or chemistry depends on these kinds of nodes on the graph, where if you disprove one node, an entire theory falls, and then new knowledge can be created. I wonder to what extent you can do that and focus on the more important nodes, which would open up entire new spaces, rather than just doing the usual thing. Because if not, I think you end up in a kind of experiment where, let's be honest, maybe 30% to 50% of all science papers are probably non-replicable and probably trash, and that's what we're training the AIs on.

    2:30:27

    Nathan Labenz: Although I bet that if not yet, then soon, they're doing the same thing with papers that they're doing with their own RL environments. These folks are not always the most socially sensitive, but they're sensitive enough to know it's not in their interest to come out and call out whatever percentage of the scientific community for having published junk. But I would bet they have a growing sense inside the companies of what scientific literature is actually reliable and what is not, and that will continue to get refined as the models get better. I would be very surprised if that's not well underway at this point.

    2:31:16

    Prakash Narayanan: Can you imagine the outrage, the absolute outrage, when the models come out and say, oh yeah, this Nobel Prize-winning theorist was wrong? And the guy is still alive and well and says, what? Can you imagine?

    2:31:34

    Nathan Labenz: Well, it's another one of these jubilee questions that's coming our way sooner or later. There's no doubt you're right that there's a lot of junk science out there. If these AIs get that good at science, which seems very likely, they'll be able to separate the good from the bad, presumably with imperfect reliability. We know Pangram isn't perfect either. We demonstrated that a little with one of our past Gonzo experiments. Looking back at my own work, I found a couple of things where it called me AI writing, and I feel confident that's not a fair description, though I'm not afraid of being Pangrammed, just in point of fact. There were a couple of false positives. That will surely happen in the science realm too, but I would bet it will be accurate enough that people will generally take it at face value when the AI says something is junk. Some people will be defending their position, and it's probably going to be tough for them.

    2:32:19

    Nathan Labenz: At a minimum, we're going to need to be graceful about that process, because who knows why all these different junk things happen. Mostly it's not fraud. There are all kinds of ways people can go wrong in the scientific process that aren't fraud, so I don't think we should take the mindset that it is by default. But watch out for a mega retraction, would be my expectation over the next, can't be that much more than a year. I have a hard time fitting into my worldview a scenario where it wouldn't be possible in the next year.

    2:33:29

    Prakash Narayanan: The president of Stanford was an Alzheimer's researcher, and some kids at the college newspaper took his old research papers and went through them. He had some diagrams and pictures, and they spotted that these were fake, basically. That drove this entire process of him trying to explain why he did it, saying it was someone else's fault and not his, and then he had to quit. He resigned. I think it was [unclear]. And that's just image recognition, not even LLMs or reasoning, just image recognition, post-AlexNet, post-ResNet, applied to these diagrams. People just haven't done it at scale yet. Or the resolution at which papers have been scanned in as PDFs is not high enough to have confidence. But if you had the high-resolution originals... which is also why there's this whole pushback: some of the journals are asking for raw data on submissions, and there's pushback on submitting raw data because people don't want it reanalyzed by someone else, the errors pointed out, and the outcomes disproven.

    2:34:59

    Prakash Narayanan: It's this whole mess. And just as we saw in mathematics, with the whole question of status and who should be allowed to do it, I can imagine some young undergrad who spends a hundred dollars and runs all of biology's papers at scale, and then it says, okay, 30% of these guys are just faking it. So I think there's going to be enormous social pushback in the research community on this, because it's going to affect a lot of careers and livelihoods.

    2:35:42

    Nathan Labenz: Yeah, it's going to be fraught, no doubt about that. One other thing I heard at The Curve this weekend was on the topic of medicine in particular. My experience using AIs for a medical situation is well documented going back a year, and models are of course much better today. Somebody who works in the AI-for-medical-advice domain said, and this is not confidential, just something I heard, that doctors have been surprisingly warm and receptive to AI, much less threatened by it than you might have thought, and much more inclined to embrace it. From my position it sometimes felt slow in the hospital. But even at the level of the AMA, which I had long assumed would come in and try to shut things down that encroached on its turf, the view I got from somebody who has been doing some of that work with those kinds of organizations was: not really.

    2:36:27

    Nathan Labenz: I came away from that conversation feeling good on the doctors, and good on the AMA, if they really are as patient-oriented as it sounds. Of course, the medical field contains multitudes, like everything else. But at a high level, I had expected walls to be put up in an unreasonably protectionist way, and it sounds like, from somebody in the know, much less of that has happened than we might have feared, and the doctors are broadly living up to their oath of putting the patient's well-being first. I thought that was really exciting, and I would hope to see the same from scientists. But it's a little more personal for the scientists. It's one

    2:37:49

    Nathan Labenz: thing

    2:37:50

    Nathan Labenz: when the AIs can come in and be super useful in medicine in general, and it's a whole new reality, but they didn't attack you specifically as a doctor who did wrong. The mega retraction watch is going to touch a lot of individual names, and that is inherently going to be a lot more contentious. But I did come away optimistic, not just about AI for medicine, which I've long been optimistic about, but also that the powers that be will not fight it and might even embrace it. We might actually all get to use it, with their blessing and support, in a way I had always dreamed of, though I was steeling myself against the possibility that it might prove harder. It sounds like it might be smoother sailing in that regard than I had dared to dream.

    2:38:45

    Prakash Narayanan: I wonder to what extent it's because the expectation that doctors be current on all of this new medical technology is so high that they've felt overwhelmed for the last few years, post-Google, basically, unable to keep up with the questions from patients who go in having Googled. And now they have something that works to keep them up to speed very quickly, and accurately too. One of the problems with Google was that the accuracy was very debatable, and things were hard to search because you had to go through all of these links, and it didn't connect to what you already knew. I think OpenEvidence solved a lot of that, and that's been pretty significant for doctors. Yeah.

    2:39:46

    Nathan Labenz: No doubt.

    2:39:48

    Nathan Labenz: I need to go prompt some AIs to make sure I'm still hitting this 5x acceleration. This was a great session today. I really enjoyed the two conversations, and a lot of interesting conversations just between the two of us as well. We'll be back at it tomorrow. We're only doing two days this week, but we'll have another good one tomorrow, and we'll just keep rolling. Thanks for being with us on AI in the AM.

    2:40:17

    Prakash Narayanan: Thank you.