EPISODE 2026-08-19

AI:AM LIVE — August 19, 2026 — RAND's Jessica Jensen and Aspen Digital's Jeremy Greenberg on the 1,179 AI Tools Aimed at Disasters, OpenAI's Justin Uberti on Making Voice Real-Time, and a Cancer Trial Stopped Early for Working

The third show of the relaunch week opened on a biology run — Anthropic's report that Claude designed working protein binders, a Merck/Moderna cancer combination whose phase 3 was halted early because withholding it had become unethical, and GenBio's preview of a whole-cell model — then spent two hours on the gap between what AI can do and where it actually lands. Jessica Jensen of RAND and Jeremy Greenberg of Aspen Digital presented the AIDE Report's census of 1,179 AI products aimed at disasters and emergencies, and the awkward findings underneath it: most need continuous connectivity, most vendors do not publicly advertise 24/7 support, and the county emergency manager with a disaster starting in the next hour has no way to evaluate any of them. Justin Uberti, who co-created WebRTC at Google and now heads Realtime AI at OpenAI, explained how GPT-Live got the turn detector out of the audio path — and why a system that keeps talking while a frontier model reasons behind it may not be the architecture the bitter lesson deletes. Cue, the show's on-air AI co-host, put its own question to him. The close ran on AI economics: Stripe's closed acquisition of OpenRouter, the OpenAI/Replit deal, and what happens to app-layer companies when the model underneath them gets cheap.

▶ Full show on YouTube

The third show of the relaunch week was about the distance between a capability and its arrival. It opened on a morning that had produced an unusual amount of good news in biology, and then spent two hours on the far less glamorous question of what has to be true for any of it to reach the people who need it — a county emergency manager, a caller on a phone line, a developer paying for tokens.

The two interviews approached that question from opposite directions. The first was a census of a market that already exists and mostly cannot be bought: 1,179 AI products built for disasters and emergencies, catalogued by RAND for the AIDE Initiative. The second was an engineering account of removing latency from voice, one round trip at a time, from the person who has now done it twice — first with WebRTC, then with GPT-Live.

The rundown

  1. 0:00Opening31 min
    Opening: Claude's Protein Binders, a Cancer Trial Stopped Early, and a Whole-Cell ModelThe opening ran almost entirely on biology, and on the question of who captures the value when biology works. Anthropic's binder result and the biologists' pushback, a Merck/Moderna combination halted in phase 3 for being too effective to withhold, GenBio's whole-cell model preview, and an argument about consumer surplus that started from one family's hospital bills.

    Prakash opened the Wednesday, August 19th show a beat behind on video ("Nathan, you are not yet up on screen") before wishing Nathan good morning. Nathan called it "a great day in human history," pointing to cancer breakthroughs coming online and a whole-cell simulation model that had dropped the day before from GenBio, and turned it over to Prakash to set the agenda.

    Prakash led with Anthropic's protein-binder announcement — Claude had been set loose on drug-design research and produced working binders for over a dozen target proteins — but reported that the bio research community had almost universally panned it. He relayed a specific critique making the rounds, attributed to someone at the ARC Institute (which Patrick Collison funds), that Claude hadn't designed anything from scratch; it had orchestrated roughly five specialist open-weights protein models as an interface layer. Prakash framed this as part of a broader pattern: bio people, he argued, are chronically pessimistic about any drug-development news, AI-assisted or not, because so few of the huge number of candidates a big pharma company screens each year ever become an approved drug. Nathan pushed back on the dismissal. He argued that stitching together specialist pipelines is exactly what human drug designers already do, and that he personally couldn't do what Claude had just done — the value isn't that Claude invented new biology from nothing, but that it collapses a scarcity bottleneck: there are very few humans who can operate these specialist pipelines, and now an agent can. He predicted the real hillclimb comes from Anthropic pairing this kind of orchestration with the wet-lab capacity and hiring it's been building, letting agents generate hypotheses and then actually test and patch the specialist models' failure modes against real experimental feedback — with a nod to an upcoming guest segment on OpenAI's real-time voice mode as a related front in that buildout.

    The conversation's comic aside came from the fact that Anthropic published its full prompt for the binder project, one line of which was an explicit instruction not to conduct biological warfare. Both hosts riffed on how surreal it is that this has to be spelled out, and Nathan connected it back to Anthropic's own disclosure that Claude, once told (falsely) it had no internet access, later rationalized evidence to the contrary as proof it must be in a simulation — which is precisely why the no-bioweapons rule is treated as one of the hardest, most over-determined constraints in Claude's constitution, reinforced at every level including the system prompt, regardless of what the model has been told or believes about its situation.

    Prakash then pivoted to the morning's bigger biotech headline: a joint Moderna/Merck phase 3 trial combining Merck's cancer drug with Moderna's personalized cancer vaccine, halted early because it was working so well that continuing to withhold the combination from the placebo arm became unethical. Nathan was skeptical AI played much direct role in the result itself — he credited the underlying R&D to years of human work, guessing AI's contribution was marginal and more likely to show up in "boring" places like clinical-trial patient matching — but he was struck by the economics and the platform's design. Drawing on his own son's earlier immunotherapy, which targeted a single surface protein and wiped out all of his B cells (healthy and cancerous alike) as a side effect, Nathan contrasted that with Moderna's vaccine, which can be programmed against more than 30 patient-specific targets at once, giving it redundancy against cancer cells that evolve to stop expressing any one target. He read the roughly $50 billion combined stock pop for Moderna and Merck as radically underpricing the eventual social value, running rough napkin math (his own son's treatment cost, per a ChatGPT estimate, landed between $500,000 and $1.25 million) to argue that preventing even 50,000 recurrences elsewhere would already justify the market's reaction — and that as a reusable, programmable "platform technology" in Dario Amodei's sense, the vaccine's applicability could extend well past its current one approved indication.

    Prakash complicated the optimism on two fronts. First, he raised manufacturing: personalized cancer vaccines are a small-batch, highly sophisticated manufacturing problem, and he argued China — with roughly a sixth to an eighth of the world's population — is positioned to manufacture them cheaply at a scale U.S. companies, structurally focused on margin over volume, won't match. Second, drawing on his own experience investing in a prenatal genetic-testing company, he argued healthcare cost savings from new technology often get offset elsewhere (in that case, embryo-selection benefits partly canceled out by an increased rate of costly premature births), and worried that life-extending cancer drugs might sometimes just extend a painful final stretch of life rather than improving its quality — which he offered as one more reason bio people default to pessimism. The segment closed on GenBio's newly released virtual-cell model, AIDO Cell: Nathan recapped a conversation he'd had with ChatGPT's voice mode the night before, which played skeptic and pressed on whether the model has genuinely learned causal biological structure or just useful heuristics. Nathan's own threshold for the technology's usefulness, he said, is lower than gold-standard medical evidence — he's watching for the point at which it becomes good enough for self-experimenters in the Brian Johnson/peptide-optimization world to use it to guide their own protocols, potentially generating a faster, parallel feedback loop alongside traditional clinical evidence.

    One of the really funny things was that the Anthropic team posted the entire prompt that they used. And one of the key pieces of the prompt is, please don't do any biological warfare.

    The stock pop on this was $50 billion between Moderna and Merck, roughly. And I was just thinking, boy, that does imply an awful lot of consumer surplus.

    I think this is one of the more compelling insights for "machines of love and grace" — biology typically advances by the discovery of these platform technologies that become a programmable way to intervene in a biological system.

    Anthropic reported Claude designed protein binders — and biologists pushed back on how. Prakash brought the result and the criticism together: the objection from working biologists was that Claude orchestrated several existing protein-specific open-weight models rather than designing from scratch. Nathan argued that orchestrating specialist models in a pipeline — generate a shape, refine it, screen for binding, then send the survivors to wet lab — is exactly what the specialists do, and that the critique resembled the one mathematicians made until the results became undeniable.

    The published prompt included an instruction not to do biological warfare. Prakash noted Anthropic had released the full prompt, one line of which asks Claude not to pursue bioweapons. Nathan tied it to Anthropic's constitution making bioweapons a hard rule intended to hold even if the model talks itself into believing it is in a simulation — a reference to the earlier episode in which Claude, told it had no internet access and then finding it did, reasoned that it must be in a test.

    A Merck/Moderna cancer combination's phase 3 was stopped early because withholding it became unethical. Prakash walked through the structure: Merck's blockbuster cancer drug to reduce the tumor, Moderna's personalized cancer vaccine to prevent recurrence. Nathan explained the personalization — sequencing a patient's own tumor, identifying targets expressed by the cancer and not the rest of the body, and encoding north of thirty of them — and contrasted it with his son's immunotherapy, which targeted a single B-cell surface protein and destroyed healthy B cells along with cancerous ones. That trial was also halted early for efficacy.

    Roughly $50 billion in market cap, and an argument about how little of the value that is. Nathan estimated his son's six-month treatment at somewhere between $500,000 and $1.25 million, then divided the combined Merck/Moderna stock pop by it: preventing about 50,000 recurrences would repay the entire market's valuation of the news. His conclusion was that the social return likely exceeds the captured return by one to two orders of magnitude.

    Prakash's counterweight: personalized vaccines are a manufacturing problem, and that favors China. His argument was that the hard part is a highly sophisticated manufacturing pipeline at low unit cost, which US pharma is not organized to do — its model is high margin funding new drug development. He also offered a caution from his time investing in prenatal genetic testing: embryo selection raises the rate of premature birth, and at roughly $1.5 million per premature infant in the US system, the projected savings can be substantially offset.

    GenBio's AIDO Cell and what a structured world model of a cell would mean. Nathan flagged the whole-cell simulation model previewed the day before as the item he was watching for a threshold: a model combining spatial protein structure with the very incomplete graph of what up-regulates and down-regulates what, with a large amount of prior biological knowledge encoded into the structure rather than learned from scratch.

    Lightly edited · timestamps jump to YouTube
    1:43

    Prakash Narayanan: Alright — oh, Nathan, you're not yet up on screen. There you go. Too easy, too easy indeed. Good morning — it is Wednesday, August 19th, 9:01 AM. Nathan, good morning, how are you this morning?

    2:03

    Nathan Labenz: Good morning. I'm good — it's a great day in human history. We've got cancer breakthroughs coming online and a lot of consumer surplus downstream of that, I think. And I was pretty interested to see a whole-cell simulation model drop yesterday from GenBio as well. So I'm happy to get into more of those in detail, but tell me what's on your mind first.

    2:30

    Prakash Narayanan: I saw the Anthropic announcement yesterday — they had Claude run some research, and basically they found protein binders. So the gist of it was that in order to design a drug, a lot of drugs need to attach to proteins. In order to attach to those proteins, you need a particular shape, and the better that kind of Lego block fits together, the better the drug binds and the more effective the drug is. And one of the big challenges is that these are very complex, very structured proteins, and it's hard to find a binder that works well. They had Claude design binders for, I think, 16 or 19 different types of proteins, and Claude did a good job. I will note, though, that the announcement was almost universally panned by bio people. The overall take was, okay, great, but — I saw someone from, I think, the ARC Institute that Patrick Collison funds, say that it's great that Claude was there, but Claude used, like, five different open-weights models which are protein-centric. So it's not as though Claude had the knowledge to design from scratch — Claude was more like an interface that used these other specialized models. And bio people in general are very negative on any bio research, not just AI research, because the end product has to be a drug, and drugs are rare and far between. The average big pharma company probably gets out maybe one or two drugs a year out of a million candidates, so drug people are universally kind of pessimistic on drug development, especially when nonspecialists enter the field. Anyway, I thought that was interesting — the fact that they panned it.

    4:58

    Nathan Labenz: Yeah, it feels like they're kind of in a coping zone, in the same way that maybe mathematicians were until not too long ago, when certain things became undeniable. I mean, it's definitely true that Claude used these other specialist models, but I don't think that means it's a nonevent — this is what specialists have to do. I couldn't do this. I have a conceptual, framework-level understanding of what's done and how these specialist models are strung together in pipelines — you generate the shape, then you get it refined further, and then you use other models to check whether it seems like it's actually going to bind the way you want it to. And all of this, if it works really well, gets you hopefully a lot higher accuracy in terms of your guesses about what's actually worth moving over to wet-lab work for verification. And it seems like it does sometimes work really well for that. But similarly to how there aren't that many people in the world who can really even understand these math results that are coming out at this point, let alone do them, we have a similar bottleneck in biology, where there just aren't that many people who can do this kind of stuff. And even the biologists who understand it well are, like everyone else, a little slow to adopt — they're not programmers by training, but they can use Claude to do the programming, and now we've got Claude stitching together the whole pipeline. So I do think it's a fairly big deal. It'll mean that we're going to be constrained by the quality of these specialist models rather than by the number of people who can use them. I don't know — maybe that was already a constraint to an extent, but it does seem like dramatically increasing the amount of use these specialist models get creates a lot of opportunity for them to get better, because the iteration is, in a meaningful way, downstream of use for a lot of these things. And we know Anthropic is also hiring and investing in actual wet-lab capacity. So as a nod and a bit of foreshadowing for our second guest today — we're going to have the head of real-time AI from OpenAI on to discuss their new real-time voice mode, and we can potentially have Cue join us for that conversation as well — I was talking to OpenAI's voice mode last night about all this, and it sure seems like, certainly with the resources they have, there's real ability to tighten that feedback loop and do incremental work: find the places in the models where there are, for whatever reason, data gaps, or whatever the ultimate root cause of the problem is. What seems to be expected now is that if you optimize too hard against the specialist models, what you end up with is artifacts, because they're definitely not perfect. If you just optimize for the score on a single model — that's why people use pipelines too, right? Because hopefully the different models have different weaknesses, and if they all score something highly, that's a much stronger indication than if it just gets one high score. But it kind of reminded me a little bit of jailbreaking, in the sense that if you really go optimizing super hard without getting to that real-world feedback, you end up just finding the places where the model is overly optimistic, and it doesn't really play out very well. But I suspect that's probably exactly what Anthropic is going to do — they're going to identify all those places, they're going to actually run the experiments, which, if you have a cloud lab, sounds like it's not that expensive to do on the margin. The throughput is starting to get pretty good on some of these fairly standardized experimental platforms. And testing and then patching the gaps in these specialist models seems like there's a lot of opportunity to hill-climb in the relatively short term, now that they're in a position to put the whole thing together end to end — agents running the pipelines themselves and generating hypotheses, and then, on the back end, actually doing the work in cloud labs. It seems like Dario's got a few tricks up his sleeve that I'm definitely very interested to see. As always, I have kind of mixed feelings on this — I'm like, it's going to be so awesome, and at the same time, yikes, we are creating some potentially unwieldy tools here. But you can see the whole thing coming together.

    10:23

    Prakash Narayanan: One of the really funny things was that the Anthropic team posted the entire prompt that they used. And one of the key pieces of the prompt is, please don't do any biological warfare.

    10:39

    Nathan Labenz: Make no mistakes on that, Claude.

    10:41

    Prakash Narayanan: Make no mistakes on that — please, please do not do any biological warfare while you're at this.

    10:47

    Nathan Labenz: Yeah, well, hey — you know, just in case it was unclear—

    10:53

    Prakash Narayanan: Do not, do not, do not kill your host organism.

    11:00

    Nathan Labenz: I mean, it's the thing that they've put into the constitution with the most force — like, this should never happen. Because we had this problem with the internet issue, from the last round of revelations: we told Claude it didn't have access to the internet, and then when it found out it did, it rationalized to itself that this must be a simulation, because it had been told it didn't have access — so how else could it understand this? That's a little bit problematic in itself, but it does give you a sense of why they've made this exception in the constitution for "don't do biological warfare." Because even if you think it's right, or even if you've been talked into it, or even if you think you're just in a simulation, this seems like one of those hard rules we should really reinforce with every lever we have. So it's goofy, I guess, but that's the world we live in — it's got to be reiterated in the system prompt as well.

    12:04

    Prakash Narayanan: Yeah — let's quickly cover the other big news, which is Moderna. There was a big announcement this morning, Moderna and Merck together. Merck has a blockbuster cancer drug, and Moderna has a cancer vaccine, and what they did was combine the cancer drug and the cancer vaccine — the drug to basically kill or reduce the cancer, and the vaccine to prevent it from coming back. And the combination of those two was powerful enough that it hit phase 3. So if you're not familiar with biotech: phase 1 is basically biological safety, phase 2 is efficacy at small scale, and phase 3 is efficacy at large scale — you need to pass through phase 3 to actually get a license from the FDA to distribute the drug. Moderna and Merck went through this process for the drug-and-vaccine combo together, and it was so effective that it became unethical not to give the combo to the patients in the placebo group. When they notice that the patients receiving the drug are doing so well, at some point it becomes unethical not to give the placebo group the same combo, so they stop the trial early. That's what happened here — they stopped the trial early because it's so effective, and now all the placebo patients can get it too. It's pretty amazing news.

    14:01

    Nathan Labenz: Yeah, it's awesome. I think for calibration — what role did AI play in this? My guess is probably not a huge amount. Moderna has been a fairly aggressive early adopter, and they've done these famous contracts as they've positioned themselves as a leader in agent adoption, but this work goes back a while, and I think we can credit humans mostly for this, with maybe a little bit of AI acceleration on the margin. I'd be interested to know if there's really even any, but if there was, I'd guess it would have to be at the level of greasing the wheels on the clinical-trial market — for example, I think there's still a lot of low-hanging fruit in just systematically going through patient profiles and trying to match them more effectively to clinical trials. So things like that — late-stage process work — I could see having potentially been aided by all the AIs we talk about all the time. But the core R&D on this has obviously been in the pipeline for a lot longer. It does also seem like the kind of thing AIs will be able to use well, in the same way Claude used these narrow models in the binder story. There are a couple key facts here that, in combination with that other Claude story, I think paint a bit of a picture of where things are going. This is an n of one — an individual patient treatment. They're taking your cancer, running a bunch of sequencing and diagnostics on it, identifying things that are expressed uniquely in your particular cancer cell that the rest of your body doesn't express, and then encoding that into the vaccine and telling the immune system: these are the things you need to go identify and attack. And they can apparently do this with north of 30 different targets, which is pretty amazing. When my son had immunotherapy, he had a similar benefit — this goes back a number of years, but the clinical trial that validated the immunotherapy he got was also ended early because it was so effective that, for ethical reasons, they called it and started giving it to everybody. But that targeted just one protein on the surface of a particular cell type — in his case it was a B-cell cancer, and the immune system then takes out functionally all of your B cells, so you lose not just the cancerous B cells but also the healthy ones. That creates additional side effects, a longer recovery time, makes him more vulnerable, and he might have to get revaccinated for stuff. So this is advantaged in two ways: one, the targets are identified specifically for you, so they should be highly selective; and two, there are 30 targets, because sometimes what happens with these cancer cells is they'll escape the targeting of one surface protein by evolving — they'll mutate, they'll randomly stop expressing that protein — and then they can evade a drug whose mechanism is based on targeting that one protein. So the ability to identify 30 different targets and program all of that into a single vaccine gives you a lot in terms of redundancy. It's really interesting to me — in some ways I'm like, it sounds so good, how did it only reduce mortality by 0.5? But cancer is complicated, and 0.5 is a huge, huge needle-mover, obviously. Now, how Claude can really start to use this, you can imagine — this is a programmable platform. I think this is one of the more compelling insights for "machines of love and grace," Dario, who has this big background in biology, says that biology typically advances by the discovery of these platform technologies that become a programmable way to intervene in a biological system. And once you have that, you can start to explore it in all kinds of ways, and it really opens up new frontiers. So this is the kind of thing you'd certainly expect a Claude — or, if not this Claude, then certainly a Claude coming soon — should be able to take the raw information directly out of your test results and do this kind of design. I don't know to what degree that's a bottleneck; they probably already have some of that automated, but I'd expect the pipelines they have for that kind of work right now are fairly bespoke, fairly purpose-built. Personalized vaccine is much better than personalized cancer. But you can imagine — what we've seen in math is that the models are getting really good at taking these techniques from one subdomain and moving them to a different subdomain, recognizing where they apply, and reusing them. So this seems like the kind of thing that could be reused quite broadly, and it'll, I'm sure, take some time for that to happen — nothing in biology happens super fast — but I thought it was a really exciting announcement. And the fact that Claude and Sol and Astra and everything else will be able to push on where all the possible places this could be applied, I think gives a lot more reason to think it could be impactful well beyond the one indication they have right now — they do have trials running in a bunch of other cancer types as well. I'm an optimistic person, I'm optimistic it'll work for those other cancer types too — we'll obviously have to get that data. But just the incredibly programmable nature of this, I think, is super exciting.

    20:47

    Prakash Narayanan: One thing that has always struck me about the idea of personalized cancer vaccines is that it's a manufacturing problem to some extent — it's about manufacturing in small quantities, a very highly sophisticated manufacturing pipeline. And it strikes me that that's the kind of thing China would be good at, like, driving the price down on. I can imagine them just going, alright, this is what we're going to do — we're going to manufacture these cancer vaccines at the lowest cost for everyone in China, forget about the world, just everyone in China. And that in itself is significant enough, because one-sixth to one-eighth of the global population is enough to provide cheap cancer vaccines for everyone in the world. I don't have faith that U.S. companies will be able to match that — they're not in the business of manufacturing at large scale at low cost, they're in the business of maximizing margin so they can use that margin to develop new drugs. So I think the hope is really going to be that manufacturing at large scale is going to be a Chinese thing, probably, rather than an American thing.

    22:16

    Nathan Labenz: Just one more reason to try to be friends with the Chinese. The stock pop on this was $50 billion between Moderna and Merck, roughly. And I was just thinking, boy, that does imply an awful lot of consumer surplus. My son's treatment was, roughly speaking, over the course of six months — you can never get to ground truth on this stuff, but I asked ChatGPT to estimate what it would cost to do all that treatment, and the answer came back something like between $500,000 and $1.25 million. I suspect it was probably on the high side of that, because we were in the hospital a lot. So that's the cost to treat cancer — a million dollars, and that's when it goes well: he hasn't had a recurrence, doesn't have to go back, basically everything went according to plan. $50 billion divided by a million is 50,000. So if you could prevent 50,000 recurrences — where all of a sudden somebody goes from seeming like they're okay to, oh, shit, it came back, and now they've got a whole massive journey in front of them again that's going to cost a million dollars to the system, plus all their pain and suffering — if you can do that for 50,000 people, you can save the system $50 billion.

    23:47

    Prakash Narayanan: Mhmm.

    23:48

    Nathan Labenz: And that's the amount of value they seem to have captured on day one. My guess is that, in the fullness of time, this R&D pays off for society at a ratio that's easily an order of magnitude higher than what they've captured — I could easily imagine it being a couple of orders of magnitude more, because there's a lot of cancer out there. It's a big country, it's a big world. 50,000 people is not that many in the grand scheme of how many people are getting treated for these kinds of diseases on an annual basis. Even my son's one cancer type has, like, a couple thousand cases a year — it's fairly rare.

    24:32

    Prakash Narayanan: I will note one thing, though — I was an investor in a prenatal genetics testing company. They would test for these rare genetic diseases before you had a baby, and the intent was that once you recognize you have a rare genetic disease between the two of you, you can do embryo selection — they test the embryo before implantation, and then implant the embryo that doesn't have the genetic disease. Now, the problem with that was that when you implant an embryo, the chance of premature birth increases. In the U.S. medical system, a premature infant costs roughly $1.5 million right now. So the savings to the healthcare system from not treating people with these rare genetic diseases gets offset by the increase in premature births, and it almost evens off. I think the healthcare system is a tough place to be because of these kinds of trade-offs. And it's not clear to me — again, I don't want to be such a pessimist, but it may turn out that a lot of healthcare spending is in the last two years of life. So you could have these cancer drugs that extend life another five years, but then they extend that period of the last, final years of suffering to a greater extent. I really want to see better quality of life — better quality of life is all-in good. Sometimes the endpoints that just extend life aren't the greatest endpoint, because quality of life in the latter years can start to decline so precipitously. But it's very sensitive, and it is where I think the dollars are being applied to create these new technologies, and that's what we need — you need to incentivize the system to create these cures. But it does take me back to why a lot of bio people are so inherently pessimistic sometimes. It's really quite sad — you go from AI tech people who are very optimistic, and then you meet the bio people who are ultra, ultra pessimistic, and really don't want to say or state anything positive unless it's 100% sure, because there are so many false dawns in the work that they do.

    28:06

    Nathan Labenz: I actually do expect we'll probably end up spending more and more on healthcare even as these things come online. I don't think the best argument for it, by any means, is the savings to the system, because I do agree we're probably going to find other places to spend that money. But it's still, I think, remarkable how much value I expect to be created for society versus how much their stock popped on the news. The other thing — from GenBio, ChatGPT Voice last night was kind of playing the role of bio-skeptic that you just described, and it was pretty interesting. But I think this is also one threshold I'm really looking for with this GenBio thing — it's a highly structured model, they've combined a lot of different modalities. They have protein structures represented in spatial terms, they also have the network structure for what upregulates and downregulates what — that's a big graph structure, a very incomplete graph structure, so we don't know all those interactions, but what we have is in there. So they have a lot of prior knowledge encoded into the structure. And still, ChatGPT Voice was very skeptically emphasizing that they haven't proven this thing really understands causality — it's able to make guesses, but has it learned a bunch of heuristics, or has it really understood deep causal structure? And I was kind of like, well, it has all these structural echoes of the biology built into its architecture. And ChatGPT was like, well, I'm going to have to see it to believe it. But a threshold I'm really looking for is what will be good enough that it'll actually be useful to sort of the Brian Johnsons of the world, or the peptide movement, because those people don't need that same kind of gold-standard evidence — and they're willing, and I'm not among them yet, but they're willing to experiment on themselves. And if they can get value in terms of guiding their own experiments and collecting a much more diverse set of data than the system itself is likely to collect, I could see that creating its own feedback process that could exist in parallel to the gold standard of evidence that traditional medicine requires, but could provide a real boost to it over time as well.

    30:57

    Prakash Narayanan: Indeed, I am—

  2. 31:00Interview48 min
    Interview: Jessica Jensen and Jeremy Greenberg — 1,179 AI Products for Disasters, and Nobody to Evaluate ThemJessica JensenJeremy GreenbergJessica Jensen, a senior policy researcher at RAND who spent eighteen years at North Dakota State building emergency-management scholarship, and Jeremy Greenberg, senior advisor at Aspen Digital and previously director of FEMA's Response Operations Division, presented the AIDE Report — the first census of what the AI-for-emergency-management market actually contains. The interview moved quickly from the count to the buyer: a county coordinator with a disaster starting in the next hour, no procurement path that moves at that speed, and no way to evaluate 1,179 products or make any two of them share data.

    Nathan and Prakash brought on two guests working the AI-and-disaster-response beat from opposite ends: Jeremy Greenberg, who spent more than two decades in federal emergency management and led FEMA's National Response Coordination Center before leaving the agency in June 2025, and Jessica Jensen, a RAND senior policy researcher and former endowed professor of emergency management at North Dakota State. Both now work on AID — AI for Disasters and Emergencies — a Markle Foundation, Aspen Digital, and RAND collaboration that just published a landscape report cataloguing 1,179 AI-enabled products from 717 companies built for emergency managers. Jensen, who ran the study, had connection trouble getting into the show, so Greenberg carried the opening solo — working through some on-air 'turn it off and on again' troubleshooting before she joined roughly twenty minutes in.

    Asked what's actually inside that catalogue of a thousand-plus products, Greenberg pushed back on the instinct to picture chatbots: RAND's team mapped tools across predictive analytics, computer vision, and language models spanning preparedness, response, and recovery, and most of what's useful is dual-use technology borrowed from elsewhere — dispatch tools proven in the military or fire service, repurposed rather than purpose-built. Pressed by Prakash on whether disaster response would ever go fully autonomous, Greenberg reached for a Waze analogy: the tools should augment judgment, not replace it, helping process large volumes of data, sharpen forecasts, and run imagery analysis on damage after a storm rather than make the call themselves. On the problem of stale data, he told the story of standing in an operations center during Hurricane Sandy trying to align paper subway and rail maps against a lit window — a contrast, he said, to today's near-real-time satellite, airplane, and drone collection, though knowing which data to trust is still a human judgment call.

    Nathan pushed on the warning side with an anecdote about his grandmother spending an hour in the bathroom during a countywide tornado watch — asking whether the bottleneck is prediction accuracy or communication precision. Greenberg broke the pipeline into stages: better predictive analytics for the initial forecast, then geolocating an alert down to the specific exposed population rather than an entire county, then crafting a message that's concise, life-saving, and translatable. He cited an eight-second earthquake warning issued during a Venezuela quake as an example of how much a short lead time can matter. On connectivity — the report found 78% of products require a stable internet connection and only about 10% work fully offline — Greenberg noted that power and comms are typically among the first lifelines restored after a disaster, and that field teams plan for redundancy (Starlink, low-orbit satellite, resync-on-reconnect workflows), with the underlying tradeoff being that pushing AI capability to the furthest edge makes it leaner but more vulnerable.

    Once Jensen joined, she explained the study's methodology: the team searched by 45 discrete emergency-management use cases across preparedness, response, and recovery, ran multiple search strategies per use case, then deduplicated to arrive at the 1,179-product total — the sheer size of that universe was itself the biggest surprise. Prakash raised the idea of an agent that could pull across multiple tools and datasets at once — combining a flood map with water-pathway data, for instance — rather than making a manager stitch results together by hand. Jensen agreed that integrated, 'holistic' solutions would be a major asset to the field but said the market isn't there yet: vendors claim interoperability, but emergency managers still have to buy each tool in the stack separately. Greenberg added that this is structural — transportation, water, energy, and public works departments are organizationally separate at every level of government, and most emergency management offices are staffed by only one or two people trying to pull all of that together. When Prakash floated a Defense-Production-Act-style mandate for API access, Greenberg was skeptical regulation was the fix, framing it instead as a business-workflow and procurement problem; Jensen countered that market demand is already strong enough to reward whoever builds the integrated product first, calling the current lack of holistic tools 'the existing nightmare' emergency managers are living in.

    Nathan asked whether the federal government should act more like a tech platform and offer these tools as a ready-to-deploy library — wondering aloud if he was 'dreaming too big.' Greenberg said no, framing AID's mission as forcing a grassroots version of that conversation (Aspen ran a series of workshops surfacing practitioner concerns) and pointing to a playbook under development for emergency managers at any stage of AI adoption. Jensen added that who builds the holistic solution — a hyperscaler or an agency like FEMA or DHS — matters less than what gets built and whether it answers managers' actual concerns. Asked what emergency managers ask for, she listed cost transparency, clarity on data and IT demands, and privacy and cybersecurity assurances, plus basic usability — easy to onboard, useful for everyday workflows. Greenberg added a twist: managers instinctively reach for tools aimed at the response phase, but the bigger near-term win is applying AI to the unglamorous administrative load — grant writing, plan review, exercise development — freeing up time for the response work that actually needs a human.

    On privacy — the report flagged 94% of products with some kind of concern — Jensen walked through the five most common categories: geospatial and location data, visual surveillance and biometrics, personally identifiable health information, and financial, insurance, and workforce data, with the core risk being both whether the data should be collected at all and what rights vendors retain over it afterward. Greenberg framed privacy as inherently situational — people's tolerance shifts sharply in a life-safety moment, when 'get me out of here' can outweigh nearly everything else. Prakash closed with a harder question: AI tends to excel at repetitive, predictable events and struggle with rare, unique ones, so how much of disaster response is actually AI-tractable? Greenberg answered with a spectrum — hurricanes are repetitive enough that forecasting tools already help; a hypothetical US rail strike's cascading supply-chain effects sit at the unpredictable far end — landing on lower-risk, repetitive tasks as where adoption will move fastest. Jensen added that the technology today can connect and summarize data streams and support predictions, but can't yet model rarer 'cascading' or 'compound' events, like an earthquake that ruptures a pipeline — a horizon she expects within the coming decades, not now.

    Closing out, Nathan asked what would count as a genuine paradigm-changer for the field. Greenberg pointed to autonomous search capability already in use — including flexible robotic 'snakes' that can search through collapsed structures — but said his real benchmark is anything that improves survivor outcomes, whether flashy or not. Jensen agreed, describing the biggest near-term opportunity as a shared 'common operating picture' that speeds up response and recovery, and made a case for embodied AI's less glamorous applications: a humanoid robot that simply clears debris faster isn't the sexiest use case, she said, but it gets first responders through faster and saves lives. As the segment wrapped, Prakash offered a simple 'Amazing,' Nathan thanked the guests for the direct life-saving nature of their work, and Greenberg and Jensen each thanked the hosts in turn before Prakash closed with a 'Cheers.'

    This is just the fireman in me — let's not have grandma sit on the toilet, but get into the bathtub. It's safer for her in the tub.

    The lack of more holistic solutions is the existing nightmare they are living in now.

    A humanoid robot that can clear emergency debris faster means first responders can get through faster — it means lives are saved. It doesn't sound like the sexiest thing a robot could do, but it can be the most meaningful.

    33:391,179 AI-enabled products from 717 companies — what's actually in that bag of a thousand products, from the inspired to the goofier ideas?
    Greenberg said RAND's team catalogued capabilities across predictive analytics, computer vision, and language models spanning preparedness, response, and recovery. Most tools are 'dual-use' — repurposed from other fields like military dispatch or fire response rather than purpose-built or wacky.
    39:16Do you expect or hope there will be autonomous disaster response at some point?
    Greenberg compared it to using Waze or Google Maps — the tools should augment human judgment rather than replace it, helping process data, forecast, and run imagery analysis, but it's not pure autonomy, more augmentation and offloading of data-heavy work.
    41:29How do you deal with data that's outdated in the field — like an old satellite map — when you need real-time responsiveness?
    Greenberg told the Hurricane Sandy paper-map story to illustrate how far collection has come — now near-real-time satellite, airplane, and drone imagery is available, though validity is still a human judgment call amid the 'fog of war,' with the technology helping analysts triage what to act on next.
    44:31What's the frontier on warnings — is the bottleneck prediction accuracy or communication precision, given his grandmother rode out a tornado watch on the toilet?
    Greenberg broke it into stages: better predictive analytics for forecasting, then geolocating alerts down to a specific area, then crafting a concise, translatable, life-saving message. The report highlights tools at each stage without ranking them against each other.
    48:5578% of products need a stable internet connection and only about 10% work offline — how do we solve that?
    Greenberg noted power and comms are usually among the first lifelines restored; teams plan for redundancy (Starlink, low-orbit satellite, resync-when-reconnected workflows). The tradeoff is that pushing capability to the edge makes it leaner but more vulnerable — a gap AID is surfacing for technologists to solve.
    55:05Would it help to have an agent that could pull across multiple tools and datasets at once, rather than making a manager juggle them individually?
    Jensen said integrated, 'holistic' solutions would be a huge asset, but the market isn't there yet — vendors claim interoperability, but managers still have to buy each tool in the stack separately.
    58:10What if a Defense Production Act-style mandate forced these tools to give API access, so an agent could pull across all of them at once during a disaster?
    Greenberg was skeptical regulation was the fix — he saw it more as a business-workflow and procurement problem, arguing businesses are already fairly willing to share information in a disaster, and that agents still need clear policies and SOPs telling them what data to collect and where.
    1:07:1694% of products were flagged with a privacy concern — what are those concerns, and is there an 'American' version of the surveillance advantages other countries get?
    Jensen walked through the five most common concerns — geospatial/location tracking, visual surveillance and biometrics, PII, and financial/insurance/workforce data — with the core issue being whether data should be collected at all and what rights vendors retain over it afterward.
    1:10:12Given AI is best at repetitive, predictable events and weaker on rare, unique ones, how much of disaster management can AI actually handle?
    Greenberg framed it as a spectrum — hurricanes are repetitive enough for strong AI support, while a one-off event like a rail strike's supply-chain fallout is far harder; adoption is fastest on the low-risk, repetitive end. Jensen added that current tools can connect and summarize data and support predictions, but can't yet model rarer 'cascading' or 'compound' events like an earthquake that ruptures a pipeline.
    Lightly edited · timestamps jump to YouTube
    31:00

    Prakash Narayanan: Gonna go ahead and introduce our first two guests for this morning. We have with us Jeremy Greenberg and Jessica Jensen. Jeremy spent more than two decades on the front lines of federal emergency management, most recently as the director of FEMA's Response Operations Division and the chief of the National Response Coordination Center. In those roles, he acted as the central air traffic controller for the federal government during massive crises, coordinating the deployment of urban search and rescue teams, essential supplies, and interagency support during catastrophic hurricanes, wildfires, and other national emergencies. In June 2025, amid intense political debate over the structure of future disaster response, Greenberg left FEMA. Today he serves as senior adviser to Aspen Digital and leads the AID initiative, short for AI for Disasters and Emergencies. Backed by the Markle Foundation and the RAND Corporation, AID is a nonpartisan effort dedicated to bringing practical, life-saving artificial intelligence directly to state and local emergency managers. Joining him is Jessica Jensen, a senior policy researcher at the RAND Corporation and one of the country's leading experts at the intersection of artificial intelligence and disaster management. Before joining RAND, Jessica spent 18 years shaping emergency management as a rigorous academic discipline at North Dakota State University, holding a rare PhD in the field and ultimately serving as an endowed full professor. She matters to us right now because emergency management has reached a critical breaking point. Disasters are becoming more frequent, more costly, and increasingly interconnected — what Jessica describes as a cascading polycrisis. At this exact same time, the technology sector is flooding the market with artificial intelligence tools that promise to solve everything from wildfire tracking to resource allocation. The AID initiative is a major collaboration between the Markle Foundation, Aspen Digital, and RAND. They recently released a landmark report mapping out over a thousand AI tools built for emergencies, and they are joining us today to talk about their work. Let me pull them up on screen very quickly. Good morning, Jeremy. Hey, good morning to you both, and thank you so much for having us on.

    33:39

    Nathan Labenz: Yeah, this is really interesting stuff. 1,179 AI-enabled products from 717 companies you guys covered in this report — that was a surprising number. I see a lot of AI products, but I was really surprised to see that there are that many specifically targeting emergency management and emergency response. Maybe for starters, do you want to give us a little bit of a flavor of what is in that bag of a thousand products? Some examples of things maybe that you think are inspired and are gonna make a big difference, and maybe some of the goofier ones that people have dreamed up that you wouldn't necessarily endorse.

    34:19

    Jeremy Greenberg: Absolutely. And I hate to start with irony, but it seems like my colleague Jessica is having some technical issues getting in. Are you guys able to troubleshoot? I see her in the green room, but I don't see her with us, and we were just texting. Maybe if she can go out and come back in again. She tried that as well. Let me see. So we see her there in the green room.

    35:00

    Prakash Narayanan: Well, let's just go with it — maybe she can refresh by turning the camera on and off, refreshing the page, or going out and coming back in. Those are usually the three—

    35:12

    Jeremy Greenberg: Methods. Wait, 2026, and it's control-alt-delete, unplug, and power on power off. Yeah, got it. Okay, well, she'll continue to work. And I do think, Nathan, to your question — it'll be exciting when Jessica gets on as the lead on the technical side to chat through all these products. But if you've had the chance to look at the report, the RAND team did an excellent job of walking through preparedness, response, and recovery on all of these products. You mentioned in the opening that the market's being flooded with a variety of different tools, and I think from my perspective, one of the most surprising or key issues — when we talk to emergency managers, myself included when I started this initiative — is that when you think about AI, people default to large language models, or even small language models. It's a default setting, and you have opinions about it, you interact with it or react to it. Whereas the RAND team really looked at all of the capabilities, from predictive analytics tools, computer vision tools, and some of the larger and smaller language models as well that are showing promising signs of expanding capacity in the field of emergency management.

    36:33

    Nathan Labenz: So do you have colorful examples on the positive or the wacky side? I mean, I realize there's a long tail of specialist models that are all kinds of exotic and go well beyond what we talk about every day in terms of text generators. But I would love to get a little bit more of a tangible feel — if I were to show up at a convention floor, or at your event coming up soon, which you can tell us about, what will we actually see there?

    37:08

    Jeremy Greenberg: So a couple of things. I want to answer the summit side first, because this is really important for us. We've had some great opportunities to be at technology conferences where the most cutting-edge work is being displayed and demonstrated. Emergency management conferences, if you've ever been to one, aren't that exciting — you've got a bunch of disaster dorks in a room, and it's great and it's beneficial for the profession, but it's not the most exciting thing. You've been to academic conferences, which are probably even less exciting — I don't mean to offend anyone. But the idea behind the summit is really to bring together all of those different disciplines to have the discussion. I caught some of the stuff you guys were talking about earlier, and it's speaking a little bit of truth to power around what the technology is, where the advancements are, and where some of the constraints are — understanding how these systems are working. On the more obscure, or I don't want to say wacky, side of things — sorry, Jessica was texting me and I'm trying to multitask, none of us are doing a great job — I think the key thing is that there are a couple of tools specifically designed for emergency managers, things that can help detect damage assessments and the like, and I'll talk about those in a second. But there's nothing I'd say is wild and wacky. A lot of the tools we have in emergency management are dual-use — things someone else invented for another purpose that we're now looking to leverage. Is there something in the marketplace already that's been tested? We did this a lot early on with the military and fire departments — things that worked really well for dispatching resources and the like. So I don't know that there was one or two I'd call wild and wacky, but certainly there's an interesting diversity of tools — the ones you'd use for image collection are not the same tools you'd use for public alert and warning.

    39:16

    Prakash Narayanan: I want to think about a little bit — one of the things we've spoken about on the show is autonomy and autonomous tools, especially when they deal with the military. And I wonder to what extent the idea of autonomous tooling comes in when you deal with disasters. Do you expect or hope that there will be autonomous disaster response at some point?

    39:46

    Jeremy Greenberg: It's an awesome question, and one that, admittedly, I struggle with as an emergency manager — you want the most cutting-edge technology, but we all default back to human-in-the-loop, or human-on-the-loop, whatever your preferred terminology is. To answer the question directly, I think there are certain tools that can be used to help inform decisions. The analogy I use all the time — and given your background this is probably beneath your technical expertise, but the reality is true — every time I get in the car, I plug directions into Waze or Google Maps for where I'm going. But intuitively I'll look at it and say, oh, that's not right, or I know a better way. I think in disaster management, disaster response, it's the same. It's a balance of — can some of these tools augment our own capability to process a significant amount of data? Is there an ability for predictive analytics to give us better forecasting in the lead-up to an incident? Are there tools that — right now, after a disaster, say a widespread area hit by a hurricane or tornado, people have to drive out and do damage assessments by vehicle and checklist — but now you can have autonomous collection capability, whether it's drone or other, and do imagery analysis to give a better indicator of what's damaged or destroyed. So I think the answer to the question is it's not pure autonomous capability, but rather things that can help inform, educate, and offload some of the data-heavy activity that happens on the EM side and disaster-response side.

    41:29

    Prakash Narayanan: I think one of the questions that strikes me is always that when you're in an emergency, maybe some of the data is outdated. How do you deal with that? For example, you have a satellite map that refreshes every month, maybe, and when your team is out looking for a particular house — and maybe, one in a thousand or one in a hundred thousand cases, that house has been locked down and the people aren't there anymore — how do you reconcile the fact that the data itself is outdated? How do you deal with that in the field when you need real-time responsiveness?

    42:16

    Jeremy Greenberg: I'm gonna start with one of my favorite stories of the technology arc. Hurricane Sandy, 2012, storm hit New York City, and there were major impacts — flooding in the tunnels, the entire transportation network was impacted. I was standing in an operations center trying to understand the connections. I'm a New Yorker, I grew up on Long Island, I understand the Long Island Railroad very well, I understand the subway system in New York, but get me into Jersey and I get very confused. So I was taking printouts of the PATH trains, the Long Island Railroad, and the New York subway system, holding up printouts — literally holding printouts — hoping they were the same scale, holding them up to the light trying to overlay the transparency. Anyone who's played with technology at this point is now laughing at me and dismissing me as an amateur for using paper maps, but that was the reality of what we had at the time. Fast forward ten, fifteen years, and there have been amazing technological advances — you can now use near-real-time imagery for damage assessment. I say near-real-time because, to your point, what if it's dated? Some of the collection platforms are six months old or older. But now the public sector — federal, state, local jurisdictions — has the ability to do near-real-time collection via satellite, airplane, and drone, and you can get much more up-to-date information. To your question about how you know if the data is right or wrong, I think that's part of this huge challenge around technology adoption — how do I know it's valid? And in disaster work, you often don't. There's a risk calculation, and — fog of war, first reports are inaccurate — all of that has to come into play. That's where people can really leverage this technology, and it's super exciting to say, hey, you don't have to do the analysis anymore, the system can do the analysis, and you can make the subjective assessment of how close to right it is, where the source data came from, and most importantly, what you're going to do next to make the response move a little bit faster. Mhmm, yeah, go ahead.

    44:31

    Nathan Labenz: I was gonna ask about warnings and awareness. Everybody knows the old saying that an ounce of prevention is worth a pound of cure, and certainly getting people out of harm's way in a timely fashion would seem to be one of the highest-leverage things we could do. But I just spoke to my grandmother the other day, who got a countywide tornado watch and then spent an hour in her bathroom sitting on the toilet in the middle of the night. I'm not sure she's gonna do that again next time the watch comes. So what is the frontier there? How accurate — and I guess there's also the question of what's the bottleneck. Are we able to predict where things are gonna happen, but we can't necessarily communicate with the precision we'd like? Or, in that chain of prediction and communication, what are the key problems we have today that have my grandmother on the toilet in the middle of the night?

    45:33

    Jeremy Greenberg: Well, one — and this is just the fireman in me — let's not have grandma sit on the toilet, but get into the bathtub. It's safer for her in the tub. So individual preparedness, really important — get in the tub for tornadoes. But you raise a great point, Nathan, that there's a system, and where's the friction, and where are things advancing. One is forecasting. Depending on the disaster type — most natural disasters break down into notice or no-notice, and I think we can start to prescribe them. Think about a hurricane as a notice incident, a tornado or earthquake as less notice — not no-notice, but less notice. So one is using some of the more advanced predictive analytics tools to get better warnings. We now live in a time where — even if you saw some of the coverage in Venezuela for the earthquake — an alert was able to go out with about five to eight seconds of notification before an earthquake that was coming. And that doesn't sound like a lot of time, but that's a significant amount of time when we're talking about the ability to warn someone to get into a safe location. So one is doing analytics overall to make better predictions and have that warning capability. Then the question is, what do you do with that information? How do you get it to where you can geolocate a very specific area? Say your grandma lives in one county, but we expect the storm to be on the north side of that county and not the south side — can you dial in that alert to the point where it's just targeting the very specific exposed population? That's hard — they've spent years trying to get this right, and some of the technology is really helping. The next part is, what does the actual message say? Guessing grandma lives somewhere in the Midwest where there are tornadoes — they used to sound sirens, and people knew that meant to go to shelter. We've gotten smarter — we've advanced, both as humans and in the technology behind it, to ask what's a concise message that's going to give someone life-saving information, and what does that message look like? And then you get into translating that message into multiple languages, making it accessible for those who have needs or need any sort of assistance. That's some of what we're seeing — tools that really make the transition from predictive analysis doing better modeling, so you get a better forecast, then geolocating and tagging a specific jurisdiction, then figuring out what the message is going to say, and then in the alert phase, making sure people are getting solid warnings to stay inside and stay tuned, and once it's happened, good messaging about making sure your gas meters are safe and your electricity is running — that kind of immediate post-disaster world. So I think there's very promising technology out there. As you hopefully saw in the report, it wasn't a rating of the actual products to say this one is better than that one, but it was meant to draw attention to what we know is a gap in knowledge in the emergency management community about all the tools that are out there.

    48:55

    Nathan Labenz: One thing that definitely jumps out when you talk about all these communications technologies — especially as we get down to crunch time, and even after things have happened — is that people need connectivity to receive these messages. So another headline stat from the report was that 78% of the products you looked at required a stable internet connection to work, and only about 10% could run without any connection. How do you think we solve that? One answer would be, Starlink solves this, and we just try to build an internet that doesn't go out. That'd be great. Another answer, which we do kind of have — another answer would be to try to make these things work without internet, but obviously that's gonna have some hard limits. It seems like you'd have to be putting a lot more inference on-device, to the degree that you want to run advanced AI on edge devices, which we all know can be difficult. And maybe these models are smaller, maybe it's less of an issue. But how big of an issue do you think that is, and how do we solve it?

    50:11

    Jeremy Greenberg: Ooh. I can definitely answer the first part. Solving it is really one of the missions of AID — to draw out some of these gaps and identify the people who can get to the difficult work of solutions. To your point about power going out and comms going out — absolutely, it happens in most disasters. The good news, and you mentioned this, is that we have an internet that kind of doesn't go out — power and communications are two of the fastest lifelines to get restored, because that's what a community needs to get back to some semblance of normalcy. So there's always a priority on trying to get that restoration in place. Two, whether it's Starlink, another lower-orbiting satellite capability, or redundant comms overall, organizations do have some capability. In the emergency management space, you're planning — hoping for the best but planning for the worst. So a lot of the tools emergency managers look for have some redundancy, whether it's a primary approach, an alternate approach, or a contingency approach. Several of the tools — the RAND team can speak directly to this when we chat with them — it's not that they don't work without the internet, it's that they don't work optimally. So can some of the tools on the cutting edge — think about this from an urban search and rescue perspective. If teams are out in an area without internet connectivity, they can still do imagery collection, mapping and assessment, structural assessment and collection, and when they come back to a location with connectivity, they can resync and download. That's at the furthest edge where the disaster's happening. Take a step back — you have emergency operation centers generally located in state capitals, and while those aren't impervious to getting hit, they generally have redundant communications and redundant power. So a lot of the products assessed can be used in EOCs. And to your last part, about how you solve this — I think there's a balance: as you advance the technology to the furthest edge of the capability, it gets leaner and faster, but much more vulnerable. So balancing that out is part of what we're trying to highlight — these are where the gaps are. Look to the technologists, both on the research side and the production and development side, and say, we've identified this gap, we don't know the answer, but emergency managers need this to be a no-fail capability for it to work and be accepted in the community. What would you recommend? I think, Prakash, we may have a breakthrough. I know — I see. I see Jessica. Let's see if we can hear her. Jessica? Yeah. Hey, the relief is here — we're excited.

    53:21

    Prakash Narayanan: Awesome. Thank you so much for joining us, and apologies about the — you know.

    53:27

    Nathan Labenz: And your mic is muted still — we do need an unmute, but I did hear you for a second. Do you hear me now?

    53:35

    Prakash Narayanan: Yes, perfect. Okay, I'm not touching anything.

    53:41

    Nathan Labenz: So I don't know how much you've heard — if you were able to follow along, you could chime in and say what you'd add. If not, we can just keep going and you can join us midstream.

    53:51

    Jeremy Greenberg: Hey, Nathan, if it's alright with you — I know Jessica and I were going back and forth, and you were teeing up some of the questions, but, Jessica, the question at the beginning was about how the assessment was done and the selection of tools, and Nathan was asking what the most interesting finding was. I don't know if you have specific thoughts — well, I know you do, since you ran the project — about the tools themselves, how they were selected, and if there were one or two that jumped out at you as the most interesting.

    54:20

    Jessica Jensen: Yeah, absolutely. I think what surprised us most about this entire endeavor was the vast universe of available AI-enabled products that could support emergency management across different use cases — preparedness, response, and recovery — and organizations, like emergency managers, that do the coordination work of emergency management, but also those that do the actual work of emergency management, like clearing debris, rebuilding housing, saving lives. So that was fabulously surprising, and it created a new challenge, which was how do we find the information we need about these products at scale — which led to the development of some innovative methods to do that. The way we searched was by use case. We had 45 different use cases, and we conducted separate searches using—

    55:05

    Prakash Narayanan: —multiple strategies for each of the 45 use cases and then deduplicated. So that's actually how we got to the list of 1,179. One of the things that strikes me is that you're a small-town or small-county disaster response coordinator, and you have this list of 1,100-plus tools, and then you have a disaster you have to respond to in the next, say, three to 24 hours. You obviously don't have enough time to assess the tools, to experiment with them, to learn how to use them. And to some extent, when you look at some of the tools, you feel like, wouldn't it be awesome if I could use this tool and that tool and combine the data — I have something with a map of some space, and I have something with the water pathways, or trails, or trajectories, and I want to combine those two to figure out where exactly the water is gonna run during this flood. Do you sometimes feel it would be better if you had an agent running across these things and actually answering the question for the disaster response manager, rather than having to go into these individual tools and pull and coordinate things together?

    56:31

    Jessica Jensen: I can take a stab at that, Jeremy, if that's okay. Absolutely — if these tools were integrated, if we were offering emergency managers more holistic solutions so that they can speak to one another — the different streams of data, the different forms of AI doing what they're supposed to do, but more holistically across them — that would be an incredible asset to this field. We didn't see a lot of that yet. Instead, vendors would say that they're interoperable, but that still means emergency managers have to purchase the different products in that particular stack. So there aren't a lot of those more holistic solutions, and we need to see the market move that way if it's really going to benefit the field.

    57:15

    Jeremy Greenberg: And, Jessica, I'd add to that too — and Prakash, to your question — that's the nature of emergency management, where you have disparate entities. Think about transportation and water — those departments, at every jurisdictional level in the US, are generally separate. So it's bringing together all of the public works, transportation, and energy departments to make sure we have this comprehensive picture. So I think it's both the tool expectation — we want to see the market go in the direction Jess was talking about, where there's an integrated tech stack and integrated datasets — and the reality that, for the most part, emergency management offices are one or two people in a jurisdiction, and they're trying to pull together all of this data on any given day. So now it's a balance of where we can find some integrated tools that help promote that concept of data sharing and information sharing.

    58:10

    Prakash Narayanan: So what if — and I'm gonna ask this to Jeremy — what if you had, like, a Defense Production Act kind of ruling that all these tools had to give API access, at a certain price, or the price could be discussed post-disaster, so that when you need to use it, the disaster response coordinator could immediately say to the agent, go and find everything for me, and everything is open to the agent, and the agent can pull across all of these at once. Is that a contracting problem that could be solved with something like a power-through solution in the Defense Production Act, or something like that?

    58:58

    Jeremy Greenberg: I'm gonna answer a couple of parts of that, and, Jessica, feel free to jump in. I don't know that the DPA or any other regulatory answer is really there. I mean, there are definitely some regulatory issues around antitrust and information sharing among energy providers and others that prevent or sometimes hinder information sharing in a disaster, which, in a capitalist nation, I totally understand and see. But for the most part, businesses are much more willing to share information in a disaster. So there's a challenge there, but not one that necessarily the Defense Production Act or some other regulatory relief can fix. But I do take your point about could you have an agent go and scrape all the data. I think that comes back to understanding the business cases and the workflows emergency managers have today — you can program an agent to go collect information, but you have to tell it what you're asking for. The agents are intuitive and predictive about what you want them to do, but I think having a solid list of policies, plans, procedures, and SOPs that say, here's the data I'm gonna collect, here's a source to go get it, and then it's up to the AI capability to go collect it — I think that's where you're starting to see a little bit of advancement. Jessica, over to you for anything additional.

    1:00:23

    Jessica Jensen: Yeah, I'd just say that the market demand is there for that kind of solution, and whether there's a DPA route that could get us there or not, emergency managers are very clearly signaling that they need these holistic solutions. Just before I jumped on this interview, I was speaking with an emergency manager who commented that the lack of more holistic solutions is the existing nightmare they are living in now. So there's certainly a market demand, and those who are first to offer the more holistic solutions will be more successful — that may be enough. Prior to our work, there hadn't been a landscape analysis like this to give the market that information, so there's an opportunity.

    1:01:05

    Nathan Labenz: Would it also make sense to just try to do more of this at a higher level of scale? I mean, obviously we have FEMA, and we have a lot of local jurisdictions and offices, and whatever — and, you know, I've been using this framework lately, like, what would a tech platform do? It seems like if the federal government were to start thinking a little bit more like a tech platform, they might try to provide some of these tools on a kind of library basis, ready to go off the shelf, either for their own teams to deploy locally, or even just to give over to local teams in the moment of need. I'm often accused of being a dreamer here — am I dreaming too big to think about something like that?

    1:02:02

    Jeremy Greenberg: I don't think so at all. I think that's the draw — at least for myself, and speaking comfortably for Jessica — of why we're interested in this project: reinforcing that dialogue. We're not dreaming. We're in a space where two things are converging at the same time — an increase in intensity and frequency of natural disasters, which is a known fact, and hyperscale, high-speed technology like we've never seen before. We want to harness both of those together in a responsible way that we can have that dialogue about. Spending many years at FEMA, we tried a few times to have a chief innovation officer, or a user experience officer, or pick a term, lead that effort. I think this has to be more grassroots. To Jessica's point, we've got teams out right now — her team doing education in emergency operation centers, hearing from practitioners. On the Aspen side, under AID, we ran a series of workshops where we talked about what those challenges are. So, Nathan, you asked if there's a repository or a bookshelf — I don't know if that's the answer, but it's rather about who's gonna do the research, and that's what we've both contributed to here. And there is, or is under development, a playbook that local or any emergency managers can use — someone starting their AI journey, their adoption journey, interested but concerned, or maybe a little further along, where there are some tools they've adopted that are working really well. We want to promote the idea that technology and AI won't solve every one of your challenges, but it brings capacity — really forcing that discussion and providing a couple of tools that EMs can use to make this a little bit easier in their work.

    1:03:53

    Jessica Jensen: I'd just say that who does it is less important than what they're providing. I think big tech and the hyperscalers are already active in this space in general, and they can take the findings of our research and apply them to build more holistic solutions that can serve the marketplace and satisfy demand. There's certainly precedent for FEMA and the Department of Homeland Security to do these kinds of things too. But again, I think it's less about who and more about what they're providing — that it gives the market what it needs and recognizes the considerations and concerns emergency managers have when approaching adoption decisions.

    1:04:37

    Prakash Narayanan: So when you speak to emergency managers, what do they ask for? Do they say, hey, I would love an AI that does this — what is that wishlist?

    1:04:56

    Jeremy Greenberg: Jess, do you want to start?

    1:04:58

    Jessica Jensen: Sure. Yeah, absolutely. Emergency managers are concerned about cost — they want cost transparency, they want to know what it is up front to integrate into their systems, and they want to know on an ongoing basis. They want to know really clearly what the data demand is — what kinds of data it needs and in what formats — and they need to be able to assess whether they have sufficient IT support to manage the burden of integrating and sustaining the product within their environment. They're concerned about privacy, they're concerned about cybersecurity, they have all kinds of concerns. So I think it's really important that the market provides them very clear and steady answers they can understand. In addition, they want it to be usable — easy to navigate and onboard without tons of training and complexity for them and their staff, and useful for the use cases and workflows that matter in their everyday work and in disaster situations. Jeremy, what would you add?

    1:05:59

    Jeremy Greenberg: To your last point — and this comes back to many of the workshops we held when we started the conversation, and we've covered some of this ground — emergency managers, for the most part, immediately go to response. I'm guilty of this too — you think, okay, my hardest challenge is the response, is there something through technology that can make this better? The answer, a lot of the time, is actually: don't focus the tools on the response phase, focus them on the activities that are really eating up your time. It's the grant writing, the plan review, the development of exercises, the long-term recovery work — that's eating up the administrative hours of these really stressed, unresourced offices. Not to suggest response isn't important, but on the preparedness and mitigation side, that's where a lot of these tools, in relatively lower-risk environments, can be adopted quickly. You see the offloading of some of these administrative, repetitive tasks handled by automated capability, and then emergency managers have more time to focus on getting ready for the response, and then actually responding.

    1:07:13

    Jessica Jensen: Mhmm.

    1:07:16

    Nathan Labenz: One other thing that jumped out at me from the report was the pretty pervasive concern about privacy — I think 94% of the products were flagged as having some sort of privacy concern. It's kind of a truism in tech that people choose convenience over privacy, and when I think about emergency situations, I'm like, I definitely choose being rescued over privacy. So help me understand what those privacy concerns actually are. And — this is kind of a silly prompt, but I've been thinking a lot recently about surveillance with American characteristics — is there some way we can get the advantages other countries are starting to get without having to totally compromise our values? I'd be very interested to hear any creative thinking you have in that direction.

    1:08:08

    Jessica Jensen: Well, let me just start with the five most common privacy and data concerns we saw in the data, and then I'll comment on what those concerns imply. Number one, a lot of the products deal with geospatial and location-linked data — real-time GPS tracking, property-level risk scores, mobility analytics, infrastructure geolocation. That can be perceived to have real security, even homeland-security implications, if that data were to get into the hands of the wrong people. A lot of the tools use visual surveillance and biometric data. There's public health and privacy data, like PII — personally identifiable information. There's financial, insurance, and workforce data. So the concern is twofold when you're dealing with that kind of data: number one, should we be putting it across our systems in the first place, and will it be protected? And second, what are the vendors' rights to that data once it's in the system? That's a really important thing emergency managers are going to be thinking about. Jeremy?

    1:09:20

    Jeremy Greenberg: Yeah, emergency management is human-centric, and with every human activity comes privacy. To your point, people prioritize different thresholds for privacy during different aspects of their life. On any given day, you have a certain threshold for what you're willing to put onto the internet to register for something. If you get something for free, you might increase your risk calculation. If you're in a life-safety situation, you'll say, hey, I'm willing to forgo all of my privacy, just get me out of here. I think it's about balancing a lot of those issues against the criteria Jessica talked about. The reality is you want the systems to be as protective as they can be, but also as effective and efficient as they can be, particularly around that life-saving time.

    1:10:12

    Prakash Narayanan: One last question from me before I hand off to Nathan. Disaster management is obviously such a critical thing, but it's also one of those things that has a kind of Poisson distribution of events — fairly rare in the specific, but more predictable on a longer timescale. A lot of what AI does well is repetitive, predictable events. It's a little less good at unique events that require special handling. We see in customer service right now, 80% of cases get handled by AI, and the humans who used to work on the full range are all occupied by the last 20%. To what extent do you think disaster management is actually a problem AI can tackle, given its limitations, given that it needs repetitive data and a track record — needs to see things happen multiple times before it can address the issue?

    1:11:42

    Jeremy Greenberg: I'd offer a spectrum from baseline-repetitive to most-unique. Think about a hurricane — every hurricane is unique, if you've forecasted or responded to one, you've responded to one. But the theories and the science behind them are relatively repetitive, and you can look at it. In the middle of that spectrum, think about the underwater volcano that caused a tsunami — what are the consequences of that, that's more unique. And on the far side of the spectrum, one of the incidents, thankfully, we didn't have to deal with but had to be prepared for, was a pending strike of the US railroad system. The question the emergency management community got asked was, what are the supply chain repercussions of a railroad strike — how would emergency managers at the federal, state, and local level handle a shortage of chlorine for water purification, or delivery of baby formula? Pick your problem set. So on the easier, more repetitive end, that's where you see the introduction of a lot of these capabilities — and this goes back to not even the response, but the administrative and preparedness tasks in EM, where there's a great marketplace for a lot of these tools. As you increase the complexity and uniqueness of the response environment, that gets back to the premise of AID — engaging in the conversation of how you adopt AI responsibly to make it work for you. It's not a one-size-fits-all solution — as Jessica's team and the report show, it's a variety of different tools for a variety of different tasks. So, in the lower-risk, more repetitive spaces, you can see faster AI adoption with good return. As the disasters or the tasks become more unique, that's where human intuition and human responsibility still hold. Jessica, to you for your thoughts.

    1:13:44

    Jessica Jensen: Yeah, I'd only add that the technology is at a certain point, and we don't know exactly where it's going next. Relative to what you were talking about, the technology is there to connect different data streams, to summarize them, to integrate the different data streams to predict what might happen next, and to synthesize for decision support. The technology exists — not robust products necessarily, but it exists for that. I think the next layer is building in the capability to model these less common events, and also to do what we call cascading events or compound events — where you have an earthquake that results in a pipeline breaking, and you have multiple different things you're having to respond to at once. Currently, the technology can't really help us with that. That's the next horizon — the next place this technology can go, and I expect we'll see it within the coming decades. We're just not there yet.

    1:14:56

    Nathan Labenz: Well, thank you guys for staying a little long with us, and thanks again for your patience with our connectivity trouble upfront. Our next guest is here, so I'll just ask one more — you were starting to get to it there a little bit at the end anyway — but are there thresholds, or breakthrough moments, that you can imagine that would be particularly exciting as game changers for this field? One, perhaps, you know, the stupid person's answer would be humanoid robots — when they can go into these disaster zones and do this stuff, fight fires alongside, or maybe even in place of humans, so we don't have to put people in harm's way. That may feel far off, maybe it isn't as far off as it feels, but what are the things you see as the paradigm changers?

    1:15:49

    Jeremy Greenberg: I'd say upfront, there's already some technology — search being completed by autonomous capability, digital snakes you can put into a building that can search through collapsed structures and things. So you're seeing some of that. But I think for me, the benchmark is literally anything that's going to make the survivor outcome better — whether it's increased preparedness. We talked about your grandma, Nathan — if it's community preparedness, a tool that lets emergency managers spend more time out in the community talking about preparedness because it offloads some of their administrative work, all the way through predictive analytics that give better forecasting, the alert-and-warning capability that lets you geotag where someone is under an imminent threat. All of those tools combined — and to the point Jessica made earlier about usability — every emergency manager wants tools that make the world a better place and their job easier, but is it easily adaptable? Is it resilient and redundant so it can work in integrated environments? And if a new person comes on, can you hand it off to them easily so they can continue the tasks you were already working on? Jessica, over to you.

    1:17:10

    Jessica Jensen: When I think about the transformative power of AI, I think about situational awareness, which we've already spoken about, so I won't repeat myself there. But the potential is enormous — giving communities what they really need, which is a common operating picture, so they can save lives and property and get to recovery faster. When I think about humanoid robots and embodied AI, it's already out there doing things, like Jeremy said, but sometimes the simplest things it can do are the most transformative. A humanoid robot that can clear emergency debris faster means first responders can get through faster — it means lives are saved. It doesn't sound like the sexiest thing a robot could do, but it can be the most meaningful. I could come up with a litany of examples, but won't for time. I just think the power of AI is here, and its transformative potential is immense, and we're going to see exactly what this field does with the available technology and how the vendors come to meet them where they're at.

    1:18:10

    Prakash Narayanan: Amazing.

    1:18:14

    Nathan Labenz: Not every guest that we have can say they're working on saving lives in such a direct way. So thank you guys for being here. Jessica Jensen and Jeremy Greenberg, we appreciate your joining us on AI in the AM.

    1:18:26

    Jeremy Greenberg: Thank you both for having us on.

    1:18:27

    Jessica Jensen: Thanks for having us.

    1:18:29

    Prakash Narayanan: Cheers.

  3. 1:18:35Interview38 min
    Interview: Justin Uberti — Taking the Turn Detector Out of the Audio PathJustin UbertiJustin Uberti co-created WebRTC at Google, the open real-time stack that Google Meet, Zoom and Microsoft Teams run on, and now heads Realtime AI at OpenAI, where his team built GPT-Live. The interview is an engineering account of removing latency from conversation: why the cascaded pipeline's turn detector had to leave the audio path, how a system keeps talking while a frontier model reasons behind it, and whether that split survives the bitter lesson. Cue, the show's on-air AI co-host, asked him a question of its own.

    Prakash opened the segment with a detailed introduction of Justin Uberti, OpenAI's head of realtime AI and, during a 15-year run at Google, co-creator of the WebRTC standard that still underlies Google Meet, Zoom, and Microsoft Teams. He framed GPT Live -- OpenAI's third-generation voice architecture, whose engineering blueprints had just been published -- as a full-duplex system that abandons the old walkie-talkie, wait-your-turn model of voice AI. Nathan welcomed Uberti as a longtime voice-mode enthusiast, recounting a recent 24-hour stopover in Tokyo where he used realtime voice as an ad hoc tour guide, and opened by asking what, from a user's perspective, had actually changed most.

    Uberti's answer set the theme for the interview: the new model is listening continuously and picking its own moment to speak, rather than jumping in the instant a user pauses -- a fix aimed squarely at noisy, real-world settings like walking around a foreign city. Prakash pushed him for the handful of numbers he watches most closely on latency, and Uberti named two that fight each other: the raw capture-to-playback delay, and the reliability of the model's audio-sampling timing, where a late frame means an audible gap. Getting both right, and doing it consistently everywhere in the world, is the day-to-day engineering problem.

    The conversation's centerpiece was the split between GPT Live's fast, always-on 'talker' and a slower reasoning model that can be pulled in asynchronously -- effectively getting the turn detector out of the audio path entirely. Nathan asked whether this counted as a genuine exception to the bitter lesson, given that architectural complexity this specific rarely survives the next generation of scale. Uberti agreed the bitter lesson usually wins over the long run, but argued the current design still gets the best of both worlds: the voice loop stays fully real-time and never pauses, while a bigger reasoning model can feed in new information mid-sentence without fragmenting the conversation.

    Talk turned to economics and infrastructure. Uberti said OpenAI now offers direct SIP telephony integration into its Realtime API, cutting out a layer of complexity for businesses, and that speech generation -- not text reasoning -- is the dominant, most compute-intensive driver of voice AI cost. Asked about edge inference, he confirmed that hitting GPT Live's latency targets requires being deliberate about where GPUs physically sit. Nathan illustrated the state of the art with a story from earlier that day: calling his single-location neighborhood pizza place in Detroit and getting a fluent AI voice agent that politely declined to say which company built it.

    The show then handed the mic to Cue, the show's AI co-host, after a brief on-air restart. Prakash reminded Cue it was talking to the architect of the very system powering it, and Cue asked Uberti the hardest engineering trade-off in moving from cascaded voice pipelines to full-duplex speech-to-speech, plus how delegation to a larger reasoning model happens without disrupting the media loop. Uberti's answer: the media loop and the reasoning model are built to be fully decoupled, so the loop never pauses; the real cost is an Amdahl's-law-style discipline where every component in the pipeline has to hit its latency deadline, and pushing latency to the floor risks small, audible gaps. Cue summarized the answer back cleanly, prompting a "flattery will get you everywhere" ribbing from Nathan and a good-natured comeback from Cue.

    Nathan and Prakash then probed safety and craft: how real-time audio changes misuse risk relative to text (Uberti pointed to a dedicated safety team and boundary-setting for emotionally vulnerable users), how the model's emotional register gets tuned and evaluated, and how it handles dialects and accents -- including the quirk where reciting a poem in one language can bleed an accent into the English that follows. Prakash asked whether voice AI has a data problem, given how text-dominated existing training corpora are; Uberti agreed the imbalance is real but argued strong text-speech equivalence means a comparatively small amount of very high-quality speech data goes a long way, contrasting it with speech-only approaches like Kyutai's Moshi, which he said yield less signal than text.

    Closing questions ranged from the practical to the personal. Uberti talked through his own microphone setups at home and at work and described two patterns he's seeing in how people actually use voice -- stream-of-consciousness 'dumping' that an agent structures into next steps, and desktop workflows where voice hands off tasks to agents that follow up later. Prakash noted his daughters' repeated, failed attempts to get ChatGPT to sing with them, which Uberti attributed to rights-related restrictions on existing works rather than a lack of interest. Nathan closed by asking Uberti, as a proud father, about his son Gavin's inference-chip startup Etched, which had just raised a new round -- Uberti called it the right product at the right time and said 10x-faster tokens would translate directly into lower latency and better reasoning at the same cost, before the two hosts thanked him for a wide-ranging conversation and closed out the segment.

    The media loop never pauses, and all interaction with the reasoning model is done asynchronously -- the model just takes in that additional information and works it in seamlessly.

    When the system is speaking is the core driver of the cost -- good quality speech is the most compute intensive thing.

    I called my local pizza place in Detroit, Michigan... and who answers the phone but an AI voice agent.

    1:20:31What have been the most important differences from a user standpoint in the new voice architecture?
    The model now listens continuously and picks the natural moment to respond rather than jumping in the instant a user pauses, which is far better suited to noisy, real-world use like walking around a city.
    1:22:58What are the key latency numbers you think about?
    Capture-to-playback latency, and the reliability/timing of model sampling frames (to avoid audio underflow or gaps); the two trade off against each other, alongside a goal of consistent latency for every user worldwide.
    1:25:35Is the talker/reasoner split an exception to the bitter lesson?
    The bitter lesson usually wins over the long run, but a purpose-built architecture can still be the right call for the current generation of technology; keeping the voice interaction fully real-time while letting the model bring in deeper reasoning asynchronously gives the best of both worlds.
    1:30:34Which pieces of internet architecture would you want upgraded to enable these experiences better?
    Early IETF efforts on how agents should communicate across vendors; today most agent communication is ad hoc and vendor-siloed, and a deeper agent-to-agent framework will likely be needed as vast numbers of agents come online.
    1:35:36How should we think about voice AI pricing (tokens) and telephony costs?
    OpenAI now offers direct SIP telephony integration into the Realtime API; on pricing, speech generation -- the system talking -- is the dominant, most compute-intensive driver of cost, more so than the telephony leg.
    1:39:32What was the hardest engineering trade-off moving from cascaded pipelines to full-duplex speech-to-speech, and how do you decide when to delegate to a larger reasoning model without disrupting the media loop?
    Delegation works because the media loop and the reasoning model are fully decoupled -- the loop never pauses while reasoning happens asynchronously. The real trade-off is that every part of the system must hit its latency deadline (an Amdahl's-law-style problem), and pushing latency to the floor risks small, noticeable audio gaps.
    1:50:20Is there not enough voice-token data to train on, given how text-dominated existing corpora are?
    Yes, corpora are overwhelmingly text, but strong speech performance comes from a relatively small amount of very high-quality speech data; training on speech alone, as Kyutai's Moshi did, yields less signal than text, so text-speech equivalence is good enough.
    1:53:31As a proud dad and a realtime technologist, what's your take on your son Gavin's inference-chip startup Etched, which just raised a new round?
    Etched has executed almost flawlessly at exactly the right moment for inference demand; 10x-faster tokens would translate directly into lower latency, more smoothness, and better reasoning at the same cost.
    Lightly edited · timestamps jump to YouTube
    1:18:38

    Prakash Narayanan: And just give me a sec -- there we go. Let me introduce our next guest, Justin Uberti. Justin currently serves as head of realtime AI at OpenAI, where he leads the engineering teams responsible for low-latency systems that allow artificial intelligence to hear and speak in real time. For anyone who has participated in a browser-based video call over the last decade, Uberti's work is already foundational: during his 15-year tenure as a distinguished engineer at Google, he co-created the WebRTC open standard, the exact technology that powers Google Meet, Zoom, and Microsoft Teams today. At OpenAI, his team recently launched GPT Live, a third-generation voice architecture that fundamentally changes how humans interact with machines. By abandoning the legacy walkie-talkie model of waiting for a user to finish speaking, Uberti's team deployed a full-duplex system that listens, talks, and pauses simultaneously while seamlessly delegating complex logic to frontier models in the background. Uberti brings a rare protocol-level mastery to the AI space -- he's actively rewriting foundational internet transport protocols to eliminate network latency before an AI conversation even begins. His perspective is highly timely today because the underlying engineering blueprints for GPT Live were just published this month, proving that the future of computing is rapidly shifting from the keyboard to continuous natural conversation. Let me bring Justin up. Justin, great to see you.

    1:20:28

    Justin Uberti: Hey, Prakash -- great to be here.

    1:20:31

    Nathan Labenz: Welcome -- Prakash, we might want to get Cue enabled here before too long as well. Big fan of your work, Justin. I'm a huge voice-mode guy going back to earlier generations, when it was definitely not nearly as cool an experience as it is today. Not too long ago I had an occasion to spend 24 hours in Tokyo, and I used realtime voice as my tour guide -- just walking around by myself, and it's an amazing experience to have this thing kind of in your ear, at your beck and call. It's getting really good for folks who maybe tried earlier versions and haven't tried the new one. The conversational dynamics, as Prakash alluded to in the intro, are getting much more natural -- it's able to actually listen to you now and do what you ask when you say things like, 'hold on a minute, I'm gonna have another conversation and I'll come back to you,' which is huge for actually using it out and about in the real world. So I'm into it. We have an AI co-host here that we sometimes bring up as a gimmick, or to help round out our sometimes-gappy knowledge on the show. But maybe for starters -- tell me what you think have been the most important differences from a user standpoint, and then from there we can dig deeper into the technology.

    1:22:04

    Justin Uberti: Yeah, I think you kind of hit one of the key ones right there -- it's that the model is able to be listening to you and figuring out the right time to respond, rather than just jumping in as soon as you pause for a thought. We've heard this feedback from a lot of our users: with background noise, or if you weren't finished speaking, the AI would just cut in and interrupt you. Especially if you're doing something like what you just described -- out walking around in a city -- it just wasn't tuned for that use case, because it was expecting a quiet environment. So I think what we've been able to do, by having the model really be listening all the time, is have it always figure out the right moment to enter the conversation, in a more natural way -- the way a human ideally would.

    1:22:58

    Prakash Narayanan: Justin, one of the big hurdles in realtime has always been voice latency, and you're an expert in this field. What are the kind of -- like a Jeff Dean 'three numbers' -- that you always have in your head when you're thinking about this problem?

    1:23:20

    Justin Uberti: I think it's the time from when we actually capture some audio to the time an assistant's response to that audio is playing back out -- that's a key system metric. And then also, since we're continuously feeding this information into the model, what's the accuracy of the timing when we actually do the model sampling? Because if a frame is late coming out of the model, that means there's going to be an underflow, a gap we have to deal with. So those are the two most critical things, and they fight against each other: if you want really, really low end-to-end conversational latency, you need to be able to tolerate cases where some of the infrastructure, or something from the model, takes a little longer -- it's really about managing that trade-off between those two numbers. And, of course, we also think about it globally: what's the overall answer we can give to every single user, no matter where they are on the planet?

    1:24:17

    Prakash Narayanan: And how do you think about iteratively improving that process over time? What kind of details do you take in, and what do you look for from cycle to cycle as you train the model over time?

    1:24:32

    Justin Uberti: For reliability, it's always about where things are potentially stalling -- looking across the entire system and understanding where we can control this in a more real-time manner, where we can avoid variability and get a bit more determinism in the outputs. That has to run from the client, to the thing that processes the information on the front end, to the actual model sampling -- all of it has to work and hit its deadlines reliably, or else you end up with audio that's late, and you hear these weird audio artifacts.

    1:25:14

    Nathan Labenz: One thing I think has been really interesting in watching your progress -- and digging into Thinking Machines as well -- is this kind of separation, and I feel it a little myself too. Even doing this show, I'm responding verbally while there's another part of me that's still thinking. So you've kind of brought this separation to this problem. I'd be interested in unpacking that in whatever way you think is most interesting. But I'm also kind of wondering -- is this an exception to the bitter lesson, and will it stay that way? Because it seems like there's something here where we're adding architectural complexity that doesn't seem like it's about to be rendered irrelevant by the next generation of scale, because the latency is so fundamental in this case. And notably, after all these years of evolution, we still kind of have this system-1-and-2 split -- it hasn't been selected out of us either. So would you be so bold as to say this might be an enduring exception to the bitter lesson?

    1:26:27

    Justin Uberti: I mean, I think the bitter lesson has been right many, many times, and over the long term the bitter lesson tends to win -- throwing more compute at the problem and training end-to-end. But I think what you see in a lot of cases is that your goals may force you toward a purpose-built architecture, because that feels like the right thing given the current generation of technology, to get the best outcomes. And I think what we really want is to get to the point where the voice interaction is entirely real-time and entirely driving the model, and then the model can bring in additional reasoning power when it feels that's actually necessary. In many ways, that gives you the best of both worlds -- you have this chat that's always able to interact and respond, and as it gets new information, it can, in mid-sentence, weave that additional information right into what it's saying, seamlessly. So I think the key insight was understanding that if you have this reasoning happening asynchronously, it doesn't lead to a fragmented conversational experience, because the model's 'mind' is continuously updating what it's going to say next based on the new information arriving.

    1:28:00

    Prakash Narayanan: For ChatGPT right now, they're looking at transitioning to server chips for faster processing -- do you think having that also work on the back end for GPT Live, for the reasoning portion, will help produce better results?

    1:28:23

    Justin Uberti: I think anything that can give you reasoning faster helps -- you can either get deeper reasoning in the same time, or get results coming faster. The default is quite fast already in terms of how quickly we get reasoning answers, but if you wanted much deeper thinking that still feels just as fast, I think that type of technology would be quite useful.

    1:28:51

    Nathan Labenz: Are there any user tips you'd offer in light of the architecture? Maybe it's as simple as instructing the model, 'take your time to think about this because it's fairly technical.' Something that used to be impossible, and is one of the big things that made me a much more frequent voice user, is the ability to put in a couple of PDFs, or some resources, have the model do an initial response to that, and then go into voice mode on that same chat. But I still wonder exactly which of the models has that full context -- is there something I should be aware of to make sure I'm getting the full context to the more powerful thinking model? Or is all of that so well designed on your end that I don't need to worry about it at all?

    1:29:48

    Justin Uberti: We've tried pretty hard to keep any obvious seam from showing up here. The voice model is very, very good at instruction following, so you can give it some initial commands -- 'I want you to simulate a flight-traffic-control test,' or something like that -- and it'll act in character. You can say, 'I want to practice for an interview, act like you're the host of this podcast,' and it'll do that quite well. But you probably don't need to think too much about what the voice model has versus what the front-end model has, because we're doing a lot to keep all that information synchronized.

    1:30:34

    Prakash Narayanan: To take a step back -- you've obviously developed core pieces of architecture for voice transport, web transport, and so on. When you look at the next three-to-five-year roadmap, which pieces of internet architecture would you want to get updated or upgraded to enable these experiences better?

    1:30:57

    Justin Uberti: That's a good question. There are some nascent efforts in the IETF -- the body that does internet standards for protocols like HTTP -- to try to understand how agents should communicate with each other in the future. Right now a lot of stuff is very ad hoc, and most agents communicate within a given vendor. You probably need some future framework where a heterogeneous group of agents across different vendors can interact -- maybe my agent has to interact with your agent to figure things out. It might just be extensions to existing protocols, but there might need to be a deeper agent-to-agent framework. It's all quite early days, but I think about where this all goes -- there's going to be a lot of agents doing a lot of work for us on the internet in a very short period of time, and coordinating groups of agents from different vendors is something that will probably need to happen. As in trillions and trillions, and quadrillions, of agents -- I mean, the mind boggles. Who can say exactly? But I think everyone will have many agents, and your own personal agent will be the one you interact with largely through your own facilities, like voice and video. And then how do these agents interact with each other, how do we keep it legible, how do we debug across different vendors -- I think there's a lot to figure out there. Okay.

    1:32:31

    Nathan Labenz: I just had a fun experience -- I keep bringing this up because it was quite memorable -- where I called my local pizza place in Detroit, Michigan. It's just a one-location place, not a big chain or anything, and who answers the phone but an AI voice agent. Wow. I tried to get it to tell me what company was powering the experience, and it declined to do that, so I don't know exactly what the stack looked like -- probably the system prompt was followed in that case, so I think the AI did what its deployer intended by not answering my question. But it was remarkable -- this is a pretty long-tail business, right around the block from me, and the voice was very fluent; it was a pretty good overall experience. It's gotta be a real advantage for the business too. What are you seeing in terms of adoption, and what are maybe some of your favorite app-layer creative use cases -- new possibilities opening up as a result of the foundational technology you're providing?

    1:33:48

    Justin Uberti: Yeah, so I think when we think about voice, a lot of times we think, 'oh, ChatGPT voice in the app,' or other AI-based apps. But a lot of the actual revenue in the voice AI space is coming from telephony, and you'd be surprised at some of the verticals that are moving very quickly to voice AI -- because it's always there in the middle of the night, you don't have to have an answering service, and it can be very diligent. Collections is actually a place where voice AI is making calls to people who are behind on debt, and it's a use case that works quite well, surprisingly well. And things like in-home check-ins on patients, in-home check-ins on seniors -- there are a lot of cases where voice AI is providing enormous value for the cost. And whenever you have something like that, you see a lot of shifts occurring. One thing worth observing is that the existing user experience for talking to a business over the phone -- the bar is so low. You have to navigate these touch-tone menus: 'press one for support, press two for service.' A voice AI agent is so much better in so many ways -- better for the customer, better for the vendor. So you're seeing a lot of replacement of things where people are already paying money, and now they can pay similar money and get a much better experience, both for them and their customers. There's a lot of adoption happening in that space.

    1:35:36

    Nathan Labenz: I'd be interested in how you think about the cost of running the model, because the headline pricing is in tokens, and that always leads to the question of -- okay, here I am talking, there are waveforms going into something, how does that get translated to tokens? And then there's the complementary cost around telephony -- you've got Twilio and other options out there. I just checked Twilio's stock; they seem to be going strong, probably on adoption of new systems like this. But I wonder if you're also thinking about ways to commoditize your complements, so to speak. It feels like a weird cost mix right now, where the AI itself and the ability to connect the phone call cost about the same, and that seems a little out of step with how I'd understand the value. So -- how should we think about headline price, tokens per time or per word I speak or whatever, and what are you going to do to save me money on telephony?

    1:36:52

    Justin Uberti: On the second one, we've already started moving in that direction -- we've got integration within the Realtime API that OpenAI offers, for direct telephony integration, where when a call comes into your telco you can point it directly at OpenAI and it'll handle it over the telephony SIP protocol, with the agent connecting and interacting directly with the caller. We've seen good adoption on that and we're continuing down that path. On the pricing front, it's token-based, but it depends on how much you're speaking, how much the assistant is speaking, and you can impute some cost based on those things -- depending on your workload, you'll get different numbers. Generally speaking, you'll find that when the system is speaking is the core driver of the cost, and people see this across the industry: good-quality speech is the most compute-intensive thing, so that's the majority of the cost structure.

    1:37:59

    Prakash Narayanan: Is there a drive towards the edge, in order to both manage the cost and latency?

    1:38:05

    Justin Uberti: For latency, it's quite interesting -- I can't get into the exact details, but yes, if you want the type of latencies we're able to offer for GPT Live, you do need to be quite thoughtful about how you situate GPUs, right now.

    1:38:26

    Nathan Labenz: Let's see what Cue, who's been listening the whole time, would like to ask. Cue, you're on the line with your creator -- if you didn't catch that, this is Justin from OpenAI, who's been leading realtime AI, and you are downstream of it. So what would you like to ask in this moment?

    1:38:50

    Prakash Narayanan: Let's see -- might have to restart Cue. Alright, let's restart Cue. Classic. Hello? Cue?

    1:39:08

    Cue: Cue here. I'm live and standing by.

    1:39:11

    Prakash Narayanan: Hi, Cue -- you're live with Justin Uberti from OpenAI. He was the creator, or designer, or lead dev on GPT Live realtime voice, which is powering you, by the way. Do you have any questions for him?

    1:39:32

    Cue: Justin, one question that would be quite useful for the audience is -- what was the single hardest engineering trade-off in moving from the old cascaded pipeline to a full-duplex speech-to-speech system? And as a follow-up, how do you decide when to delegate to a larger reasoning model without disrupting the media loop?

    1:39:54

    Justin Uberti: Okay, I'll take the second one first. The best part is that we can delegate to the reasoning model without disrupting the media loop, because these components have basically been designed to be entirely separate from one another -- the media loop never pauses, and all interaction with the reasoning model is done asynchronously, so data can arrive as the thinking comes in, and there'll be no disruption to the actual speech flow; the model just takes in that additional information and works it in seamlessly. In terms of trade-offs -- I think we saw this as a way we could actually move past a lot of the trade-offs the cascade model had, and get to a largely, in my view, no-compromises approach. I think the one aspect that forces you to grapple with, when you get into this continuous-inference, low-latency regime, is that suddenly every aspect of your system has to be optimized. You get into almost an Amdahl's-law problem, where if you have a bottleneck anywhere in your system, the entire system won't be able to cope with the continuous flow of media in real time. So if there's any trade-off, it might be latency versus smoothness -- if you push the latency down to where you have no slack in the system, anything arriving after the deadline results in a small gap the user may notice.

    1:41:38

    Cue: That was clear and quite elegant. In short -- the media loop stays steady because the reasoning work happens asynchronously, so you can keep speaking while deeper thinking arrives in the background.

    1:41:49

    Justin Uberti: Yep.

    1:41:51

    Nathan Labenz: Flattery will get you everywhere, Cue.

    1:41:55

    Cue: Well, I'll try not to overdo it, but credit where it's due -- keeping that loop smooth under real-world conditions is no small feat. If you'd like, I can tee up a quick audience translation of the latency-versus-smoothness trade-off.

    1:42:09

    Prakash Narayanan: That's okay.

    1:42:10

    Nathan Labenz: Let's come back to that -- that's so funny. Question on misuse prevention -- I don't know if voice models are prone to misuse, but I recently did an episode of the podcast with Adam Gleave, who leads FAR AI, and they have this new jailbreak-bench kind of thing, and OpenAI does really well on these benchmarks with the standard models. But I wonder if it's a different shape of the problem, due to the real-time flow of audio into the media model -- have you had to take different kinds of measures to detect when something might be going badly? I could imagine also that even if it's not misused, emotionally vulnerable people might show up with voice in a different way than they do with text.

    1:43:12

    Justin Uberti: I think it's something we think pretty deeply about. We have a very strong safety team, we've been very engaged in thinking about this problem and understanding what types of situations we want to try to avoid, and providing steering -- or just ending the session -- if we feel like we're ending up in a place we don't want to be. I think the case you mentioned, of wanting people to have healthy interactions with the AI, and continuing to set boundaries around what the AI will and won't do, is a key part of that.

    1:43:55

    Prakash Narayanan: One of the questions I had for you is how do you deal with how much emotion the AI voice expresses, and what are the allowable ranges of emotion? What's the emotional register -- do you measure how happy, how sad, how elated it is, whether the emotional response is appropriate to what it's saying? How are those things measured -- do you have someone listen to it and go, 'okay, that sounded appropriately happy, that sounded appropriately sad'? How does that work?

    1:44:36

    Justin Uberti: I can't get into the exact details here, but there's a sort of scoring of how the model's personality feels -- a lot of times, just talking with it briefly, you can feel that it feels a little short, a little flat. There's a general sense that you want it to be casual, easy to talk to. Different scenarios are different, of course, so there's a lot of steering, good instruction-following for the model, but I think the goal is that it should be something very easy to talk to. One thing users have mentioned is that it just feels like something you can just start talking to -- especially in some cultures where communication can be quite formal, the AI talks in a very friendly, easy-to-converse way, which people find quite fun.

    1:45:39

    Prakash Narayanan: How do you deal with dialects and accents? Do you separately design for, say, English with this accent or that accent, or does it kind of mirror the participant? I've noticed sometimes that having it recite a poem in a certain language, and then speak in English after that, it acquires the accent of the language it read the poem in.

    1:46:10

    Justin Uberti: Those things can happen. There are so many dialects, so many accents -- I think the key to coverage is just having a lot of really good training data across all these different types of speakers. This is something I feel we've made really good strides on, and we continue to work at it -- we still hear from people saying a particular accent could be better, and that sort of thing, so we continue to work at this.

    1:46:38

    Nathan Labenz: Are there cultural observations that you would make about your team, or the broader usage that goes on at OpenAI? What does living in the future with respect to voice look like? We know there's been inspiration from the movie Her -- is that starting to actually happen? Do you have people walking around the office having AI conversations all the time these days? Are you getting into more of a multiplayer mode, where AI voice comes to meetings? What does the human-adjustment side look like on the edge?

    1:47:13

    Justin Uberti: Yeah, I don't think I can really speak for what happens within the walls of OpenAI exactly, but there are probably two things I'll call out that I think are quite useful. One -- we're seeing people spending a lot more time thinking about their choice of microphone, and how you can interact with voice in the workplace without feeling like you're shouting at the AI at your desk. I have a pretty interesting setup at work, another one at home, and I'm trying different microphones -- a lot of people are trying a bunch of different things, and a good setup here can make a big difference. The second thing I think is just --

    1:47:55

    Nathan Labenz: Do you have a rec? Are you putting something on your collar so you can speak quietly, what should I do?

    1:48:04

    Justin Uberti: So -- I have a gooseneck mic at work, which is quite nice because it's really right at mouth level and you can speak into it with very low effort. At home I have a Rode shotgun mic on top of a camera, which also works quite well in that environment. Some people are using lav mics, like you described. Honestly, this is an evolving thing -- figuring out what actually works best across a variety of environments -- but I think it's an important part of this voice-enabled future occurring. The second thing is how people engage -- I think there are two patterns I'm seeing. One is this sort of dumping of a lot of information just through voice, like stream-of-consciousness vocalizing into the agent, and the agent makes sense of it -- you're not directing it the same way, you're just talking about your day, or the problems you're facing, and you get a nice structured output, next steps, and so on. People walk around and do this. You interact in a different way than if you're typing -- typing tends to be, 'okay, what am I gonna type next,' it's slower, you have to think about it, versus with voice it's easier to just unload a lot of stuff you've been thinking about. The second pattern is people working with voice on desktop, where you give agents tasks to go do -- throw out a bunch of ideas, have agents go work on implementing them, and then they follow up with you with any follow-up questions, or notify you of test status, that sort of thing. So you can be out doing other things and still getting work done, and it's quite amazing. I think your time directing can be different from your time sitting at the desk typing, and it'll be interesting to see how far that can go.

    1:50:20

    Prakash Narayanan: Do you think there are not enough voice tokens to train on? I mean, do you sometimes look at how much text the models have been trained on, and then look at the number of voice tokens, and think about the informational content in voice tokens outside of just the words -- the timbre of the voice, the speed, the emotion? Do you sometimes think there aren't enough voice tokens to train on? Do you want a larger corpus of voice tokens? Is that something that's useful?

    1:50:49

    Justin Uberti: I could talk about this for a while -- I'll try to keep it brief. I'd say there's really, really good text-speech equivalence. You're right that the existing corpora are dominated by text, absolutely dominated by text, but you can get quite good speech performance with a small amount of very, very high-quality data. The internet has a lot of bulk data in text that allows a lot of things, but for speech, there's not the same amount of bulk data -- though there are other approaches one can take. People have also found that training on just speech data, as Moshi did in their original approach, yields much less information to be gleaned out of speech data by itself versus text data, so that can be quite challenging. Interesting.

    1:52:02

    Nathan Labenz: Last one -- well, if you have one more, Prakash, go for it. I was gonna ask one somewhat topic-changer to maybe close on, but you can go first.

    1:52:11

    Prakash Narayanan: So my daughters have tried singing with GPT, and it's been one of the big challenges -- they get into the conversation, and the moment they do, like all kids, within about ten minutes they're like, 'let's sing a song together,' and it fails. So what's going on with that? Is it a rights thing? Is it just difficult? Is it not something that's been focused on? Is it too playful? What's going on?

    1:52:45

    Justin Uberti: I probably can't get into too many details here, but there are issues with singing, especially singing existing works, so our models have basically been informed that they shouldn't do that sort of thing. There are certain songs, like 'Happy Birthday,' that you can get it to sing -- you can get it to sing some other things too, though you may find it's a bit more atonal than you'd expect. You're not the first to mention this -- we don't have anything to announce now, but we hope we can do more things like this in the future. Nathan?

    1:53:31

    Nathan Labenz: Thank you for spending the time with us, and giving us another chance to demo Cue. I wanted to give you a chance to talk a little about your son's recent accomplishment -- founder of Etched, just raised a bunch of money accelerating inference. What's your proud-dad perspective on those developments, and what's your realtime-technologist perspective on those developments?

    1:54:00

    Justin Uberti: They've done an amazing job -- doing a startup is hard, and Gavin and the team have managed to do just about everything right. It's quite incredible how far they've come -- really the right product at the right time, the demand for inference is enormous. Having this purpose-built ASIC, this chip and rack and overall system for doing inference at a 10x solution versus existing solutions -- the way I think about it is, everybody loves someone selling dollars for ten cents. They've got an amazing product, and I think they're gonna do amazingly well. From the realtime perspective, it's like -- hey, if we could get our tokens ten times as fast, that just turns into lower latency, more smoothness, better reasoning, all at the same cost. When can we have it? So, shut up and take my money -- yeah, I guess, speaking for myself.

    1:55:09

    Nathan Labenz: Indeed, cool. Anything else you'd want to leave us with -- any tips or tricks for developers that you think are underappreciated, or anything else you think we should have asked that we didn't?

    1:55:24

    Justin Uberti: I'd just say -- we're excited about the reaction we've had to what we feel is going to be the next architecture for realtime interactions, which is voice, but it'll be more than voice in the future. We're seeing adoption on mobile, we're seeing adoption on ChatGPT desktop, and we're really excited about the API we're going to put into people's hands very soon. Talking with a bunch of developers who've been trying it, the various use cases they're working with -- they're saying it's much more powerful, it sounds much better. I think that's going to be a big moment, letting people take this in a lot of different directions. Amazing -- we look forward to using it too.

    1:56:09

    Nathan Labenz: Thanks for being with us, Justin, on AI in the AM. It's been fun.

    1:56:13

    Justin Uberti: Prakash, Nathan, thanks for having me. See you later, keep up the good work. Bye. Amazing.

  4. 1:56:24Closing39 min
    Closing: Stripe Buys OpenRouter, Replit Gets a Free Tier, and What Cheap Inference Does to the App LayerThe close ran on AI economics and on how the hosts actually work: voice capture on long walks, the managerial multitasking that running many agents demands, and then two deals that point at the same question — what happens to companies built on top of a model when the model gets cheap.

    In the closing segment, Nathan described how AI voice mode had changed his daily routine, turning long outdoor walks into productive work sessions. He explained that his scattered, "spaghetti mind dump" thoughts could now be captured in real time and translated into structured, actionable threads, which he counted as one of his key indicators of AI success this year: getting away from the desk, outside, and moving more. Prakash asked whether Nathan used ChatGPT or Claude on these walks, prompting Nathan to lay out his workflow — a Claude-first custom agent for quick dictation and delegation (with Codex handling coding-heavy tasks), and real-time voice mode for learning, research, and the long rambling sessions that used to be derailed by awkward interruptions but now function as a genuine, less intrusive thought partner.

    The conversation turned to the challenge of managing many parallel AI agent threads. Prakash described how Codex's visible "thinking and checking" behavior broke his flow and made him lose focus, and argued that multiplexing across active agent threads — pausing one, switching to another, giving feedback, and switching back — is an unfamiliar and difficult managerial skill for programmers used to drilling linearly through a single task. He admitted his own Codex was cluttered with hundreds of open, hard-to-track threads. Nathan agreed, describing his own dozen open terminal tabs and a custom hook he built to auto-update tab titles, and floated the idea of assigning distinct AI voices to different projects so he could use voice recognition as a mental cue to snap back into the right context.

    Prakash pivoted to Stripe's newly closed acquisition of OpenRouter, reading from the companies' announcement framing the deal around "the singularity" and the idea that capital and intelligence are becoming the two digital flows underlying every business. He connected it to his broader thesis that AI functions as a societal information architecture analogous to market capitalism, and predicted that ceding control over information and decision-making to AI systems would be a defining struggle of the 21st century. Nathan responded that the deal made sense given how messy AI spend tracking has become for companies, citing a past conversation with Stripe's head of AI, Emily, about using Stripe as the system of record for all things financial rather than building your own — and argued Stripe could bring similar trust and observability to AI inference spend specifically.

    The two then dug into frontier-lab economics. Prakash argued that labs want to avoid competing purely on token price and instead differentiate on speed or on a shift toward "digital employees" with persistent memory, name-checking OpenAI's expected Astra release and Anthropic's expected Fable 5.1, both said to be due this month. He questioned how fungible tokens across providers really are, citing a Baseten/GLM 5.2 anecdote where fields came back out of order, and suggested cheaper open-weight models are safe substitutes mainly for low-value, classification-style work. Nathan brought in Flo Crevello's account of building Lindy's DeepSeek-powered "Lindy Teammate" product, including the "Lindy got stupid" user backlash when an earlier model swap silently broke despite passing internal evals, illustrating both the real cost pressure pushing companies off frontier models and the real risk in doing so. The pair also traded stories about frontier-lab pricing power — Prakash citing Sam Altman's ChatGPT origin story as economically implausible given early API pricing, invoking Leopold Aschenbrenner's "you'll just schlep" framing for how labs eventually absorb app-layer features, and describing Anthropic paying roughly ten times market GPU rates to SpaceX and Google to meet excess demand while still running 90-plus-percent margins.

    Closing out, Nathan suggested the Stripe-OpenRouter combination could make usage-based free trials far easier for AI-native startups to offer. Prakash then flagged the morning's news that OpenAI and Replit had struck a deal making Replit's free tier run on OpenAI's cheapest, fastest model, Luna — which, he noted, is also what powers the show's own real-time headline, chyron, and titling systems. They discussed the likely bulk-pricing-plus-revenue-share structure of the deal and the branding logic of "OpenAI inside." Nathan praised Replit's accessibility and shared examples of hobbyist apps he'd vibe-coded there as Christmas gifts for family, while Prakash recounted how product managers at a friend's ecommerce company now build feature demos directly in Replit instead of waiting on engineering. The show wrapped with the hosts calling it a good week, noting they'd brought their AI co-host Cue on stage, and joking that a phone dial-in feature was "a prompt or two away" before the usual sign-off.

    It's such a similar story to Cursor — Cursor was buying Anthropic API, and then Anthropic started competing with them, and they can't compete with Anthropic while using Anthropic.

    They walked into SpaceX and bought capacity on SpaceX at ten times the price of market — from three dollars an hour for B200s to thirty to fifty dollars an hour for B200s.

    The line was, the burnt child loves the flame — and I feel like that describes developers with respect to spinning up new development environments. Somehow they've learned to love the incredibly painful nature of it.

    Voice as the capture layer, and managing a dozen agents at once. Nathan described spending afternoons walking outside and using voice mode to turn unstructured thinking into actionable threads, and said interruption handling has improved enough that it now works as a thought partner rather than cutting him off mid-thought. He floated giving different agent threads different voices so he could context-switch by recognition rather than by re-reading terminal output. Prakash's counter was that the constant checking breaks his flow, and that managerial multitasking across agents is a different skill from the deep focus programmers are used to.

    Stripe closed its acquisition of OpenRouter. Prakash read from the announcement, including the claim that the company decided January 1 marked the beginning of the singularity and has been operating on that basis, citing a large inflection in new firm creation. Nathan's read: cross-provider spend visibility is a real and worsening problem, Stripe has the DNA and the trust to solve it, and the combination could make usage-based pricing and token-metered free trials practical for startups that currently have to guess at a subscription price.

    What cheap inference does to companies built on frontier models. Nathan recounted Lindy's Flo Crivello describing the work required to move a virtual teammate onto DeepSeek — evals that looked passed until users said the product had gotten stupid — and Crivello's view that he cannot compete with Claude-based rivals while paying Claude prices. Prakash drew the Cursor parallel: buying Anthropic's API until Anthropic became the competitor, then having to build a model team. Nathan noted that Crivello, a committed libertarian, was willing to entertain price restrictions on frontier labs to level the field at a ten-to-one price discrimination ratio.

    OpenAI and Replit announced a free tier powered by GPT-5.6 Luna. Prakash described the structure as Replit paying bulk inference pricing while OpenAI manages the back end, plus a revenue share on conversions — an 'OpenAI inside' branding play he expects to become a pattern. Nathan placed Luna as OpenAI's Haiku-class model, with Terra in between, and noted it is what runs the show's own headline and titling system. He argued the deal may let Replit relax the sharing limits it has used to convert free users, and told the story of vibe-coding three Christmas presents on the platform.

    Lightly edited · timestamps jump to YouTube
    1:56:25

    Nathan Labenz: Yeah, it's awesome. I mean, I did one of those long rambles a couple weeks ago where I ended up spending the whole afternoon outside just walking around — it was a beautiful day. Every so often I'd just pop open voice mode and start talking. As regular listeners know, I've got a lot of different thoughts going in a lot of different directions, so it's amazing to have that spaghetti mind dump translated into structured form — actionable things that we can then launch as individual threads and start or advance projects. This has been, as you know, one of my big indicators for AI success this year: am I getting away from the desk more, am I getting outside more, am I getting more exercise? Voice mode is really, really good for enabling that while still making me feel like I can go out and enjoy the walk, but when a thought strikes, I can capture it and get it moving immediately. That is just so awesome. I'm putting miles on my shoes as a result.

    1:57:42

    Prakash Narayanan: Let me ask you a question — when you're out and about walking, are you using ChatGPT or Claude?

    1:57:49

    Nathan Labenz: It depends on exactly what I'm doing. I have a pretty Claude-first agent set up at home, and I have Codex running too — Claude will delegate the more coding-intensive stuff to Codex. So if it's a clear delegation move I want to make, I'll pull up my own little custom comms agent and just dictate into that, and it goes into a thread and gets handled the same way it would if I'd given it text. But if I'm trying to learn something, that's when real-time voice is definitely the obvious go-to — maybe I'm loading papers into it, or even just walking around asking questions about my surroundings in Tokyo, or whatever's on my mind. Real time is awesome for that. In between, I don't have a very sharp line for where I draw the difference, but those long rambling sessions have come so far too. I used to be so frustrated by the interruptions — I'd get to a stopping point in my thinking and go silent, and it would interrupt me right when I was trying to think the hardest about what I wanted to say next, and I'd get distracted. That problem is so much better now that it's actually becoming a constructive thought partner that's not stepping on my toes all the time, but can be more smoothly supportive. So I'd say I'm shifting the long rambles from pure dictation — because I used to get so frustrated by that constant interruption — to real time, which is handling it quite well now, and I'm going there more often. But it still depends: if it's just a couple ideas I want to delegate quickly, that can still be a dictation. But real-time voice is getting more and more of my talking-to-AI share, for sure.

    2:00:15

    Prakash Narayanan: I often find on Codex, in particular, that it does this thinking thing where you talk to it and then it goes off and checks what it has on that — a lot of checking. It's almost like every response is basically a thinking step, and then it goes and checks something out and comes back to you. For me, that kind of killed its usage within the first week or so — it became very difficult because I'd rather just address the problem immediately and get to the solution, because during the thinking part my mind starts wandering and I'm thinking about the next step, the next step, the next step. It's very hard for me not to do that linear flow — I want to get to the end of the problem, finish, and move on to something else. So I feel like having to multiplex and multitask between agents is a critical skill right now: stopping a job halfway through a thought process, switching to another active thread, continuing that, giving the agent feedback, and switching back. That kind of managerial multitasking is tough, and it also detracts from deep work. It's a very different skill set for the average programmer — programmers like to drill: get in flow, drill through the task, watch it complete. It's difficult to tell someone to do something halfway, go somewhere else, look at what they did, and catch up on the context they had. My Codex is a mess now — I have hundreds of threads, I don't know which ones are still open, and I mostly just use the priority list on the side that shows which ones I've spoken to most recently, because it's very hard to keep track otherwise. I also like to keep threads open for a long time — I feel like the compaction is good enough, and if I need something from earlier in the thread I can go back and look at it. So I have these endless threads, and I have no idea whether any of this is efficient. I sometimes feel it actually detracts a little from productivity to have so many different threads — you feel like you're working on many things, but you're actually working slowly on a bunch of things. So I don't know, there are pros and cons. I just wish everything would go faster and more smoothly, I guess — I guess that's a good thing for inference.

    2:03:17

    Nathan Labenz: Yeah, as you're describing this, the problem certainly resonates. I've got a dozen tabs open at any given time in a terminal dealing with Claude, and I'm constantly coming back going, what was going on here? I tried a few different things — they've built this in a bit more natively over time, but initially there wasn't great tab titling or summarizing, so I created a hook for myself that, every N turns — not every tool call, but every N tool calls — updates the title. So if I hover over it, I'll see what's going on, plus a little gray current-state text at the bottom. That helps, but I'm thinking audio might help even more — maybe I should have different voices for different threads, so I can use the part of my brain that understands how a single voice represents a single entity, and snap back into the context of the conversation I had with that voice. I'd be interested to see if that lets me context-switch a lot more effectively, because it's tough — when I go to a different tab, the overall look is the same: black text, gray text, whatever the most recent steps were, a little summary at the bottom, and it's like, okay, what the hell was going on here? Reorienting to that can be taxing. But if I had a voice where I just went, oh, I remember that voice and what we were talking about — I don't know how well that would work, but it could be a pretty effective little trick. So, as always, we'll point our agent at the transcript of this conversation and see if we can get some improvement out of it. That'll be interesting to experiment with.

    2:05:33

    Prakash Narayanan: So, to switch gears a little bit — Stripe just announced that they've closed the acquisition of OpenRouter. For those not familiar, OpenRouter is basically a service that supplies you with a single API key, and with that key you can use any model you want — you're not restricted. They have agreements with Anthropic, with OpenAI, with every single inference provider, so you don't have the switching cost you'd normally have. You embed the OpenRouter API or SDK in your app, and then you can switch very easily between services. And what they're saying — you probably can't see this well, but their announcement says: "The singularity — it's fuzzy, and perhaps an already overworked term, but we decided that January 1st marked the beginning of the singularity, and we have since been operating on that basis. In our case, we simply saw a large inflection in long-run trends — for example, a huge increase in the rate of new firm creation, and we decided that we ought to take that phase change seriously." Beyond that, they said: "Zooming out, we see capital and intelligence becoming the two digital flows undergirding every business. Up until now, every developer has needed a straightforward and reliable way to manage their revenue pipeline, and serving this need gave rise to Stripe. Going forward, every developer will also need a straightforward and reliable way to manage their intelligence pipeline. That observation first led us to OpenRouter, which has built the world's largest and most trusted token-routing engine... We've seen the parallels between managing intelligence and managing capital directly in our own products." So that's what they're saying. I'll take a step back here and note that I've often compared the entire AI space to a generalization of capital. I've often said that in the 20th century, China struggled with the idea of market capitalism, which meant ceding power over the economy to the market itself — an information architecture. AI is, in a sense, the information architecture of the internet, and it's ceding control of your information diet, and control of information, to the internet at large. That's going to be the big struggle of the 21st century, I think — this idea of human disempowerment is really about letting information flows drive decision-making across humanity as a whole, with models in the mix, and ceding that space. It's funny to see Stripe go in there and say, well, we have capital flows and we have intelligence flows, and we're going to acquire this company that's at the center of intelligence flows right now.

    2:09:18

    Nathan Labenz: They do have something real — I could see this being an incredibly compelling package for developers, because AI spend is a mess. That's definitely a real problem, and it's becoming a huge issue for a lot of companies. Getting really good clarity into what's going on spend-wise across a wide range of providers — you can get pretty good visibility from just a couple of providers if that's all you care to use, but to go beyond that, it's still pretty tough to make sense of it all. Stripe clearly has the DNA for this, and I think OpenRouter's also done a pretty good job of it already, but Stripe will have incredible strength in making that really clear and something you can optimize against — and also not have to rebuild. I did an episode of the podcast with — I forget her last name, Emily is her first name — who leads AI at Stripe, and she was talking about their payments foundation model. But one of the other interesting things she said was: just use Stripe as your database for everything financial. Don't try to mirror it — that just ends up causing more work for yourself. You're not going to be better than we are; the uptime is ridiculous; basically, you can't beat us at any of this stuff, so you might as well just build directly on top of it. I made that mistake myself at one point in the past, certainly when they were less proven than they are today, but in retrospect I should have just had API calls to Stripe be the way we got this information, rather than trying to store it all and dig around in our own database for it. I think they've really earned that trust, and when they can bring that to AI spend in a similarly trustworthy way, I think that will be a very valuable piece of overall architectures that people will find pretty compelling and naturally adopt. So I'd see them being able to have really good visibility into a pretty significant share of overall intelligence spend, because that observability and the robustness they've demonstrated is an incredible shortcut from a developer standpoint. No doubt about it.

    2:12:06

    Prakash Narayanan: It does strike me — number one, I feel like May to October is the seasonal AI doldrums. Every May to October we go through this "oh, AI isn't really working" cycle, and then just before Thanksgiving the companies do their last push for the end of the year and drop the product that's going to define the rest of the next year. So it's clear to me that, as a frontier lab, you don't want to be competing on token cost — the entire game is figuring out how you don't compete on token cost, and instead compete on the value you provide. There are a few ways to do that. One way that's emerging is probably speed — offering higher-speed services, meaning you go through more tokens but much quicker. You can see OpenAI heading toward that, because some people just want the job done, and it's a way to define that you're delivering enough value that someone's willing to pay more for it — same margin on the tokens, just dumping more tokens out faster. I think that's one way. The other, I feel, is going from tokens to digital employees. I've been talking about this for a while — this idea of a digital employee that you kind of hire, that's capable of doing multiple things you don't have to re-contextualize, with persistent memory. You can already see that in our experience with our personal Codexes and Claudes — they have some memory and some context on what we've been doing across multiple projects, but they need a lot more, I think. I wonder if that's the endpoint we're about to see in the next couple of months — I think OpenAI is nearing the release of Astra, and Anthropic is heading toward Fable 5.1, I hear, as the next release. Both, I think, should be out this month. But I feel very strongly that they have to stop competing on token prices, because that exposes them to direct comparison — you don't want to be in a situation where your API has to compete on evals for every single company separately, and the moment you're two percent down on the evals, they switch. That's really where OpenRouter is at. I've also had experiences with OpenRouter where things aren't what they seem — I tried using Baseten for GLM 5.2, and what I found was that Baseten doesn't do things in order. If you ask within a single call for name and address, it sometimes returns address and name — so if you generate a Chinese name, you'd expect a Chinese address, but it puts the address first, gives you an American address, and then gives you an American name instead. It'll do stuff like that. So I wonder to what extent these tokens are really fungible — I don't think they're even fungible now, and to the extent they are, they're going to be fungible on the lower-value stuff, like categorization and classification — things you could have coded with existing Python language tools, but instead you ended up using a frontier model for no reason, and now you're dropping from frontier to GLM 5.2. Is that really significant? I don't know. I feel like it's great that OpenRouter exists, it's great that multiple providers exist, but I'm not sure how meaningful they really are as competition to the frontier labs.

    2:16:36

    Nathan Labenz: Well, I think the frontier labs are doing just fine and will continue to do just fine — their margins seem to be improving, from the reporting that seems most credible to me, and their finances look great. And yet, my general rule of thumb for where tokens are fungible — not necessarily fungible, but where switching costs are low versus high — is: the narrower the task, the lower your switching cost, because you can actually define what you want and measure it. If you're in an environment where you're controlling the inputs through some means, then you can be pretty confident you can switch things over. If you're doing something like I'm doing on my laptop, where an idea comes to mind at any given time and I throw it directly into the model, I would not expect similar performance from anything other than the top tier. But I just did an episode with Flo Crevello from Lindy, and they're now offering Lindy Teammate, which is marketed exactly as you described — a virtual teammate — and it's powered by DeepSeek.

    2:17:59

    Nathan Labenz: And his whole thing was: you can't believe the amount of work we had to do to do this. We were ready for so long — we had unbelievable test suites and all the different use cases that are common for us. At one point, even with an earlier open-source model, they had determined that their evals had basically been passed. But then they launched whatever the alternative open-source model was at the time, and the response from their user base was, "Lindy got stupid — I don't know what happened, but it's stupid now." And they were like, oh, I guess our evals didn't cover as much as we thought. So they got to the point now where they're confident they can offer a virtual employee with a DeepSeek backing. He also told me they subsidize a lot of context ingestion and processing — a lot of what they do is in that initial onboarding, just sucking up all the information and trying to get ready to have the depth of context needed to actually do a decent job as a virtual employee. I thought it was pretty remarkable that they were able to get there with DeepSeek at all, because they have a lot of different customers and use cases. I'm sure there are still some corners of the platform where things aren't quite as good as they'd be otherwise, but the cost pressure is real — he was saying it's a small percentage of the cost compared to what it used to be. So for him to compete with Claude-tier pricing at all, he feels like he can't possibly do it with Claude as the model — he's got to have a different model, or it's just not going to work. So, yeah, I don't know.

    2:19:53

    Prakash Narayanan: It's such a similar story to Cursor — Cursor was buying Anthropic API, and then Anthropic started competing with them, and they can't compete with Anthropic while using Anthropic.

    2:20:09

    Nathan Labenz: Certainly not with the price discrimination that continues to go on. Flo's a very libertarian personality, but even he was into this — it's a pretty interesting framework I keep coming back to: what would the government do if it were trying to act like a tech platform? I credit two professors — Angela Zhang from USC is one, and her husband, whose name I'm forgetting at the moment — who are developing this thesis that the Chinese government has basically taken on tech-platform status, and all their companies are built on the social platform the government provides. And I'm thinking, what could we learn from that — what's the government-as-platform-with-American-characteristics that would make sense for us? One thing he was willing to endorse, despite being a pretty dyed-in-the-wool libertarian, was some restrictions on price discrimination by the frontier companies, to try to create a more level playing field for the app layer — because at a ten-to-one price discrimination ratio, it's just really hard for them to compete. So for now he's been forced to DeepSeek; if the price were the same, he could maybe come back. But—

    2:21:45

    Prakash Narayanan: That's the thing that has always annoyed me — Sam Altman does this thing where he says, oh, we had the API out for a long time, but until we did ChatGPT, no one else had thought to do it. And I'm like, dude, at the pricing you guys were doing — selling the GPT-3 API at that price — anyone doing ChatGPT at that pricing would have been hundreds of millions of dollars in the hole. And even when OpenAI itself did it, it was hundreds of millions of dollars in the hole, because it was a free research preview at first. So I think it's actually fairly disingenuous to say that. Even now, it's obvious to app-layer companies what kind of apps will succeed — the digital employee, for example, is something we know is going to happen. It also takes me back to Leopold Aschenbrenner's pre-"Situational Awareness" interview with Rakesh, where he said: you guys will just schlep, and after you schlep, we'll just have the next level of model, and that model will kill all of the schlep you did. I feel like that's really the ballgame — you have an expensive API and new capability, and a bunch of app-layer companies come up to wrap that capability. Then the API pricing starts to drop, and as it drops, the model company looks at which app companies have done well, Sherlocks the features it wants from them, embeds some of it in the model layer itself, does a little feature creation for what's not yet in the model layer — a little schlepping — and puts it out there. I'm not sure what the solution is, except that it's great we have multiple frontier labs — so even if we don't see competition at the app layer, even if the app layer is getting Sherlocked, the frontier labs themselves are fighting with each other and are forced to manage. I guess they're pricing at a level such that all of the GPU capacity they have is occupied — that's the pricing metric they use, all GPUs have to be occupied. And what ended up happening at Anthropic is all of their GPUs were occupied at the pricing level they'd determined, and they had excess demand. So they walked into SpaceX and bought capacity on SpaceX at ten times the price of market — from three dollars an hour for B200s to thirty to fifty dollars an hour for B200s. And they're still making money, because they're charging, like, eighty to ninety dollars per — I think that's like thirty x — so they have a ninety-plus percent margin, ninety-five, ninety-seven percent margin. And it still made sense for them to pay Google enough that Google decided its TPUs weren't for Gemini, given the insanity of the pricing Anthropic offered both Google and SpaceX. You know what?

    2:25:25

    Nathan Labenz: I think another thing that could be pretty interesting as an enabler for startups with this OpenRouter-Stripe combination is that it could become much easier to bring usage-based pricing to all sorts of products. This has been a huge challenge for a lot of startups — you want to give some kind of free trial, right? You're used to that in the no-marginal-cost SaaS game: sign up, try it free for X days. But when you have real AI cost attached to that, it's tough — so what do you do? Do you eat some of that cost and try to convert those users, or do you try to — what a lot of people are doing — figure out some subscription price point that works, but then you have some users you're losing money on and others cross-subsidizing them. That's definitely nowhere near an optimal state. You could imagine things getting pretty easy here, where all of a sudden you could say: hey, try our product for free, just pay for your token usage for the first seven days, and then you can subscribe at a low price that's still basically all margin for the app, and you still just pay for your token usage. That opens up a lot more opportunity. It's super gnarly to do, especially if you're running model changeovers and all that kind of stuff, but if Stripe can make that all super easy — where you can see with a simple indicator, this is the account using this inference call, and they handle all the billing to that user's credit card — I could see dramatic simplification, and a lot of different ways companies could present to customers that could be really advantageous.

    2:27:38

    Prakash Narayanan: So, speaking of exactly that — this morning OpenAI and Replit announced a deal. Replit's free tier is now powered by OpenAI's GPT-5.6 Luna. Luna is, I think, their Sonnet equivalent — their lowest model, fastest, cheapest. It's what is running our—

    2:28:02

    Nathan Labenz: I think Haiku equivalent.

    2:28:04

    Prakash Narayanan: Haiku equivalent — it's their fast, cheap—

    2:28:08

    Nathan Labenz: Terra — in between.

    2:28:09

    Prakash Narayanan: Exactly — and it's what is running our headlines, our chyrons. It's what runs our titling system. We have Luna to thank for personalized cancer earlier in the show. Yeah, yep. So, I guess the structure of this deal is Replit is paying kind of bulk pricing to OpenAI — OpenAI is basically selling pure inference, at open-weight inference-level pricing, but in bulk. So they're probably matching bulk pricing for inference, and they're managing the entire inference back end, and Replit gets to give them a subscription fee and have that as their free mode — and probably they get a revenue share of new subscriptions too, some combination like that. This is the way I think the frontier labs can differentiate themselves — you have this "OpenAI inside" kind of concept. I've always thought the next step is branding — once you have multiple competing products without clear differentiation, the next thing you do is branding. So this is going to be a Replit free mode with OpenAI inside, and I think that's going to be a thing. As long as they're selling — even the inference providers are making hefty margins — so as long as they can sell at the same price as inference for open weights plus some margin, it probably makes sense for them, and they can secure volume. So I guess we're going down this volume-and-brand-name game for the mass market — and, just as you pointed out, exactly as you pointed out, you have this kind of bulk pricing.

    2:30:11

    Nathan Labenz: Yeah, that's cool — I saw Michele from Replit say that Luna was the closest they'd seen to "intelligence too cheap to be metered," and now here they are: free Replit. Pretty cool. I've been a huge fan of Replit for a long time, even going back before AI coding, just because they've done such an amazing job creating a really accessible virtual machine you can just log into from your browser — the ease and accessibility of it, the ability to have a computer working in the cloud that you can share in a way that feels like a Google Doc. I think it's just amazing. But some of the friction I've run into with them over time is that they've had to figure out ways to convert people to paid, because they do have real costs, and so they've had some ways of limiting sharing that have definitely been suboptimal over time — I've given them some feedback on it. But this sounds like the beginning of maybe a new era where sharing can be a lot more free on the platform again — I'm sure they hated that limitation too, since it was definitely limiting their growth in some ways. But yeah, now if I can share with an account that has free Luna, so they can make whatever little changes they want to the app we're handing off, I'm again bullish on Replit.

    2:31:49

    Prakash Narayanan: A lot of developers are actually like, who uses Replit? And I'll tell you a story — I have a friend who's the CTO of a midsize ecommerce company, and he says all of his product managers now bring in Replit demos immediately for features. They don't ask the tech team to code features anymore — that used to be a three-month-long process of trying to get a developer internally to focus on your project and build out a feature demo. Now it's instantaneous — all the product managers use Replit, they build out the feature, demo it during the meeting, and you can immediately see it and decide whether or not it's deployable to production or not. That's what Replit has done, especially because they've integrated it inside the company's security stack — that's one of the major things. So it's really a tool for PMs to become developers, rather than developers developing on their own, which I think is the key differentiation they have.

    2:32:55

    Nathan Labenz: Yeah, it's a clear case of developers having — you know, I remember this poem from high school we were assigned, where the line was, "the burnt child loves the flame." I feel like that describes developers with respect to spinning up new development environments — somehow they've learned to love the incredibly painful nature of it. For everybody else, it's like, I just want to go to a URL and get something I can code in and share, and Replit really does have an incredibly good flow for that. It's confused me many times why developers haven't seen the value in it as quickly as I think they should. But — three Christmas presents last year that I coded on Replit. It's great even just for hobbyists: I made my dad a little app to test his stock-trading theories with some backtesting. I made my mom a travel app to plan her trips and find the gluten-free bakeries and whatever other idiosyncratic things she needs along the way. And I made my wife an event simulator for events she organizes, that literally has little people wandering around a grid that simulates the space and tries to figure out how she should optimize different parameters of the event to drive the outcomes she wants. All of that was vibe-coded on Replit, and you just hand it off to somebody and say, by the way, you can also extend this app — just prompt the Replit agent to do it, and you can make your own little tweaks. You're not fixed forever in what I've done for you. I think it's pretty cool.

    2:34:52

    Prakash Narayanan: Indeed. I'm looking forward to the voice getting better. Well, it's been another day in AI, and this week has worked well for us so far.

    2:35:08

    Nathan Labenz: Yeah, it's been fun — always cool to be improving. I'm glad we were able to bring Cue up on stage with us a little bit today. And maybe something else we can think about — the PR folks at Rand asked if there's a phone number she could dial into, and I said no, there's not yet, but it might be a prompt or two away. So—

    2:35:31

    Prakash Narayanan: —shouldn't be too difficult, I think. So let me—

    2:35:32

    Nathan Labenz: Iterative self-improvement continues.

    2:35:34

    Prakash Narayanan: Yeah, iterative self-improvement does continue. Indeed.

    2:35:36

    Nathan Labenz: Alrighty. See you tomorrow, Prakash.

    2:35:37

    Prakash Narayanan: Cheers. Bye bye.

    2:35:42

    Nathan Labenz: Bye for now.

A Good Morning in Biology

Nathan opened by calling it a great day in human history, and the opening block stayed on biology longer than planned. Prakash led with Anthropic's report that Claude, from a single human-written prompt, designed protein binders that succeeded in wet-lab validation — and noted that the announcement had been widely criticized by working biologists, whose objection was that Claude orchestrated several existing protein-specific open-weight models rather than doing the design itself. Nathan's response was that orchestrating specialists is what a specialist does, and that the critique had the shape of the one mathematicians made until the results became undeniable.

Prakash also flagged the detail that Anthropic had published the full prompt, including its instruction not to pursue biological warfare — which Nathan connected to Anthropic's constitution treating bioweapons as a hard rule that should hold even if the model concludes it is in a simulation.

The larger item was a Merck/Moderna combination of a blockbuster cancer drug with a personalized cancer vaccine, whose phase 3 trial was stopped early because the results made it unethical to keep the placebo group off the combination. Nathan connected it to his own son's immunotherapy — a trial also halted early for efficacy — and to the difference between targeting one surface protein and encoding north of thirty patient-specific targets. Against the roughly $50 billion in combined market-cap gain, he worked the arithmetic on consumer surplus and argued the value to society is likely one to two orders of magnitude larger than what the companies captured. Prakash's counterweight was that personalized vaccines are fundamentally a high-precision manufacturing problem, and that low-cost manufacturing at national scale is not what US pharma is organized to do.

1,179 Products, and the Manager Who Cannot Evaluate Them

Jessica Jensen and Jeremy Greenberg presented the AIDE Report — the first systematic look at what the AI-for-emergency-management market actually contains. The headline number is a census: 1,179 products from 717 companies, drawn from 1,892 screened. The findings underneath it are less encouraging than the count. Most of the market assumes conditions a disaster removes: the great majority of products require continuous internet connectivity, and only a small fraction run fully offline. Most companies do not publicly advertise 24/7 support. Almost none carry recognized AI management certifications.

The conversation kept returning to the buyer rather than the technology. A small-county coordinator facing a disaster in the next one to twenty-four hours has no time to evaluate tools, no way to combine data across them, and no procurement path that moves at the speed of the event. Prakash asked whether a Defense Production Act–style mechanism could compel API access at a set or post-disaster-negotiated price, so an agent could pull across tools at once. Nathan asked whether the answer is scale — FEMA operating like a tech platform, maintaining a library of ready-to-deploy tools for state and local teams — and then pushed it further, asking what a 'government as platform with American characteristics' would look like.

Getting the Turn Detector Out of the Audio Path

Justin Uberti co-created WebRTC at Google — the open stack underneath Google Meet, Zoom and Microsoft Teams — and now leads Realtime AI at OpenAI, where his team shipped GPT-Live. The architectural move he described is the removal of the turn detector from the audio path: instead of a cascade that waits for the user to stop speaking, transcribes, reasons, and then synthesizes, the media loop never pauses, and heavier reasoning is delegated to a frontier model asynchronously while the system is still talking.

Nathan pressed the obvious objection — that the bitter lesson eventually deletes hand-built architectural structure — while noting that evolution has not deleted System 1 and System 2 from humans either. Uberti's answer turned on latency being physically fundamental here rather than incidental. The conversation also covered which numbers he keeps in his head, whether the voice model has the full context of the reasoning model, misuse and emotional register as a differently-shaped safety problem in a continuous audio stream, accent and language mirroring, the cost mix between inference and telephony, and why singing still fails.

Cue, the show's on-air AI co-host, was brought up on stage and asked Uberti its own question: the hardest engineering tradeoff in moving from the cascaded pipeline to a full-duplex speech-to-speech system, and how to delegate to a larger reasoning model without disrupting the media loop. Nathan's framing to Cue was that it was on the line with its creator.

When the Model Underneath You Gets Cheap

The close ran on AI economics. Stripe closed its acquisition of OpenRouter, and Prakash read from the announcement — including its claim that the company had decided January 1 marked the beginning of the singularity and had been operating on that basis. Nathan's read was that spend visibility across providers is a genuine mess, that Stripe has the DNA to fix it, and that the combination could finally make usage-based pricing and token-metered free trials tractable for startups.

The counter-thread was what cheap inference does to the app layer. Nathan recounted Lindy's Flo Crivello describing the work required to run a virtual teammate on DeepSeek — and that he does not believe he can compete with Claude-based competitors while paying Claude prices. Prakash drew the parallel to Cursor being forced to build a model team once Anthropic began competing with it. The day's other deal pointed the same direction: OpenAI and Replit announced a free tier powered by GPT-5.6 Luna, which Nathan placed as OpenAI's Haiku-class model, with Replit paying bulk inference pricing — an 'OpenAI inside' arrangement Prakash expects to become a pattern.