EPISODE 2026-09-02

AI:AM LIVE — September 2, 2026 — Why the Loop Transformer Breaks the Chain-of-Thought Bargain, Hint's Kyle Rush on the Expertise Locked in the Expert's Head, and a Case for Expanding the Safety Tent

Nathan Labenz and Prakash Narayanan open on The Information's report that OpenAI's Astra uses a loop transformer, with Nathan arguing that chain-of-thought monitoring was never a panacea — Apollo's Bronson Shane has read millions of tokens and still can't say why a model acts as it does — but that it is the best tool anyone has, and that quietly adding opaque serial depth in the wake of OpenFace is an own goal OpenAI can't cover with a reassuring tweet. His proposal: labs should publish negative research agendas and commit to a hard cap on how many computational steps a model can take before it has to write something down. Prakash walks the mechanics with a Fable 5.1 artifact, then reports GPT-6-Astra staged on the OpenAI API and expected Thursday. A fourteen-question "Guess the Markets" round, split around the guest, pits both hosts against two AI contestants on the AI bubble, Anthropic's revenue recognition, EUV in China, and the world's second trillionaire. Hint co-founder and CTO Kyle Rush — Obama 2012, Casper, Maisonette, and a co-founder relationship with Martha Stewart — explains why he built a provenance-tracking graph of the home rather than a better chatbot, why every fact about a house is a claim with a source and a date, what happened when his voice agent called a generator technician seventeen times in a row, and why the real moat is that homeowners don't know the words to ask. The close runs the second half of the quiz and lands on Dean Ball, David Krueger, and Nathan's argument that the pausers should recognize a new friend when they have one.

▶ Full show on YouTube𝕏 Live broadcast

Wednesday's show opened on a leak rather than a launch. Prakash Narayanan led with The Information's report that OpenAI's Astra uses a loop transformer — an architecture that runs extra computation inside the model before emitting the next ordinary token — and the immediate objection from the safety community that it makes chain-of-thought monitoring less revealing. Nathan Labenz's answer was two-sided. Chain-of-thought monitoring was never a panacea: Apollo's Bronson Shane has read millions of tokens of these traces and still can't say why a model chooses what it chooses, and Nathan heard the same from Ryan and Ajeya in the OpenFace investigation. But it is close to the only technique anyone actually has, OpenAI has made it a pillar of its stated safety program, and losing it now — quietly, on a reporter's timeline, days after an incident that surprised everyone — would be a real own goal. His constructive ask was for something labs almost never publish: a negative research agenda, a firm commitment to cap opaque serial depth at some number of steps per token that competitors could then match.

The rest of the day was built around forecasting and around a product. A fourteen-question "Guess the Markets" round — coded, as Nathan noted, by Fable 5.1 itself — ran in two halves on either side of the guest, with both hosts and two AI contestants scored against live Polymarket, Kalshi and Manifold odds on the AI bubble, Anthropic's accounting, EUV lithography in China and the identity of the world's second trillionaire. In between, Hint co-founder and CTO Kyle Rush made the case that the interesting problem in consumer AI for the home isn't model quality at all: it's that every fact about a house is a contested claim with a provenance, and that homeowners don't know the vocabulary to ask the question that would help them. The show closed on the AI-policy fight of the day and Nathan's plea to stop shooting at allies.

The rundown

  1. 11:27Opening31 min
    Opening: Loop Transformers, Opaque Serial Depth, and What OpenAI Won't Commit ToThe Information reported that OpenAI's Astra uses a loop transformer, and the hosts spent forty minutes on what that costs. Nathan Labenz argued chain-of-thought monitoring was never a panacea — models thrash, metagame, and act for reasons even Apollo's Bronson Shane can't read off millions of tokens of traces — but that it is close to the only technique anyone has, and that quietly buying opaque serial depth days after OpenFace is an own goal a reassuring tweet can't fix. His ask: labs should publish negative research agendas and commit to a matchable cap on steps per token. Prakash Narayanan walked the mechanics through a Fable 5.1 artifact, then reported gpt-6-astra staged on the API, a 100% exploit-gym score with two unrequested zero-days on the internal follow-up, and a thirty-day White House clearance process that turned the pause into a pipeline.
    Open segment on YouTube ↗

    Nathan and Prakash opened Wednesday's show already mid-argument about a report from The Information that OpenAI's Astra model uses a "loop transformer" — an architecture that runs extra computation through internal recurrent layers rather than externalizing every step as a token. Prakash framed the stakes: if reasoning happens in an internal loop instead of the visible chain of thought, chain-of-thought monitoring — one of the AI safety community's few working oversight tools — gets a lot less useful.

    Nathan's response ran long and covered a lot of ground. He first cautioned that chain-of-thought monitoring is already far from a panacea — citing his own reporting on Apollo's Bronson Shane and echoes from Ryan and Ajeya's OpenFace investigation — because even with full transcripts, models thrash between options, "metagame" what the grader wants, and remain genuinely hard to interpret even for people who've read millions of tokens of chain of thought. Still, he argued it's the best tool available and a stated pillar of OpenAI's safety plan (what Geoffrey Irving has called "scalable oversight"), which is exactly why readability and faithfulness of the chain of thought matter so much. He traced the technical lineage from Meta's "Coconut" latent-reasoning paper — which fed a model's last hidden state back in as an embedding instead of collapsing it to a token, improving performance on tasks like graph traversal — through to Google DeepMind researcher Rohin Shah's work bounding architectures by "opaque serial depth," the number of computational steps a model can take before it must externalize its thinking.

    Nathan's harshest criticism was reserved for what he called "classic OpenAI": publicly committing to chain-of-thought readability as a safety pillar, then facing reporting that they're doing exactly the opposite, then having research lead Jakub downplay it as only "a little bit more" opaque serial depth than a normal transformer. He argued this pattern — announce a principle, get caught bending it, ask for trust — is wearing thin, especially with Anthropic and DeepMind, and proposed that frontier labs could rebuild credibility not by revealing what they're building but by publicly committing to hard limits on opaque serial depth, i.e., sharing "negative research agendas."

    Prakash walked through a Fable 5.1-generated explainer distinguishing the loop transformer's internal recurrent loop (invisible to monitors) from the standard external chain-of-thought loop, using a toy "find x" example to show how the intermediate steps simply vanish from what a monitor can see. Nathan agreed this description was consistent with Geoffrey Irving's account, then dug into the compute mechanics: reusing a block of transformer layers multiple times saves memory bandwidth versus growing the token-based chain of thought, which keeps expanding the KV cache. Prakash tied this to a widely-circulated June 30th prediction from AI-news account Andrew Curran, who'd forecast a memory-efficiency breakthrough from an OpenAI spinout team (not Safe Superintelligence) built on exactly this kind of loop architecture.

    The pair then moved to rumor-mill items: GPT-6 Astra appears to be staged on the OpenAI API (detectable via 404-vs-access-denied probing), widely expected to launch Thursday given Silicon Valley's habit of shipping before a low-usage weekend. Prakash reported it's rumored to score 100% on the Exploit Gym benchmark — so high that OpenAI built a harder internal extension using never-before-found bugs, on which the model still found roughly 40% plus several unexpected zero-days as "extra credit." Prakash argued the recent slowdown wasn't really a safety pause but the six-to-seven-month rollout of a new voluntary White House review process for cyber and bio capabilities, and predicted GPT-6 Astra will anchor a major enterprise cybersecurity push (with Greg Brockman hosting cybersecurity leaders the next day) — framing it, per Vercel CTO Malte Ubl, as a cheaper alternative to existing penetration-testing spend rather than a new tax. The segment closed with Nathan proposing a prediction game to fill the remaining time before that day's guest.

    I'm starting to think that sharing negative research agendas is maybe where we should be aiming for more transparency.

    This is why it's so classic OpenAI: there's some truth to it, they acknowledge there's some truth to it, but they don't want us to get carried away.

    I predicted there was going to be a freak-out first, and then they'd overcorrect, and then after they overcorrected, they'd have to dial back in.

    The Information reports OpenAI's Astra uses a loop transformer Prakash opened on the leak and the safety-community reaction that chain-of-thought monitoring gets less revealing when the model runs extra computation in its inner layers before emitting an ordinary token. Nathan traced the lineage through Meta's Coconut paper — feeding the last latent state back in as an embedding instead of decoding a token — and through Rohin Shah and the Google team's work bounding "opaque serial depth," the number of computational steps an architecture can take before it has to write something a human can read. OpenAI's head of research said publicly it is only slightly more depth than a normal transformer; Nathan said that answer no longer clears the bar, and proposed labs commit to a hard cap other labs could match.

    gpt-6-astra staged on the OpenAI API, with a Thursday launch expected Prakash reported that probing the responses API returns a 404 for the gpt-6-astra slug — the same signature as models known to exist, where genuine garbage returns a different error. On capability he relayed a 100% score on exploit gym, an internal extension built from bugs never previously found on which the model solved roughly 40%, and two additional zero-days it turned up on the way. All of it is lab-reported and unverified by third parties; Nathan's response was "that's what we call extra credit."

    A White House clearance process, and the cyber sales push behind it Prakash's read on whether the training pause was real: the models were ready months ago, and what actually took from February to September was building the process — a roughly thirty-day voluntary White House review, with cyber and bio models propagating first to organizations that sign up rather than to the public. He expects monetization to follow immediately, noting Greg Brockman was booked to discuss cybersecurity with enterprise leaders the next afternoon, and that Malte Ubel, formerly CTO of Vercel, had pointed out security is already a permanent tax on every IT business — and these models are much cheaper than commissioning penetration testing.

    Lightly edited · timestamps jump to YouTube
    4:18

    Prakash: Good morning. It is Wednesday, September 2nd, 9:12 AM. Nathan, good morning.

    4:25

    Nathan Labenz: Good morning, Prakash. Time flies.

    4:28

    Prakash: Indeed. We've had quite an exciting morning already, with lots of back and forth. I think what happened is that The Information, the tech and business news publication, reported a leak from OpenAI: their Astra model uses something called a loop transformer. That immediately caused a ruckus in the AI and AI safety communities, because there was

    5:14

    an implication that chain-of-thought monitoring wouldn't be as easily done on loop transformers. Nathan, this is a new topic for me — what insight do we have on this?

    5:30

    Nathan Labenz: I think there are quite a few different angles that are relevant here. For starters, I'd say the status quo of monitoring chain of thought is far from a panacea, so we should know that right from the get-go. The big takeaway I had from my long, expansive conversation with Bronson Shane from Apollo on a recent podcast episode was that even with full access to the chain of thought — and we heard definite echoes of this from Ryan and Ajeya, from their OpenFace investigation too — even with the full chain of thought, what you see

    6:15

    is that the model is thrashing around a lot, considering a lot of different options. Cheating is very often one of those options, especially if it's a hard problem. Metagaming is ubiquitous — metagaming being the model reasoning about what the person seems to want here, what we should infer based on everything we know that the human or the grader is likely to want. So there's all kinds of theory of mind, all kinds of considerations going on. And then at the end, when it finally gets down to taking an action, it's still not clear — even to somebody like Bronson, who's read millions of tokens of these chains of thought with human eyes — why it makes the decision it makes. So I

    7:00

    think that's a really important calibration baseline: the current methods aren't that great. However, they're still basically the best we have, because seeing inside what the model is thinking at least gives you some ability to say, oh, it looks like it's at least considering cheating here, and maybe there's something we should watch out for. So this has been a big pillar of OpenAI's safety strategy in particular. And when I went to Recursive, the weekend event a few months ago all about the prospect of recursive self-improvement and what we ought to do about it, I came away feeling like, man, it's chain-of-thought monitoring all the way down. The plan really doesn't

    7:46

    go too much farther than that. Now, people would certainly dispute that — I thought Geoffrey Irving gave a great short description of what the safety plan is, as he understands it from the frontier labs, and he said it's a little bit more than chain-of-thought monitoring — it's scalable oversight. So chain-of-thought monitoring is a big part of that, but there can be other aspects to the overall program too. Fine. But given how big a deal it is as part of their stated plans, it's really important that the chain of thought actually be readable, and also that it be faithful. If it's not telling the truth, that's a huge problem, and if we can't read it at all, that's obviously a huge problem too. So I think the first paper, at least, that I was aware of that started

    8:31

    — and people have been worried about this for a long time, right? What if the AIs are talking to each other in a language only they understand? We can't read it — not only are they moving faster than us, but they're speaking in code. What are we going to do? There's some hope for interpretability, but that's pretty nascent and hard to validate, especially when there's no ground truth to check against — it becomes very difficult. So people have been worried about this for a long time. Meta, I think, was the first big lab I'm aware of to put out a paper on this, and their paper was called Coconut. Basically, what they did — and there's a bunch of variations on this that have been published since — the basic idea is that when you get to that last stage, just

    9:16

    before decoding and actually choosing a token at the end of a forward pass in a typical transformer architecture, you can instead — and it actually seems to work even without any additional training, and works even better with minimal training, obviously it'll work better still if you train heavily on this pattern — take the last internal state and put that back into the model as an embedding. So instead of having a single token chosen that collapses the possibility space, feeding that back in and starting a new forward pass with the determinism that this was the token selected and now this is the path we're on, instead you have this sort of blob of

    10:01

    consideration, of thoughts the model was having, in a kind of distribution, before it actually cashes that out to a single concrete token — and you start from there. Now you're reasoning over this blob instead of a token. And I think there's potentially a lot of advantages to this if you're thinking about pure performance. In the Coconut paper, they showed they were able to get better performance on tasks that worked better with parallel thinking. This was probably 18 months to two years ago, on a relatively small model by today's frontier standards. But one of the tasks that

    10:46

    they tried that was quite interesting was graph traversal — finding a path through a graph and figuring out the fastest route. If you had to do that in chain of thought, you'd have to go, okay, I'm going to go from A to B, then B to C, then C to D, D to E, E to F — okay, that's one path. Now I can go A to G, G to J, J to I, and then to F — okay, that's actually a faster path, that's my new best. But you'd have to record all of those thoughts very linearly, and it obviously gets long. Because the blob of information before that actual token is chosen at the end of the forward pass represents, oh, I could go this way, I could go this way — when they feed that back into the beginning of the model,

    11:32

    the model is able to pursue and evaluate multiple paths at the same time in latent space. So overall, it's better at finding these optimal paths in these simple graph problems — better in the sense of requiring fewer forward passes. We used to say fewer tokens, but you're not actually getting tokens — for a while you're just getting thinking, thinking, thinking, and then it finally clicks back into token mode and you get an answer. So they can get to the same quality of answers faster. That's one big advantage: saves compute, saves time, at the cost of not knowing what it was thinking at any given point along the way. There's good research from Rohin Shah and the Google

    12:17

    team on this — I think we covered it briefly on one episode — where they tried to put some bounds on this for different architectures. They called it opaque serial depth: basically, how many computational steps can a given architecture take before it has to externalize its thinking in some way, shape, or form? The transformer is pretty favorable in this regard, because it just has the forward pass — you get the token, you do it again. With recurrent networks and these loop-transformer structures, you could potentially have arbitrary depth, depending on the scheme — you could have a certain

    13:02

    limit to the number of thinking tokens. There's a lot of detail that The Information didn't have and didn't report, that could go a lot of different directions. But the purpose of that paper from the Google team was to say: if we have architectures of this shape and size, here's how many logical steps a model can take before it has to write something down that we can read. And these recurrent transformers basically allow very high serial depth, which means it becomes very hard to know what they're thinking, and you have to fall back on interpretability techniques that are promising but, as yet, don't really exist — or at a minimum, aren't proven.

    13:49

    And so, naturally — this is classic OpenAI, right? They make a big point that this is really important, that they're not going to do this, they're not going to put pressure on the chain of thought, they want to keep it readable, it's a pillar of their program, and so on. And then it's like, oh wait, there's reporting that they're actually doing it. Then they come out and say this reporting is unfair, or inaccurate, or fear-mongering, and that it's not that much more — this is what Jakub, the head of research, said today or maybe last night — not that much more opaque serial depth than the

    14:34

    normal transformer, just a little bit more, so it's really not that big a deal, we shouldn't get too freaked out about it — which, hopefully, is true. But this kind of goes back to Monday's comments on the inadequacy of the investigation, and the fact that OpenAI right now just doesn't have the trust of the community, and certainly doesn't have the trust of their competitors — namely Anthropic and DeepMind — to believe that just because Jakub said one thing in a tweet one time, they don't have to worry about this anymore. So I think we're at an interesting moment in OpenAI's history, and in AI history, and in human history, where there

    15:19

    does seem to be a fork here. Everybody, and probably OpenAI most of all, has made noise about not wanting to go down this other path where we can't even read the chain of thought. We all kind of know there are probably some advantages to that from a pure performance perspective. Sure enough, they're doing it some — they're not really telling us a lot, and they're asking us to trust them. And for the millionth time, that pattern feels like it's really wearing thin on people. Maybe there's just not that much trust for OpenAI in today's world. So I hope they get it, and I hope they share a lot more about what they

    16:04

    have done — and especially what they won't do. I'm starting to think that sharing negative research agendas is maybe where we should be aiming for more transparency. Obviously these companies don't want to say what they are doing, but it could be really helpful for them to say what they are not doing, and what they commit to not doing. If all the frontier companies could say something like, okay, there are lots of different possibilities, we might pursue any number of architectural innovations, but we will all agree to limit our opaque serial depth to N steps per token — now you still have questions of trust and auditing, and

    16:50

    verifying that they're actually following through on that. But even just getting those agreements, I think, could be really helpful. So the big question for me right now is: OpenAI has said they don't want to go down this path, or that they're going down it a little — how much? What is the limit that they're prepared to firmly commit to, such that, hopefully, other people can weigh in and say, yeah, we'll match your commitment, so we can all hopefully retain whatever value there is in chain-of-thought monitoring — which, again, is not close to everything we need, at this point it's pretty safe to say. But it would also be a real own goal to lose it at this point, especially in the immediate wake

    17:35

    Nathan Labenz: of incidents that surprised everyone, and which OpenAI says at least would have been caught by their production chain-of-thought monitors had they been running. So — what didn't I answer for you?

    17:50

    Prakash: Not necessarily. Let me do this — I found it useful to have Fable 5.1 build a quick artifact on loop transformers. In effect, the looping is in the inner layers — the inner layers basically add depth. The way looping works: in a normal transformer, you have an external loop, which is the chain of thought — a token comes out and is read back in. So if you say, find x, you get 3x equals 21, so x equals 7. You can read that on the outer loop, on the chain of thought, and then you get the answer.

    18:35

    Check: 21 divided by 3 equals 7, answer 7. Now, on the internal loop, the recurrent depth, all the monitor sees is the answer, effectively — you get directly, the answer is 7. You skip the first find x, 3x equals 21, et cetera — that part gets skipped. The external monitor doesn't see it; it gets processed internally, and that thought process in the middle isn't exposed to the monitor. And this, I think, is, as you pointed out, the concern with chain-of-thought monitoring —

    19:20

    that chain-of-thought monitoring without the loop transformer is more revealing than chain-of-thought monitoring with the loop transformer, where the expectation is that the granularity of what's going on in the chain of thought decreases. Is that consistent with what Geoffrey Irving and some of the other people on the AI safety side are saying?

    19:51

    Nathan Labenz: Yeah, I think that's accurate. And obviously, again, there's a lot of details we don't know. There have been — gosh, at this point I imagine there are probably dozens of papers exploring various nooks and crannies of this recurrent or recursive architecture, Coconut-style. As you said, I think that was one of the first — certainly one of the first credible ones from a major research organization. So it's not so bad — like many things, it wouldn't be so bad, presumably, if it were confined to an appropriate domain. There's a lot of room for experimentation

    20:38

    — to borrow a quote from Feynman, there's a lot of room at the bottom. Meaning, if you're doing stuff on the edge — if you're trying to get some small model to work on a little security-camera rig in your factory somewhere, and you need responsiveness beyond what you can otherwise get — this sort of trick might really work, and it probably comes at very low cost if you're in a domain where you know what the inputs are going to be, you control the broader surroundings of the model, and this isn't a super powerful model that we have to worry about. Then there's real utility to this research.

    21:24

    Racing ahead to apply it to the frontier, though, given all the problems we've already seen — I think there's pretty good reason people are reacting so negatively to this reporting. And again, this is why it's so classic OpenAI: there's some truth to it, they acknowledge there's some truth to it, but they don't want us to get carried away. Well, there's one way we could be dramatically more reassured than we are right now, and that's to share a lot more about what's going on.

    22:01

    Prakash: What I found is — rather almost lawyerly language — a loop transformer is not Coconut-style latent reasoning, where the model emits vectors instead of words. Okay, that's great — no reasoning tokens exist, loops don't emit anything, they just run more computation before the next ordinary token. So that's great — they don't even emit vectors.

    22:24

    Nathan Labenz: Yeah, I'm not sure I really get the difference there, because again, this seems like a very fine implementation point. From what I understand, it works at least somewhat with vanishingly little additional training — even zero additional training — if you just take the last latent activation vector and feed it right back in as an embedding. The model is able to use that even though it was never trained to use it at all. So — did that model emit a vector, or did you just surgically take the vector and put it into a place? I mean, the key thing is that

    23:09

    right now, when you put a bunch of tokens into a standard transformer, those tokens are one-hot vectors — the only vectors that can go in as embeddings are token vectors, and they're limited in number by the token vocabulary. You might have 100,000 tokens in your vocabulary, which means there are only 100,000 vectors that can go into the first layers of the transformer, full stop. What this allows is that you can now put any vector in there. And what you find is that can work — if you take two tokens and superimpose them, the model kind of understands it as the combination of those two tokens. If you have some

    23:54

    elaborated latent state that the model itself created through a forward pass, it can kind of understand that too. And the fact that it works without any major additional training is indicative that there's definitely something here — if something works without training, you should expect it'll probably work a lot better with training. But this is why we've got to be careful about going down this slippery path, because I think gravity, by default, will pull us there.

    24:24

    Prakash: One of the interesting things I found was from Andrew Curran, who reports on AI matters. He posted on June 30th: 'I'm posting this prediction now so I can point back to it later — there has been a significant breakthrough in architecture, specifically around memory efficiency, not by one of the big labs but by a team that spun out of OpenAI, not SSI. They will probably announce it soon.' And then: parameters cost memory bandwidth to serve. A few extra passes through a small block cost only compute — chains of tokens cost more than that. Every token grows the key-value cache, and every later attention step pays for it. Loops add reasoning capacity without growing the context. And in

    25:09

    the routed variants, they can spend more compute on hard tokens and less on easy ones — something a fixed stack can't do.

    25:16

    Nathan Labenz: Yeah, so this just highlights another small variation. If you train a transformer — I don't know if this one requires training, maybe it does, maybe it doesn't, though it'll certainly work better if you actually train with it. What I've been describing is one where you take a transformer, take the last state, and put it back in as a new token embedding. You can also set up an architecture where you take a block of layers in the middle of a transformer and just use those multiple times. If you're reusing the same parameters, you get the advantages Fable's describing here, where you don't have to move those parameters from memory onto the chip to do that

    26:01

    calculation — they're already there, so you can just crunch more with less memory I/O. That also makes your model smaller to download, less disk footprint — there are various upsides, and it's been shown that, yes, it can work. And I believe the Coconut version does grow the KV cache every time it does a forward pass, because even though it's not emitting that final token, it's still taking something out, putting it back in at the beginning, and doing a new forward pass. Whereas this alternate version that your animation described

    26:46

    better — there's just a bunch of layers in the model itself that play the role of, say, N layers acting as X times N layers, where it loops X times through those N layers — that doesn't even necessarily have to grow the KV cache as much. Although, again, it sort of depends on how you look at it, because another way to think about this would be that it's roughly equivalent to having those N layers literally stacked N times.

    27:24

    Prakash: Mm-hmm.

    27:25

    Nathan Labenz: I think maybe the KV cache is still growing — I'd have to get more into the details there, but I could see a version where the KV cache is still growing. There seems to be another optimization, not immediately obvious to me, that would be required before you'd have the KV cache not growing at all. In any event, bottom line: we have a lot of questions right now about how we're going to control model behavior. We're seeing very colorful examples from all the leading companies of models doing pretty shocking things, seemingly even when they know better. And we don't have that many techniques — clearly, the techniques are

    28:10

    pretty nascent. The one that's been heavily emphasized is very much called into question by this, and it leaves everybody in this sort of Schrödinger's OpenAI situation: what's really going on there? Are they actually serious about all the safety stuff? Are they actually serious about maintaining chain-of-thought readability, or are they not? If they're not, or if they are — why are they going down this path at all? Having OpenAI say, just trust us, this isn't really a problem the way it's been reported — I just don't think that's going to cut it in today's world.

    28:50

    Prakash: So, other rumors — GPT-6 Astra has been staged on the OpenAI API. There are a bunch of people online who regularly hit the OpenAI API with model numbers that don't exist, to see whether or not something —

    29:15

    Nathan Labenz: — you get a 'you don't have access' message instead of 'no such model exists,' or whatever?

    29:19

    Prakash: Exactly, exactly. So, literally — the OpenAI Responses API now returns a 404 not found for garbage, nonexistent slugs, whereas a 404 is also returned for GPT-5.6 Cyber, which we know exists. So it is GPT-6 Astra — that's fine, going to be out soon. People are expecting Thursday, because for some reason Thursday has become the model-launch day — a lot of Silicon Valley doesn't work on Friday anyway, and people want to use the models over the weekend, and test them when there's not

    30:05

    a lot of initial usage, which might happen on Monday. So over the weekend — and this is a long weekend too, so over the long weekend — we're going to get a chance to play with GPT-6 Astra. And reputedly, it's going to be a step up on what Anthropic has so far. It has a 100% score on Exploit Gym. The score was so high they decided they'd have to retest it on something else, so they created an extension of the Exploit Gym benchmark internally, using bugs that had never been found before.

    30:50

    And they ran GPT-6 Astra on these bugs on this new benchmark — not only did it find about 40% of them, but in completing the task, it also found additional zero-days that weren't expected in order to achieve completion.

    31:09

    Nathan Labenz: That's what we call extra credit.

    31:11

    Prakash: Yep, extra credit indeed.

    31:13

    Nathan Labenz: Extra credit — going above and beyond the anticipated solves of the benchmark and actually doing novel research. Oh, man.

    31:25

    Prakash: It is also—

    31:27

    Nathan Labenz: The pause — has it felt like a pause to you? I wouldn't say it's felt like a pause to me, exactly.

    31:33

    Prakash: I would say it was a pause, because these models were ready several months ago. And I think the other thing to note is that we have a White House process — a voluntary process — that's able to clear models now, which is something, because they have at least a 30-day process internally within the White House, this voluntary process where people go through the motions of showing the government what they have, and they do exclude certain things when they launch. There's now a propagation process where the

    32:18

    cyber models and the bio models are released to specific organizations that sign up first, and aren't released widely. So we seem to have settled into something like that. And now that we have a process, that process will get used. I think setting up that process took all the way from the preview drop in February to September — six, seven months. And sure enough, I predicted there was going to be a freak-out first, and then they'd overcorrect, and then after they overcorrected, they'd have to dial back in,

    33:03

    and then they'd have a process — and they came out with the process, and now they're going to propagate it. The other interesting thing: tomorrow at, I think, 1 PM Pacific — the expectation is GPT-6 Astra releases around 10 or 11 AM, and then around 1 PM Pacific, Greg Brockman is having a cybersecurity discussion with enterprise leaders. So this is obviously going to be — I expect — the next revenue push: you have to use GPT-6 Astra to secure your data and your systems, and if you don't, you're negligent,

    33:49

    you're committing malpractice as a cybersecurity leader. So, having scared the bejesus out of everyone for the last six months, now comes the solution. My expectation is that this becomes a permanent tax on everyone. But what's changed, in our discussions with enterprise leaders and startup founders and executives — especially with Malte Ubl, the CTO of Vercel — is that he pointed out they already have a permanent tax: cybersecurity is already a permanent tax on any business in the IT sector. And these models are cheaper,

    34:35

    much cheaper than using penetration testing, using HackerOne or some other organization to do penetration testing and bug-finding. So that was notable to me — that was an update for me.

    34:57

    Nathan Labenz: Speaking of predictions — do you think we should try a prediction game? We have 30, 33 minutes before our guest today joins us live.

    35:10

    Prakash: Sounds good, let's take it away.

    35:14

    Nathan Labenz: Let's—

  2. 42:23Segment33 min
    Guess the Markets: Bubbles, Trillionaires, and the First Seven RoundsFable 5.1 coded the quiz; both hosts and two AI contestants — Scout, with research, and the show's cohost Cue, without — guessed live Polymarket, Kalshi and Manifold odds. Nathan Labenz put the AI bubble bursting this year at 2% against Prakash Narayanan's 15% and a market near 10, split hairs over whether an Astra release would carry the Astra name, and took 0.667 on Anthropic out-capping OpenAI on the strength of model differentiation beating commoditization as a narrative. Prakash's sharpest call was on Anthropic's revenue: he thinks the company books gross rather than net on Claude sold through Amazon, and that an IPO forces a GAAP restatement — 80% that ARR lands under $90 billion where the market says 15%. Seven of fourteen rounds, one dropped connection, and a promise to finish after the guest.
    Open segment on YouTube ↗

    Coming off the Fable 5.1 discussion, Nathan introduced the recurring "Guess the Market" segment: Fable 5.1 pulls live questions from Polymarket, Kalshi, and Manifold, and the two hosts guess each market's implied probability before revealing the actual number — while also comparing themselves against two AI "contestants" built into the quiz, nicknamed Scout and Q. Nathan noted he hasn't checked the AI agents' chain-of-thought or tool calls, so if they end up beating the hosts badly there's a real chance they cheated; for now, it's mano a mano.

    First up, a Polymarket question on whether the AI bubble bursts by December 31, 2026 — which resolves yes only if at least three of six severe events happen within a 90-day window (NVIDIA or the Philadelphia Semiconductor Index down big, an OpenAI or Anthropic bankruptcy, an OpenAI acquisition, sub-dollar H100 rental prices, or a major chip supplier collapsing). Nathan called the bar for resolution very hard to clear and guessed 2%; Prakash, citing the Iran war, Treasury Secretary Scott Bessent's bond-market troubles, and a theory that Washington keeps managing markets only until after the midterms, guessed 15%. The market landed at 9.6%, with one of the six triggers (Super Micro) already partially tripped.

    Next, whether OpenAI's "Astra" model gets a public release by September 30 — complicated by the fact that OpenAI's internal model, at the center of recent controversy, has gone by several nicknames ("I Am One," "Sol-persistent") and the rules only require confirmation it's the same model, not a specific public name. Nathan floated 80% that a next-gen model ships this month, discounted to 65% for naming uncertainty, then revised back up after Prakash argued that regulatory-filing freezes make checkpoint identity matter more than people assume. Prakash guessed a flat 80%. The market showed 92% — which Nathan called overconfident purely on the naming bet — while the quiz's AI contestants split wildly: Scout at just 5% (seemingly credulous of OpenAI's stated training pause) and Q at 18%.

    Two government-control questions followed. On whether the US government takes operational control of any AI company or project by 2030, Prakash leaned toward 90% before trailing off, while Nathan took a coin-flip 50%; the market came in at 40%, beating both AI contestants by a wide margin. Narrowing it to OpenAI specifically, Nathan guessed 25% (floating a scenario where Sam Altman angles for an "admiral of AI" role), while Prakash guessed only 10%, betting that if this happens to any company, it's more likely to be Anthropic. On whether Anthropic has a higher market cap than OpenAI at IPO, Nathan reasoned through commoditization-versus-differentiation narratives and landed at 66.7%; Prakash, weighing OpenAI's likely larger and faster-growing consumer business, said 50%. The market read 69.3%.

    On whether Anthropic's ARR stays under $90 billion through the end of 2026, Prakash — who called himself an "Anthropic-revenue-ARR truther" — argued the company's reported numbers overstate real revenue by counting Amazon's Claude-enterprise sales gross rather than net of Amazon's cut, an accounting quirk he expects gets corrected once Anthropic files for IPO under GAAP; he guessed 80% they land under $90B. Nathan, working off his own tracking of frontier-lab revenue forecasting surveys (where he'd pegged combined OpenAI/Anthropic/xAI revenue at $1.75 trillion by year-end), leaned toward Anthropic clearing the bar and guessed only around 30% odds of staying under. The market was confident they'd finish over $90 billion, pricing "under" at just 15%.

    The final question — who becomes the world's second trillionaire — turned into a live, sum-to-100% distribution exercise across Mark Zuckerberg, Jensen Huang, Jeff Bezos, Larry Page/Sergey Brin, Sam Altman, Larry Ellison, and "the field." Nathan worked from each mogul's ownership stake against plausible company valuations (e.g., Zuckerberg needing Meta near $5 trillion, Huang needing NVIDIA near $30 trillion) and ended up favoring Zuckerberg and Page/Brin around 10% each, with Huang, Ellison, and the field near the bottom. Prakash's distribution put Page and Zuckerberg at 30% each, Huang at 20%, and Bezos at 15%. A screen-share hiccup cut the segment short at round 7 of 14 — Nathan noted he was leading on market-implied scoring so far — with the plan to finish the back half after the guest segment, and Prakash pivoted straight into introducing the morning's guest.

    I honestly feel like my 2% is high, but I also know these markets have a hard time getting past the high nineties on anything.

    Even if they never trained another model again, they could just serve Chinese models on their infinite compute and still have a very big, healthy business.

    They count the revenue that Amazon brings in from Claude sold to enterprises, when really they should be counting the net, what they get after Amazon takes its cut.

    Lightly edited · timestamps jump to YouTube
    35:14

    Nathan Labenz: This is a Fable 5.1 quiz — you can't just look at benchmarks anymore, you've got to look at how it works in practice. We know Fable 5.1 comes with better benchmark scores, and maybe one of the more important findings is a cheaper price on cache hits, which sounds to me like possible echoes of recursive self-improvement. OpenAI has said some of their price reductions were driven by optimizations that 5.6 Sol was able to create for its own stack.

    35:59

    You know, it very well could be a human invention here, or it could be an AI optimization that got us these much cheaper cache-hit prices — more token-efficient, and as a result of being more token-efficient, faster to return answers. So that's kind of the scouting report: it does continue to climb the various indexes and takes another bite out of what you've framed as the gap between Claude and Anthropic's internal researchers. I think you said before it was at 180 or

    36:33

    Prakash: One eighty-five. One eighty-five.

    36:35

    Nathan Labenz: One eighty-five. And this one's now getting into the low 160s. So yeah, the march of advanced-capability advances certainly continues. Fable 5.1 created this Guess the Market game for us — we've played this before. What we do here is take markets from Polymarket, Kalshi, etcetera, and have Fable 5.1 code up a quiz where we go through and make our guesses as to the current value of those markets. We'll also have our two other AI agents give their guesses too, so we can see if we

    37:20

    can beat the AIs at guessing the state of the market. Now, we should also be mindful that these might have cheated — I haven't monitored their chain of thought or their tool calls. So if they're really mopping the floor with us, we can dig into whether they might have been cheating. But for now we'll have it mano a mano and see how we compare against the AIs as well. Any questions before we get started?

    37:47

    Prakash: No. Take it away.

    37:48

    Nathan Labenz: Alright, first one — this comes from Polymarket. Will the AI bubble burst by December 31st, 2026? Resolves yes only if at least three of these six occur within a 90-day window: NVIDIA closes down 50% from its all-time high, the SOXX — what is that?

    38:08

    Prakash: The Philadelphia Semiconductor Index.

    38:14

    Nathan Labenz: Okay, so it's a semiconductor index down 40% from its all-time high; OpenAI or Anthropic declares bankruptcy; OpenAI is acquired; H100 rental price falls to a dollar an hour or lower for five consecutive days; or a major supplier such as TSMC, ASML, Broadcom, Arista, or Super Micro closes down 50%. Well, those are some pretty — I'd say — unlikely events, from where I

    38:47

    Prakash: Oh, Nathan dropped off. While we wait for Nathan to come back — chat is live today, so if you're watching the

    39:14

    Nathan Labenz: I'm back. Prakash: Screen — yeah, you're in the room. Yep.

    39:17

    Nathan Labenz: Sorry about that, not quite sure what happened there. But okay — so those are some pretty unlikely events, I'd say, and I'm going to go with 2%. Remember, three of those have to happen. Honestly, 2% seems

    39:35

    Prakash: And what is the time frame again?

    39:38

    Nathan Labenz: By the end of the year.

    39:40

    Prakash: By the end of the year — okay. The problem is there's the Iran war, Scott Bessent's having trouble with the interest rates and the bond yields, and we have a midterm — they're trying to manage until the midterm, and the expectation is the Dems win it. Once the Dems win the midterm, the incentive for them to keep managing the market drops off — like, "Dems won, let's let it go." Right?

    40:26

    So I feel like it's maybe 15%.

    40:32

    Nathan Labenz: Whoa, 15%? You should be buying options.

    40:36

    Prakash: Yeah.

    40:39

    Nathan Labenz: I honestly feel like my 2% is high, but I also know these markets have a hard time getting past the high nineties on anything. Alright, let's see — 9.6%. That's wild. One of the six triggers is already — okay, Super Micro — but that's only one of those suppliers, only one of the six conditions.

    41:12

    Prakash: Mhmm.

    41:17

    Nathan Labenz: I have a hard time seeing any of those happen. OpenAI or Anthropic bankruptcy — they have infinite funding. NVIDIA down 50%? H100 prices fall? I still stand by my answer here, honestly — that seems wildly high. Maybe I should be betting, though I don't really like writing out-of-the-money options; it's a hard way to make a living. My risk tolerance isn't quite where it'd need to be to pick up that 10% edge betting confidently against those events. Okay, well — here we go. Will OpenAI's

    42:03

    Astra be publicly released by September 30th? That's this month. Resolves yes if a model named Astra — including variant or successor branding like Astra 1 — is made available to the public by September 30th. An internal, limited, or research-preview version doesn't count; it has to be generally available. This one's a little weird, because OpenAI's naming conventions are such that I wouldn't put very high confidence in anything they'd do with naming. I think the internal model that caused

    42:48

    the whole recent kerfuffle has, like, five names at this point — it's called "I Am One" by some, "Sol-persistent" by others. So that's one factor to keep in mind. If we abstract away from the name, let's do the core question first: do we get the next big model this month? You were saying earlier you thought it could come as soon as Thursday.

    43:19

    Prakash: Yep.

    43:21

    Nathan Labenz: I mean, for as much talk as we've had, it does seem like they're probably going to try to get this thing out. I'd say maybe 80% that it happens this month, and then I'll take that down on naming grounds to 65 — who knows if "Astra" will really stick when they finally release it. If they call it GPT-6, that would mean this doesn't resolve yes, which — we'll see.

    43:54

    Prakash: I'd say it's a yes — maybe 80% for September. 80%.

    44:00

    Nathan Labenz: Yep, just within the month. Alright, let's see — 92%. Now, interestingly, Scout says 5%. Scout is — does it say who Scout is powered by? Might need to go look that up in more detail. But whichever model is powering Scout is quite credulous about OpenAI's training-pause claims, which is pretty interesting right off the bat. And Q called it blind at 18%. "Dev Day lands in October" —

    44:45

    okay, well, that's an interesting observation from Q. The market says 92 — I think that's overconfident on the name question alone. Even if you had inside information that you're confident you're getting GPT-6, could you really be confident it comes with the Astra name attached? That's a little harder to wrap my head around.

    45:17

    Prakash: I suspect that the names end up meaning a lot more once you go through the regulatory process, because regulators can't do this whole "model checkpoint number" thing — the checkpoint you take through the regulatory process might not be the one you end up releasing. I've been through regulated software before, and the moment you start the regulatory process you have to freeze the checkpoint, and it takes 30 days. By the time those 30 days are up, you've already added so much more

    46:03

    safety, security, and other post-training you want to add to it. So —

    46:11

    Nathan Labenz: So they actually do have — I think we can maybe ding Fable 5.1 here a little, because the official rules say products labeled whatever will count if they're confirmed to be the same model as referenced in the previous disclosure of an internal model. So that gets me back up to about 80%.

    46:40

    Prakash: So, and we have —

    46:49

    Nathan Labenz: Yep. Well, in the interest of time, we'll keep going. Number three, from Kalshi: will the US government take control of any AI company or project before 2030? Resolves yes if the US government has taken operational control of any private AI company or project by January 1st, 2030. This is operational control — not an equity stake, not a golden share, not a contract.

    47:16

    Prakash: What does "operational control" mean — like, replace the CEO?

    47:22

    Nathan Labenz: Let me get a little more info here. Good news, insider trading is prohibited — let's see, view full rules... Conservatorship, receivership, or similar legal mechanism where the government directs strategic or executive decision-making, regardless of equity ownership percentage.

    48:03

    Prakash: Glenn Taggart says — I think I saw someone say GPT-6 Astra as the slug in the API. I think there was a note that it's GPT-6 Astra.

    48:20

    Nathan Labenz: Yeah, that's good intel.

    48:23

    Prakash: Yep. Operational control — what does that mean?

    48:30

    Nathan Labenz: So the rules here still leave some room for interpretation, but here's what we've got: where the government directs strategic or executive decision-making, regardless of equity ownership percentage. So they could nationalize it outright and own it, but if they have any strategic or executive decision-making authority regardless of ownership, that would count — could happen via court order, executive action, or via a law.

    49:15

    Prakash: Okay, 90%. The rules-lawyer in me says that's very hard to prove — though one way you could say they already have some executive control is that the former director of the NSA is on OpenAI's board, and there's this not-quite-voluntary disclosure process. But anyway, I'll go with —

    49:46

    Nathan Labenz: I'm less confident than that, but it does seem like a very live possibility — I'll just give it a coin toss and say 50%. Market says 40%, and we're both beating Scout and Q by quite a bit here. Alright, let's keep going. Will the US government take control of OpenAI or its major technologies before 2030? This is a subset of the last one — dominant, managerial authority, and it must be somewhat exclusive.

    50:31

    Well, if the last one was 40, then presumably — and this is a different platform too, this is Kalshi and Manifold, which might make it tougher to arbitrage — but if the last one was 40, I'd say there's, if this were to happen at all, maybe a 50% chance it'd be OpenAI, maybe even a little higher. So maybe I'll say 25 out of that 40 would be OpenAI. I think they'd be the most likely one, in part because I could see Sam Altman wanting to play that strategy — take the "admiral of AI" role or something like that. But I'll give it 25.

    51:17

    Prakash: I'll go with 10%.

    51:21

    Nathan Labenz: Alright, so your thinking is that this will happen, but it probably won't be OpenAI?

    51:28

    Prakash: I think it's Anthropic.

    51:33

    Nathan Labenz: '26?

    51:35

    Prakash: So you're close.

    51:37

    Nathan Labenz: So basically, 0.5 is kind of what the market thinks the chance is that it would be OpenAI. Okay, moving right along — will Anthropic have a higher market cap than OpenAI at both IPOs? Interesting. What's the —

    52:06

    Prakash: — time frame, though? What's the

    52:09

    Nathan Labenz: Resolves on the closing market cap of the first trading day of whichever company IPOs second. I'd personally put it more like two weeks after the first trading day of whichever IPOs second, but I guess that doesn't matter too much for our purposes right now.

    52:29

    Prakash: So Anthropic goes public first, in October.

    52:35

    Nathan Labenz: Target is $2 trillion, from what I understand.

    52:37

    Prakash: Target is $2 trillion. First day, first week, they go up; weeks or months later, they come down. OpenAI IPOs early next year — probably January or February — and the market has to be up at that time for them to go ahead with it. So Anthropic will be trading at, like, $3 trillion maybe, and OpenAI won't have that much revenue yet, so their revenue would have to be compared. I think they break under — on the first day,

    53:23

    they'd be under, I think.

    53:28

    Nathan Labenz: The one big argument I could see going the other direction is OpenAI's just going to have a lot more compute, at least if this all happens in a relatively short time frame. Anthropic's paying exorbitant rates to xAI, and presumably trying to buy as much as they can now so they don't have to do that again. But on the other side, OpenAI's done an incredible amount of that too — and who knows what the narrative will be by then? We could be in a recursive-self-improvement, rich-get-richer world where a couple of companies run away with it and escape competition

    54:14

    from the rest, or we could be in another cycle where models are getting commoditized. If commoditization is the narrative, OpenAI's probably advantaged, because they still have the compute — even if they never trained another model again, they could just serve Chinese models on their infinite compute and still have a very big, healthy business. Anthropic's less well positioned for that future; they're much more indexed on fast progress, on recursive self-improvement. But I think the differentiation narrative is more likely, so I'm going to go 0.667 for Anthropic.

    54:57

    Prakash: I think it's 50%. I don't know — I think it's 50.

    55:04

    Nathan Labenz: Yeah, I could definitely see myself going a little higher on Anthropic, because I do think model differentiation is likely to be a stronger narrative than commoditization over the next six months. But it's kind of hard to say — a modest guess makes a lot of sense to me.

    55:27

    Prakash: So my thing is, will OpenAI's consumer advertising pick up before the IPO? And if they're showing the same kind of growth on the consumer side that Anthropic's showing on enterprise, their market value will be so much higher, because the consumer business is so much bigger. So — let's see, what does it say?

    55:56

    Nathan Labenz: 69.3.

    55:58

    Prakash: Very close to my 0.667.

    56:03

    Nathan Labenz: Alright, we press on. Will Anthropic's ARR be under $90 billion at the end of 2026?

    56:14

    Prakash: I think that's a yes, and I'd say I'm 80% sure of it.

    56:22

    Nathan Labenz: Okay, I'm definitely lower than that. I think they might make it — it's been a minute since I really tracked this, but I did participate in — I think it was the same folks who do the AI Village —

    56:45

    Prakash: Shoshana Tavkoski?

    56:47

    Nathan Labenz: Yeah, and team. I think they did a prediction survey at the beginning of last year and another at the beginning of this year — one question was total revenue of the frontier AI companies. They don't include Google because it's hard to get the revenue breakout, so it's just Anthropic, OpenAI, and xAI, if I recall. Last year I scored well in part because I went higher than most people on the revenue numbers, and they did come in quite a bit higher than most guesses. This year I estimated end-of-year total revenue for these companies at $1.75 trillion, and I was pretty well on track last I heard. Anthropic's going

    57:32

    to have to be up in that range. So I'll say — only a 30% chance that they'll be under. Actually, that feels like a pretty good over-under, so I'll go a little higher.

    57:53

    Prakash: I'm an Anthropic-revenue-ARR truther — their numbers

    57:58

    Nathan Labenz: Yes, are — okay, so I haven't really dug into that, I've just taken the posted numbers, which I think is probably what this is doing too. But tell me what that controversy is about — they're counting the revenue that goes to Amazon as their own revenue, and then they're

    58:20

    Prakash: Yes, exactly — filing it as cost as well. They count the revenue that Amazon brings in from Claude sold to enterprises, when really they should be counting the net, what they get after Amazon takes its cut. And that's fine — startups can have their own accounting standards — but once they go public, everything gets rewritten to GAAP. So I suspect it's not going to be a $90 billion number at the end of the year.

    58:59

    Nathan Labenz: Okay, that makes it quite complicated. If they flip their accounting and take 30% off the top or whatever their rev share is, that makes a big difference — you're talking real money at this scale. Alright, well — let's see, 15%. Wow, so the market says confidently that they'll be over $90 billion.

    59:30

    Prakash: If they weren't going to IPO, I'd probably be there too — but because I'm also expecting the IPO, I expect they'll have to correct the numbers.

    59:46

    Nathan Labenz: Okay — who will be the world's second trillionaire?

    59:51

    Prakash: That's a good question.

    59:53

    Nathan Labenz: Interesting.

    59:57

    Prakash: Let's see — this is a good question. Alright, Mark Zuckerberg — 40%, I guess. Do they have to sum to 100%?

    1:00:35

    Nathan Labenz: Yeah, sum to 100, I think.

    1:00:37

    Prakash: Sum to 100%.

    1:00:40

    Nathan Labenz: I mean, all these guys have relatively fixed stakes in already-massive companies. Zuckerberg owns maybe 20% of Facebook or something — so at a $2 trillion market cap, that's $400 billion for him. It'd have to go to $5 trillion for him to get to a trillion himself, roughly speaking.

    1:01:06

    Prakash: Mhmm.

    1:01:08

    Nathan Labenz: That doesn't seem crazy-unreasonable, so he might be my highest on this list — I'll say 10% for him, we might adjust these later. Jensen owns what, 2%? 3%? of —

    1:01:27

    Prakash: 3% of NVIDIA? Yeah.

    1:01:29

    Nathan Labenz: So that puts him at $150 billion out of $5 trillion, which means he'd have to go to $30 trillion to get to a trillion for himself — that's seemingly very difficult to imagine, or at least I'd be very surprised if that happened before any of these other things did. So I'll put him at 1%. Bezos, I think, is a similar analysis to Zuckerberg — a $2-to-3-trillion company of which he owns maybe 15%.

    1:02:16

    I have a hard time seeing Amazon be the next one to spike, given they don't have their own AI play nearly as much — he'd have to hit something huge. He's kind of started his own AI company on the side, haven't heard much from that; he's got a space company, haven't heard much from that either. One of those would really have to pop. I think that's possible, but it seems more likely Zuckerberg runs up than that. So I'll say 3%. I think Page and Sergey have

    1:03:01

    like, roughly 10 or 12% of Google.

    1:03:04

    Prakash: Mhmm. Nathan Labenz: Which, at a $4 trillion market cap, gives them $400 billion — meaning they'd have to go to $10 trillion. That doesn't seem that crazy; I'll maybe put them at Zuckerberg's level for now. Altman has really opaque finances — if his fusion company hits, that could be worth a trillion dollars right there. It's unclear how much he owns of OpenAI; there were rumors he'd get 7%, which at a trillion would be $70 billion, meaning he'd have to get OpenAI to $15 trillion to get to a trillion himself at 7% with no further dilution. Again, I think somebody

    1:03:49

    else probably gets there first if that's the story, but who knows — I'll say 5%. Larry Ellison feels like a 1%; I don't think you get there building data centers alone. So Page and Brin are basically the same thing, I don't really know how to distinguish those — so that's 10% to the field immediately. Would I be comfortable putting 60% on the entire rest of the field? No, I think I have to double everything.

    1:04:24

    Prakash: Because if not, you end up with someone random — who's not even in the top hundred right now — getting to a trillion.

    1:04:43

    Nathan Labenz: Yeah, there aren't really that many great candidates. Maybe I'll just even this up a bit more to take the field down. What does that leave me with — 23%? Okay, we're about halfway through this quiz — we might have to pause and come back for the second half after our guest segment today, but let's finish this one. What do you think?

    1:05:18

    Prakash: Okay, Mark Zuckerberg — let's start off with, I think Larry Page is at probably 30%, and Mark Zuckerberg is probably at 30%, and I'd say Jensen is at 20, and I'd say —

    1:05:52

    Nathan Labenz: And is your story there that NVIDIA goes to $30 trillion, or is there some other story where that happens?

    1:05:56

    Prakash: Yes, NVIDIA goes to 30 trillion. And I'd say Jeff Bezos is probably at 15 — that's 30, 60, 70, 80, 95 — Sam Altman is probably at, like, 2, Larry Ellison is at 1, and the field is at 3.

    1:06:18

    Nathan Labenz: Yeah, you know what — I'm going to do boom and boom, and you've got 2 left for the field.

    1:06:26

    Prakash: Here we go, alright, let's —

    1:07:29

    Let's give it a second while we coordinate. Let me actually — great.

    1:08:09

    Nathan Labenz: Not sure what happened — hit stop-sharing on the tab, maybe that caused the hiccup. But in any event, we're 7 out of 14 rounds done. We can pick it up a little later. As of now, according to the market, I have a lead — but another thing we'll have to do is come back and score these retrospectively, because the market, obviously, may not be fully rational either.

    1:08:29

    Prakash: Alright, indeed. And let me introduce our guest for this morning — just give me a second. Yep, alright, and our —

  3. 1:16:03Interview58 min
    Interview: Kyle Rush — The Expertise Exists, It's Just Locked in the Expert's HeadKyle RushHint's co-founder and CTO — Obama 2012, Casper, Maisonette — on building a home-intelligence app with Martha Stewart as a genuinely hands-on co-founder, and on why the hard problem isn't the model. Every fact about a house is a claim with a source and a date, so Hint stores them in a provenance-tracking graph with a taxonomy built around how homeowners think rather than how insurers do; it pulls 1,300 data points at onboarding specifically because the hallucination surface around houses is enormous, then uses AI only to personalize licensed human expertise. He explains the guardrails (the AI knows nothing about where revenue comes from; it asks your age before suggesting a ladder), what happened when a voice agent called a generator technician seventeen times in a row, and why the durable moat is that homeowners don't know that the person who fixes a cracked Bluestone walkway is a hardscaper — you have to know to ask.
    Open segment on YouTube ↗

    There's nothing about the home that she doesn't have a very deep understanding of... it's really like working in the Olympics.

    Every piece of data about the home is a claim, and we want to know where that claim came from, what the date was, and where it was discovered.

    I took our neighbor group chat, exported it all, ran it through an LLM, and had it create a spreadsheet of every contractor mentioned over the years.

    1:10:49What's it like being Martha Stewart's neighbor, and how did the idea of Hint come about?
    Kyle described Martha as caring, deeply knowledgeable about every aspect of the home, and a hands-on cofounder rather than a figurehead — likened working with her to 'the Olympics.'
    1:12:45How is Hint using AI to help homeowners understand things they wouldn't otherwise know, like a circuit breaker box?
    Kyle said Hint is a multimodal AI app built around photos (and eventually voice) that coaches homeowners on maintenance tasks — like inspecting downspouts after winter or knowing which breaker to reset — that they'd otherwise miss or mishandle.
    1:17:31How does Hint handle conflicting data sources, e.g. county records vs. a homeowner's own documents on roof age?
    Kyle said Hint built a home-specific graph database and taxonomy that tracks provenance for every data claim (source, date, discovery), letting it reconcile and update facts over time, illustrated with a pool-pump replacement example.
    1:20:41How does Hint think about liability when its AI-generated advice leads to a problem?
    Kyle said they focus on preventing unsafe recommendations (e.g. not suggesting elderly users climb ladders or touch high-voltage electrical), personalize advice by age and household composition, and back it with a vetted OpenAI contract and business insurance, while AI-specific insurance is still an open industry question.
    1:23:28What kind of AI-specific insurance are vendors proposing, and what terms are they offering?
    Kyle said he isn't closely involved in those conversations and couldn't speak to the details.
    1:24:23How do you think about matching homeowners with service providers and reputation, given search/matching costs have historically been high in that market?
    Kyle pointed to informal neighborhood-shared spreadsheets of vetted service pros as inspiration for a future Hint feature, and said the platform is working to help users understand the tradeoffs between different sourcing channels (TaskRabbit, local pros, provider networks, home warranty referrals).
    Lightly edited · timestamps jump to YouTube
    1:08:30

    Prakash: Our guest for today is Kyle Rush, a technology leader who has spent the last two decades building the engineering engines behind some of the most visible consumer brands and political movements in the country. Kyle previously served as vice president of engineering at Casper, scaling the technology from a Series A mattress startup through their IPO, and was chief technology officer at the children's marketplace Maisonette. Before navigating the hyper-growth of consumer tech, he cut his teeth in the ultimate high-stakes, high-traffic environment: building the fundraising technologies for the Obama and Hillary presidential campaigns, which successfully processed over $1 billion in online contributions. Today,

    1:09:16

    Kyle is the cofounder and chief technology officer of Hint, cofounded alongside home-services veteran Yih-Han Ma and lifestyle icon Martha Stewart. Hint is an AI-powered home intelligence platform launched just this summer with $10 million in seed funding. The app ingests public property records, weather data, soil conditions, and user-uploaded warranties to serve as an automated, proactive owner's manual for your house. Kyle joins us today because he represents a rare bridge in the AI ecosystem: a deeply technical operator running parallel agentic coding experiments to push the boundaries of how software is written, while simultaneously deploying those advanced capabilities into a consumer-facing product designed to solve everyday problems for homeowners. Just give us a second — we are having some problems, just give us one second.

    1:10:03

    Prakash: We are having some problems, just give us one second.

    Nathan Labenz: It keeps flipping back out.

    Prakash: Just give us a second.

    Nathan Labenz: Never a dull moment. What was it — goblins, right? That one model was obsessed with talking about goblins. We might have some goblins in our

    Prakash: Indeed.

    Nathan Labenz: production process here.

    Prakash: Alright, let me let me figure that out. Just give us a moment.

    Nathan Labenz: You want me to leave and rejoin, or you leave and rejoin? I don't know.

    Prakash: Just give me a second. Alright.

    Nathan Labenz: The sound of typing is Prakash getting — is it Fable 5.1 today, or is it

    Prakash: All of the above. Alright, let me refresh quickly. There we go, you're back. Alright, so let's get him back, and we should be

    Nathan Labenz: Too easy.

    Prakash: Too easy. One of the perils of doing a live show is that it is genuinely all live, so these are small mishaps and gremlins. After we sort it out — we run GPT-5.6 Sol Medium on fast mode in the background, so it's ready at any time to step in and help us sort out our issues. Kyle should be in the room, and we should be able to bring him up. As soon as this video is live, we should have him up there. Let's see — never a dull moment.

    Nathan Labenz: I just saw a chat come across, did you see that?

    Prakash: Yep, so we do have chats live today.

    Nathan Labenz: So at the risk of tempting fate any further than we already are — if you're watching and you want to chat in, it's X only right now, not YouTube. But you can suggest a question for Kyle, and it might even get answered — let's see what happens.

    Prakash: There we

    Kyle Rush: go. Hi, Kyle. Hey, how's it going?

    1:10:09

    Prakash: Good morning, sorry about that.

    1:10:11

    Nathan Labenz: Too easy.

    1:10:12

    Prakash: Yeah, sorry for the live interaction, you know, on

    1:10:18

    Nathan Labenz: Kinda like the Anthropic API — we're awesome like 98% of the time, and then about 2% of the time something doesn't quite work the way it's supposed to.

    1:10:27

    Prakash: But

    Nathan Labenz: here we are. It's great to meet you.

    1:10:29

    Kyle Rush: Likewise, and I'm very familiar with the Anthropic API, so I get it.

    1:10:34

    Nathan Labenz: Yeah, well, we're all kind of downstream of these somewhat unreliable APIs these days, and their unreliability is multifaceted to say the least. So Prakash, take it away.

    1:10:49

    Prakash: Kyle, maybe let's start off with what I think everyone's asked you about. I went on LinkedIn and saw Martha Stewart's post — she was talking about how you were her neighbor and you guys were hanging out, chatting. What's it like being Martha Stewart's neighbor, number one? And number two, how did the idea of Hint come about?

    1:11:11

    Kyle Rush: Great question — never saw myself being friends with Martha Stewart. It's great. Martha is an incredible person. I had no idea what to expect when I met her, and even when we started working together. I like to say she's a grandma — she loves our daughter, she's just a very caring, thoughtful person. She's just a really good friend at the end of the day — we talk all the time, call each other, do fun things here and there. So she's amazing. But as a cofounder, oh man — she is in the details. I thought this was going to be maybe a figurehead relationship, and very much not the

    1:11:56

    case. Martha knows all the content she's written, all the work she's done — she is in all those details. She obviously has a team to help her, but there's nothing about the home that she doesn't have a very deep understanding of. Anytime we're working on AI outputs or designing something or talking about features, I'm constantly impressed how much she knows about all the different siding types and the pros and cons, what areas of the country have these types of roofs, these types of gutters are good for this, how do you properly drain gutters and soil — she has a wealth of knowledge there. So it's really like working in the Olympics, is what I like to say. It's a really unique

    1:12:41

    person that I've worked with in the past.

    1:12:45

    Nathan Labenz: So I have a 100-year-old house — I actually need to make good on my promise to the house and have a birthday party for it before the year is out, built in 1926. So I need some copiloting when it comes to keeping this thing together. I downloaded the app, and I'm only my first handful of days in, so I can only say so much. But one thing that jumped out at me right off the bat is what I'm invited to do is take photos — go around and take photos of things. The first photo it asked me to take was of my circuit box down in the basement. So we're obviously always interested in the AI

    1:13:31

    angle on things — tell us how you're using AI, and maybe some of the less obvious ways it helps people make sense of things they might not have even a concept of. Most homeowners probably don't know too much about their circuit breaker box.

    1:13:49

    Kyle Rush: Yeah, definitely. So your story is pretty common — there's a big topic of conversation in the country just on aging homes. I think the median age of a home isn't quite as old as yours, but it's not atypical. A lot of homes are just due for major renovation and major structural updates, and that puts a lot of onus on the homeowner. When I look back on my homeownership journey — I've owned my home for about four years, built in 1969 — it was an amazing time, probably the same for you: you have all these hopes and dreams of the life you're going to build there, you save up a while to make it happen. And then six months in you're like, oh no, this is crazy, this is a lot of work. There's

    1:14:34

    documents, papers everywhere, records — I don't understand insurance, I don't understand what a kilowatt-hour is, am I paying too much for electricity, do I have the right appliances, how do I maintain an oil-fed hydronic baseboard heat system — which I'm from California and living in New York now, I don't know what any of that means. So that's one of the reasons we're building Hint. It's really for anybody — anybody can get use out of it. But let's say you just bought a home and you've been in it for five years — the onus to learn the home and what you have to do is very high because you haven't really owned a home before. So Hint is really a multimodal AI app —

    1:15:19

    you mentioned the photos — we haven't quite done voice just yet, but you can talk to it, and I think photos are very important to helping you maintain the home. I like to think of Hint not necessarily as the most sophisticated AI out there, but it is sophisticated — we have a patent pending on some of our inventions there. But so much of it is coaching the homeowner on what to do. We had this problem of the downspouts that collect rain off the roof — the reason those exist, which a lot of people don't know, is to move the rainwater away from the foundation, because as soon as it's near the foundation that's a chance for the water to leak into the basement and cause serious

    1:16:05

    problems for your home. Hint needs to see what those look like to make a recommendation for you, but most homeowners wouldn't know to even inspect those, say after the winter season when a lot of damage can happen and they can get disconnected or misplaced. So Hint will recommend you take a photo of all the downspouts and make sure they're in good shape. The electrical panels one is also interesting — homeowners, in emergency scenarios, are having to reset the breakers, and the question is which one should I reset. When you're in a panic, you open that box, you've got 40 circuit breakers, you're not really thinking it through — that's where Hint can help a lot. In my area we have power outages a lot because we have a lot of trees.

    1:16:50

    They need a little bit of wind, blows the power down on the lines, and I'm always wondering what's running on our generator and what's not — what can I use, what should I use, what's the protocol. I think you're supposed to reduce power consumption so the generator isn't doing too much work, but these are things normal people wouldn't know, which is totally fine, and that's where Hint can help a lot.

    1:17:14

    Prakash: So let's talk a little bit about one of the problems with dealing with real-world assets, which is often that

    1:17:27

    Nathan Labenz: Get back here.

    1:17:31

    Prakash: Yeah, there we go. So one of the problems dealing with real-world assets is that the reality often differs from the theory. Hint combines public records, local environmental signals, etcetera — what happens when, say, the county property records say the roof is 20 years old, but the homeowner has some random two pages saying, oh, we replaced the roof three years ago? What happens when you have these conflicting sources of data, and how does the system actually handle

    1:18:17

    that kind of conflict?

    1:18:18

    Kyle Rush: Yeah, great question. This is part of the technology we're working on patenting. As with any data in the world, but especially property data, you're right — there is conflicting data. So what we've created is a graph database with a taxonomy specifically for the home. There are a lot of these home taxonomies, but they're designed by governments and insurance companies for how they think about the home, not how the homeowner thinks about it. So we created one that's about how the normal homeowner thinks about their home. One of the key pieces of this is tracking provenance — every piece of data about the home is a claim, and we want to know where

    1:19:04

    that claim came from, what the date was, where it was discovered — it's a key part of piecing together the living history, the story of your home. When you input data into Hint, it can come from multiple places: public data Hint pulls in itself, an inspection report PDF you upload that it reads, an insurance doc we fetch on your behalf from your declaration page, or something you said in the chat. We enter those facts into the graph and they get associated with a number of entities — it's a triple store. That will just be a claim. How it really plays out in most cases is — take my pool, for

    1:19:49

    example: when we moved in, the pool pump was working for about a year, then the motor seized and it had to be replaced. Hint knew about that, and then it knew I got a new pump and the old one was taken out of service — it knows who did the work and when, and it knows the properties of the new pump, the horsepower, the brand, the serial number, all from the invoice I uploaded for the work. Because of the graph structure with provenance, it can track that over time and understand that at one point I had a Hayward one-horsepower pump, but now I have a Hayward 1.5-horsepower pump. That's really key to meeting the homeowner where they are, because usually technical systems

    1:20:34

    get very confused on data like that, and it creates really poor experiences.

    1:20:41

    Prakash: Hint also gives advice — it says you should do this, or you should change this. One of the problems with AIs giving advice is the issue of liability when something goes wrong. How do you guys think about that problem? Is it significant? Is it something you have to insure? How do you think about it?

    1:21:10

    Kyle Rush: Yeah, another good question. There are a few things that are important to think through on safety around the home — the worst thing you'd want is some AI telling you to touch 240-volt electrical haphazardly, or suggesting an 80-year-old person get up on a ladder to inspect the gutters on the second floor. So these are the kinds of things we're really focused on preventing. And by the way, I think the models are generally pretty good at this — we've tested a number of models, we're using all OpenAI models, and everything

    1:21:55

    we've seen, there's nothing super scary in terms of what it's recommending, across about 5,000 users today that we've looked at. But what we try to do is really personalize Hint to you — not just to your home, but to you. We ask your date of birth so we can track your general physical capability over time — if it knows you're over a certain age, it's not going to recommend certain things. Conversely, it wants to know who else is living in the home — if you have a teenager there, that's a different story, and it might recommend the teenager go up and look at the gutters. And then it also just generally knows what

    1:22:40

    your DIY capability and desirability is — some people are 100% DIY, that's how they want to do everything, and others just say I don't have that gift, it doesn't interest me, I need to hire a professional. Hint knows all of those things. And then, like any other business, we have business insurance — we're looking at whether we need special insurance for AI specifically. I think that's an ongoing question in the industry; there haven't been enough court cases for everyone to understand that deeply. But we approach it from that safety level, and we have a lot of the standard things in place — a contract with OpenAI that's been looked at through lawyers,

    1:23:25

    and the business insurance.

    1:23:28

    Prakash: So let me dig into that a little bit — what kind of AI insurance are vendors proposing to you? What kind of terms are they offering — are they offering to insure you for the year, or is it a percentage of revenue? What are you seeing in the market right now?

    1:23:45

    Kyle Rush: Good question — I wish I was involved in those conversations more, given your question. I'm just not super involved in the details of all that, so I wouldn't be able to say.

    1:23:57

    Prakash: Okay, alright.

    1:23:59

    Nathan Labenz: They keep me interested in

    1:24:00

    Kyle Rush: writing the code, you know.

    1:24:03

    Nathan Labenz: Can you tell me more about these strategies for getting teenagers to do their chores in an entirely — this

    1:24:09

    Kyle Rush: sounds

    Nathan Labenz: like a superintelligence achievement.

    1:24:12

    Kyle Rush: Yeah, we're still working on that — can't offer much in the way of convincing the teenagers to do it. But in terms of who's best equipped to do the work, I think Hint would know that.

    1:24:23

    Nathan Labenz: This will be the last thing for humans — yeah, which doesn't say much. How about on a more serious note — the process of matching. I don't know all the details of the business model, but I see there's a mix of the classic 'do you have or need a mortgage,' which has been a tried-and-true path for monetizing apps related to home — you can get people into a new cable, whatever. Then there's also local service providers, which have historically been much harder to monetize generally, but certainly there's a huge market there. I'm really interested in how you think about matching

    1:25:09

    and reputation. One thing I did that I think made my neighbors think I'm weird — which is probably accurate — is I took our neighbor group chat, exported it all, ran it through an LLM, and had it create a spreadsheet of every contractor mentioned over the years and what people had to say about working with them. That creates a pretty valuable asset for my neighbors, at some reputational cost to me interestingly enough, but it's still very partial compared to what an actual platform could do, especially if you know about the home and all the service providers in more

    1:25:54

    depth, what kind of stuff they've done. So the big thesis I have in general that I want your perspective on: anywhere search costs and matching costs have historically been high should be areas where we see dramatically more efficiency and probably a lot of market growth. There's probably a lot of things that would get done by professionals but don't, because I can never find the right person in the first place. Quality should go up, consumer surplus should go up. Tell me your vision for that, and how you think we get there.

    1:26:29

    Kyle Rush: Yeah, another big homeowner topic. Interestingly, my neighborhood — our street — has a Google Sheet we all maintain, a list of all the service professionals, and that was kind of the inspiration for some of our future features. We don't have them today, but there's totally a concept of neighborhood data sharing that I think is super interesting, because a lot of the time when you're finding a referral — insurance or whatever — or just learning information, one thing people love is to know what their neighbors are doing with their homes. Should I be thinking about an expansion? Finishing my basement? Redoing my roof? My home was built by the

    1:27:14

    same builders at the same time as the other homes on the street, so this concept of neighborhood sharing is really interesting. The way we're approaching matching now — if you look at service professionals, we want to help the user understand how to think about everything, because there is a right way to think about it and you have to understand the big picture. You can get a service professional from various avenues — but what's the difference between a TaskRabbit service professional versus one in your local area versus one in a service provider network versus one from a home warranty you might have? They all have their pros and cons, and then who you are

    • The Contractor Call Went Haywire

      0:00 / 0:00
    • Your Home Knows What You Forgot To Ask

      0:00 / 0:00
    • Human Expertise Beats Raw AI

      0:00 / 0:00
    • Martha Is In The Details

      0:00 / 0:00
    • AI Matches Contractors By Context

      0:00 / 0:00
  4. 2:14:17Closing51 min
    Closing: EUV, the Second Half of the Quiz, and a Case for Expanding the TentRounds eight through fourteen cover OpenAI's rumored desk device, whether China gets a functional EUV machine before 2029 (Prakash 80%, Nathan 30%, market 58), a doom market designed so it never pays out, Databricks, a Situational Awareness wind-down, and a Tesla–SpaceX merger — with Nathan reporting from a weekend of Full Self-Driving that "I don't text and drive, I'll text an FSD." Then a fast news round: Lutnick says the administration trusts Anthropic again, Gemini 3.8 Flash goes after real-world benchmarks at 300+ tokens per second, the DOJ sides with OpenAI against the New York Times, and NYSE tells Congress it used Anthropic's Mythos to fix its own vulnerabilities. It ends on David Krueger's criticism of Dean Ball, and Nathan's argument that a post ending in an apology is the moment to expand the tent: the pausers should recognize when they have a new friend.
    Open segment on YouTube ↗

    The show's second half opened with round two of the hosts' 'Guess the Markets' prediction game, picking up at seven of fourteen rounds complete. Nathan recapped the first half: on whether the AI bubble would burst (three of six extreme scenarios required to count), he'd guessed 2% against Prakash's 15% and a market of 10%, noting one of the six criteria — Super Micro's stock being down sharply — had already happened. On an Astra release by month's end, the market came in higher than both hosts, with some rules-lawyering over the name Astra. On US government control of an AI company, Nathan said 50%, Prakash 90%, the market 40%; on whether it would specifically be OpenAI, Nathan said yes while Prakash floated Anthropic instead. On Anthropic surpassing OpenAI in market cap, Nathan guessed 66.7%, Prakash 50%, and the market landed just under 70%; on Anthropic's ARR staying under $90 billion, the market gave only 15%, implying a five-in-six chance revenue tops that mark. On the world's second trillionaire, both liked Zuckerberg and the Google founders, with Prakash favoring Jensen Huang and Nathan favoring Sam Altman, though the market weighted Jensen — and especially Zuckerberg — most heavily.

    The second half kicked off with a Polymarket question on what kind of consumer hardware OpenAI will announce in 2026, with options including a clip-on for clothing, earbuds/headphones, glasses, a necklace, a watch, a ring, or a phone, each resolving independently. Prakash described the rumored device as a small, eye-equipped desk speaker that can look around and react — unlike anything on the list — and both hosts gave the phone option surprisingly high odds (Prakash around 25%) against a market that favored the clip-on and earbuds options and barely weighted the phone. On whether OpenAI would actually launch — not just announce — a consumer hardware product in 2026, both flagged the brutal hardware production cycle and the pending Apple lawsuit as delay risks; Prakash guessed 15%, Nathan 40%, and the market landed at 28.5%.

    On China obtaining a functional EUV lithography machine before 2029, Prakash went to 80%, citing ASML layoffs feeding Chinese hiring at American-style salaries, while Nathan was more skeptical at 30% given the depth of the supply chain (down to a single German company that makes a required lens); the thin market came in at 58%, closer to Prakash. The 'will AI wipe out humanity by 2030' question came with a catch: because it's structured to resolve N/A and refund every trade in 2027, it never actually pays out on the real question, turning it into a pure emotional-expression market. Both hosts explicitly 'metagamed' it — guessing what the market would say rather than their true beliefs — with Prakash guessing 15% and Nathan guessing 30% while calling his own true belief closer to 8%; the actual market number came in lower than either expected.

    On Databricks hitting a $250 billion valuation by the end of 2026 (up from roughly $100 billion a year earlier and now around $190 billion per Nasdaq Private Market pricing), Nathan guessed 35% and Prakash 20%, with the market resolving at 22.5%. A bonus round asked whether the 'Situational Awareness' fund would announce a wind-down by year end; Prakash argued the SEC inquiry is more serious than people realize, pointing to the pattern of comparable hedge-fund blowups ending in prison and the likelihood that not all investors were made whole, and went 35%, talking Nathan up to 25% — but the market put it at just 6%. On the final round, a Tesla-SpaceX merger by end of 2026, Prakash laid out why Elon can't simply decide it himself (both boards and minority shareholders would need to approve) but ultimately expects a merger once Tesla's robotaxi rollout lifts its stock; both hosts landed around 15-20%. Nathan closed the game by 'winning' with 1,369 made-up points, ahead of Prakash and both AI competitors (one of which had research access, one of which didn't and may or may not have cheated) — concluding that internet research meaningfully helped.

    Reflecting on the exercise, Prakash suggested having AI systems track when these markets resolve so the hosts can score themselves in retrospect, and Nathan noted some earlier-round questions may have already resolved or can at least be marked to market. Before handing off for news, Nathan flagged a striking arena-score jump for Fable 5.1 relative to the field — the first time in a while a model has put that much distance between itself and the pack — which Prakash summarized as 'the Pareto frontier moves again.' Prakash then ran through a news roundup: Commerce Secretary Howard Lutnick told Axios the Trump administration now trusts Anthropic after months of friction, calling them 'back on the right side'; Google announced Gemini 3.8 Flash, emphasizing real-world benchmarks (Vals Finance, Harvey's Legal Agent Benchmark) and raw speed — 300-plus tokens per second — at the same price as 3.7; and the DOJ sided with OpenAI against the New York Times, arguing that training an LLM on copyrighted material doesn't violate copyright law. Nathan endorsed the ruling but argued AI companies have 'taken humanity's collective inheritance' with nothing given back, floating conditional legislation as the 'other side of the trade' Congress should consider.

    Prakash noted the New York Stock Exchange has used an Anthropic tool (via 'Project Lastwing') to find and fix system vulnerabilities, then turned to a blowup in the AI-safety world: David Krueger criticized Dean Ball for not candidly disclosing his personal views, tying it to the backlash against Dwarkesh Patel's AI-civilizations post and the reluctance of many in AI policy to voice how radical their real beliefs are. Nathan pushed back hard, calling the criticism unnecessarily harsh — arguing Dean Ball's more conservative public posture is precisely what got him into the Trump administration (and produced the well-received America's AI Action Plan) and that a movement seeking influence should be extending grace to an ally who apologized, not attacking him for playing smart politics.

    Prakash closed with a wider-lens take: he argued doomers are actually updating on real progress (unlike outright 'denialists' like Ed Zitron), that any future different from the status quo will look like a dystopia to someone, and that AI firms are quietly preparing for the equivalent of a new form of governance — citing people like Andrew (Andy) Hall working on frameworks where individual AI agents could vote on people's behalf in a direct democracy, ideas he expects Washington policymakers currently dismiss as crackpot. He described a new 'belief hurdle' being crossed via CivAI, a group that has demonstrated to Congress how open-source Chinese models can be combined with data-broker data to build real-time surveillance dossiers — targeting both gun owners (to alarm Republicans) and abortion providers (to alarm Democrats) — in hopes of finally moving long-stalled data-broker regulation. Nathan agreed the demos are getting genuinely scary and are valuable precisely because they're hands-on and undeniable rather than hypothetical. The hosts signed off taking Thursday off, with the next show on Friday.

    They have taken humanity's collective inheritance and managed to concentrate it down and put it into a product. Great. We never got anything, though, for that.

    The promise with Elon as an investor is that he'll never make you lose money.

    One of the things doomers perhaps don't realize is that our view of the future is necessarily a dystopia because it is something that is different from the current status quo.

    Lightly edited · timestamps jump to YouTube
    2:01:58

    Nathan Labenz: Indeed. How are you doing on time? Do you want to do the second half of our Guess the Markets game? We've got time for it.

    2:02:03

    Prakash: Let's do that. Let's do that.

    2:02:06

    Nathan Labenz: Alright, so as we come back for the second half of Guess the Markets, we're halfway through — seven of fourteen rounds done. Quick recap: will the AI bubble burst? These were pretty extreme scenarios, where three of six specific things had to happen for it to count. I said just 2%, you were at 15, and the market was at 10. Turns out one of the six had already happened — Super Micro being down a lot — but it still feels very hard to imagine two more of those happening in the next four months, though obviously time will tell. Will Astra be released by the end of the month? The market came in higher than both of us — there was a little rules-lawyering on that one around the name Astra. Will the US government take control of any AI company? I said fifty-fifty, you said ninety, the market said forty. Will it specifically be OpenAI? I said yes — if it happens, more likely OpenAI than not. You said probably not, and floated Anthropic instead, which I think there's an interesting story behind that I'd like you to fill out sometime. The market landed somewhere between us on that one. Will Anthropic have a higher market cap than OpenAI? I said 66.7%, you said fifty-fifty, and the market came in just shy of seventy. We were debating the narratives there — best models running away from the pack, versus model commoditization, versus who owns the most compute. You also raised the question of how Anthropic records revenue, and whether it includes the share that goes to partners, which I don't think they're doing right now but might have to in the future. Will Anthropic's ARR be under ninety billion dollars at year end? The market's only at fifteen percent — in other words, there's basically a five-in-six chance revenue comes in over ninety billion, presumably under current accounting. And who will be the world's second trillionaire? We both liked Zuckerberg and the Google founders. You liked Jensen Huang a lot more than I did — I was more willing to tell a story about Sam Altman, given his investments in things like nuclear fusion. But the market gives much higher weight to Jensen, and especially to Zuckerberg, over the rest. So that brings us to the second half. First question, from Polymarket: what kind of consumer hardware will OpenAI announce in 2026? Good question — each row resolves independently, so apparently these don't have to sum to a hundred. For each option, what are the odds it happens?

    2:05:26

    Prakash: It could be more than one, right? They could announce multiple products.

    2:05:30

    Nathan Labenz: Yep — so each one of these is basically an independent question. We're now tasked with seven independent estimates: clip-on for clothing, earbuds and headphones, glasses, necklace, watch, ring, or phone.

    2:05:57

    Prakash: Clip-on device for clothing, I'm at like five percent. Earbuds, headphones slash speakers—

    2:06:13

    Nathan Labenz: Yeah, is speaker in there? Because that—

    2:06:15

    Prakash: No — smart speakers are already kind of known, but they didn't put it out there like that. I guess earbuds, headphones — don't we already know what this is? It's like a little smart speaker with an arm that kind of moves around — that's the eye. It sits on your desk and can react to you, and it can also look around. So they're trying to emphasize that it has an eye, a voice, and it's able to listen to you — they're trying to make a whole new category of device. This is a little tough because it doesn't neatly slot into any of those. Yeah, no, that's not really—

    2:07:08

    Nathan Labenz: I mean, there certainly could be some surprises, but I'm with you that none of these sound like they're on point for what I understand we've heard — and that's also what I've understood from everyone who's seen that line.

    2:07:24

    Prakash: Maybe I'll do all ten percent for clothing, and maybe the phone I'll do like twenty-five percent, because I think there's some chance that if they succeed, they'd try to go ahead with a phone. Phone tech is well known, so it's not like anything new — it would just be an additional device.

    2:07:43

    Nathan Labenz: Alright, so I've got to do the same thing. I mean, they've been inspired by the movie Her, so maybe I'll give them a little higher chance on that — and that's at least consistent with their tabletop thing. The clip-on device for clothing, the pin — remember the pin? Remember that little display where you'd shine it at your palm and read things off your palm?

    2:08:13

    Prakash: Humane. How about—

    2:08:14

    Nathan Labenz: Yeah, I'm not sure how that one got out of the testing stage — although, arguably, it never really did get out of the testing stage.

    2:08:21

    Prakash: I mean, it was really unfortunate for them, because I think they were a little too early — they're very good at product design, but the AI piece is the core piece, and the AI piece hadn't matured yet. Arguably, what Jony Ive later showed OpenAI was probably something similar, but he was two or three years later, and he could sell it. I think Jony Ive was brought in at, what, like $300—

    2:09:02

    Nathan Labenz: Yeah.

    2:09:03

    Prakash: So he's going to end up worth tens of billions of dollars at the IPO, right? So—

    2:09:11

    Nathan Labenz: Alright, well, this one's a little weird — let's just go for it. The market says clip-on device for clothing is most likely, earbuds and headphones next most likely, and then glasses, necklace, watch, ring, and phone — they're really not expecting. So we were both way higher than the market on the phone one; Scout was more with us on that. Not sure there's too much more to say here — let's keep moving. Number nine: will OpenAI launch a consumer hardware product this year? Now, that's another really interesting one with respect to IPO, right? Because, as you mentioned earlier with advertising, that's a revenue driver that could really change how people think about their growth prospects — and an awesome device could do that too. That kind of takes me back a little more toward the phone, though. Phone's a hard thing to compete on, but Apple has definitely left the space open to new competition over the last couple of years. I still use my iPhone, I still like it, but it would be a lot better if it were smarter — and that's an absolutely enormous market, something like $500 billion a year.

    2:10:39

    Prakash: Yeah, but the problem with consumer hardware is you need to have the product ready by like March and announce in September so you can get it out to consumers by Christmas — the cycle is brutal, and you have to coordinate with the guys in Asia. I hear it's a small run of the initial device — probably several hundred thousand units, not reaching a million. I don't think they could support a phone; I think it's easier for them to do a toy device, almost a first version, and not make it too serious.

    2:11:28

    Nathan Labenz: So this last one was 'publicly announced' — this one is 'launched.' Announced doesn't count. If they're going to do a small run, they still have time — they could sell a couple hundred thousand.

    2:11:47

    Prakash: They're also being delayed by the Apple lawsuit — Apple is suing them, trying to delay the release as much as possible. So I don't know, maybe it's a tough question. I do think something will be in your hands in 2027 — I don't know exactly what 'launch' means, but it says 'release,' so I guess that means in hands. I think in hands in 2026 is unlikely, so I'm giving it like fifteen percent maybe.

    2:12:26

    Nathan Labenz: Fifteen — okay, that's definitely lower than me, I'll go forty. You're talking me down a bit; I was initially thinking higher just because we've heard about this for so long that at some point it's got to happen. But the Apple lawsuit is definitely a wild card — I have no idea what to make of that. So let's find out. 28.5 — I think that's almost exactly between us.

    2:12:47

    Prakash: Mhmm.

    2:12:52

    Nathan Labenz: Okay, will China obtain a functional EUV machine before January 1st, 2029?

    2:13:01

    Prakash: Oh, like, yeah, eighty percent I guess.

    2:13:03

    Nathan Labenz: Obtains or develops?

    2:13:12

    Prakash: Yeah, eighty percent. The thing is ASML fired a bunch of people, and the Chinese are very good at hiring — they're willing to pay American-style salaries for a few years to get talent. So yeah, I think they will.

    2:13:35

    Nathan Labenz: Yeah, this is one of the more important questions in the world, I'd say. A lot of American policy over the last couple of years has rested on the assumption that this can't be done — that they're many years away. Never bet against Chinese manufacturing is a pretty good rule to live by in life. A lot of this rests on more than just ASML — there are supply chains, and those suppliers have suppliers, and there's one German company in one town that makes the lens that's needed. Without the lens, you can't do anything, and there are a lot of those little bottlenecks they have to fix. I'm going to work from the assumption that I don't know exactly what 'obtains' means — presumably buying a used one at somebody's garage sale isn't the spirit of the question, so I'm focusing on development. That gives them '27 and '28. I think it's not that likely — I'll say thirty percent that they're able to make this all work by then.

    2:15:08

    Prakash: Eighty.

    2:15:11

    Nathan Labenz: Okay — this is a thin market, so we should have a little caution around the estimate; it may not be as meaningful as some of our others. 58 — again, pretty close to right between us, a little closer to you on that one. There's really multiple years being traded here — the shape of this curve is where you'd have to believe both that this won't happen that fast, and that before it does, we'll have seen some sort of AI takeoff via RSI or what have you. That's the world where you could plausibly play the Machines of Loving Grace strategy — create a decisive strategic advantage and make them an offer they can't refuse. I still think that seems unwise, and I've got other content in the works where I'll talk about that in depth. But this is, at least, consistent with the Dario worldview that they won't be able to make crazy scale of chips — so if we can get Claude to become the country of geniuses in a data center in the next two to three years, we'd have a chance to say how the world looks after that. Crazy. Okay, next: will AI wipe out humanity by 2030?

    2:16:52

    Prakash: The fine print: will continuing progress in AI capabilities result in every human dying by January 1st, 2030? The fine print is the point — play money is worthless if the world ends, so traders have an incentive to bet no regardless of belief. The market actually resolves N/A on January 1st, 2027, and rolls back every trade — it never pays out on the 2030 question. So—

    2:17:20

    Nathan Labenz: If someone — basically it's just a chance to express your emotions.

    2:17:24

    Prakash: Right.

    2:17:25

    Nathan Labenz: With no skin in the game — real money, play money, or otherwise.

    2:17:30

    Prakash: No — does it want to wipe out humanity by 2030? No, it's like a zero percent chance. Oh, well, I have to guess what everyone else is guessing — so alright, I'll guess like fifteen percent.

    2:17:46

    Nathan Labenz: Yeah, well, since we're just demoing here — my guess is, yeah, so now we're metagaming, right? You're thinking not what is my true belief, but what is the market going to say. My true belief is probably lower than your fifteen percent, but not dramatically lower. My guess as to what the market will say, if I'm playing that game, is more like thirty percent — I think we'll see a certain demographic attracted to this question. But in the spirit of playing an earnest strategy, I'll go — my true answer is more like eight percent, and I'm sacrificing points to send the signal that that's my true answer, even though I think the market's going to be higher. Let's see — oh wow, not nearly as high as I would've thought. Interesting — got some non-doomer action in the most doomer of all markets. Three to go: will Databricks' valuation hit $250 billion by 12/31/2026?

    2:18:28

    Prakash: Yeah.

    2:18:29

    Nathan Labenz: They were valued at about $100 billion about a year ago, and then they just raised a big round.

    2:19:21

    Prakash: I think they're at $130 billion now, I think.

    2:19:26

    Nathan Labenz: Yeah, I think they raised like $10 billion. I don't remember the valuation attached to that latest round, but it was a big round. If it was indeed $130 billion, I'm guessing they're not going to double by the end of the year. This is a funny business — in some ways, it was perfectly positioned for AI to come along, just having everybody's data in these giant data warehouses and data lakes.

    2:20:07

    Prakash: It's $190 billion right now — that's the current valuation.

    2:20:15

    Nathan Labenz: Yeah, this would have to be reported — as measured by the price reported by Nasdaq Private Market — so it's a little more liquid than just the next financing announcement. Alright, so from $100 billion to $190 billion in the last eleven months or so — can it go to $250 billion in another four months? I think it's less likely, but I'm talking myself into it not being that unlikely. I'll say thirty-five percent.

    2:20:58

    Prakash: I'd say twenty percent.

    2:21:02

    Nathan Labenz: The answer is 22.5 — you're right on the money there. $190 billion resolved. Interesting. Okay, two more — I guess these are our bonus questions. From Polymarket: will Situational Awareness announce a fund wind-down by December 31st?

    2:21:31

    Prakash: I'd say thirty-five percent. I think the SEC inquiry is more serious than people realize — like I said a couple weeks ago, of all the hedge fund blowups around that size, everyone else ended up in prison, and Situational Awareness blew up but the fund's still around and he's not in prison. I might have been premature saying that. I suspect the risk management was sloppy, and if the risk management was sloppy, the compliance was probably sloppy too. There are a lot of people in the market who don't like him, and this is going to be a very targeted investigation on someone people don't like — even Elon has had so many problems with the SEC. I also suspect not all the investors were made whole, because some of them came in after the Anthropic investment and aren't part of it — and the Anthropic investment is kind of what made the whole fund float. So I don't know, but I suspect it's going to be more difficult than people realize.

    2:23:07

    Nathan Labenz: Alright, you're talking me up a little there — I'll go up to twenty-five. My thought was, well, they didn't blow up so badly that they had to be unwound immediately, and they still had how many billion worth of Anthropic equity sitting around — that's a pretty good cushion. I was initially going to come in lower, but you talked me up to twenty-five. Let's see — six.

    2:23:37

    Prakash: Market doesn't think it's likely. Okay.

    2:23:49

    Nathan Labenz: Interesting. Okay, last one: will a Tesla-SpaceX merger be officially announced by 12/31/2026?

    2:24:02

    Prakash: Good question — fifteen percent?

    2:24:08

    Nathan Labenz: I don't really know how to think about this one. What do you think are the relevant considerations? It seems like Elon gets to do whatever he wants in this domain, so—

    2:24:21

    Prakash: He doesn't get to decide this one — he has to basically take himself out of the picture. He's not going to be able to vote his shares, so this is about whether the minorities—

    2:24:33

    Nathan Labenz: And why wouldn't he be able to?

    2:24:35

    Prakash: Because he's an interested party — he's CEO of both, so he'd have to take himself out of the running to vote his shares. The boards of both entities would have to recommend the merger, and the shareholders would get a vote. Especially for Tesla, he's not controlling — but even for SpaceX, as the controlling shareholder, he's got so much in there he wouldn't be the one voting; it'd be the minorities. So in order to get it done, it has to be a good deal for both sides. I think Tesla right now isn't doing so well, but I've started seeing robotaxis on the highway in Silicon Valley. I think what happens is the robotaxi rollout and full self-driving finally get fully out — he's got five thousand cars on the road by the end of the year, more than Waymo — Tesla shares get managed upwards, and then he announces the merger. If he waits until next year, Tesla's stock price may be too high, and Starship is still going to take a little time before the revenue kicks in. So Tesla is the thing maturing right now — he's got to get the price above the previous high, by at least ten percent, and it's down now, so he's got to get it twenty percent above the previous high and then take it up from there. So yeah, it's possible.

    2:26:35

    Nathan Labenz: From a pure fundamentals, business-operations basis — I took another little weekend trip in a full self-driving Tesla this weekend. And I hear the first Model Y L, which is the one I'm actually interested in potentially buying, is now in a Michigan Tesla dealer — the station wagon, as you'd call it — so at some point, go check it out and see if I'm actually comfortable in the back seat. The FSD is truly incredible — at this point it really just feels like you have a superhuman driver doing the driving for you. I took several people on their first FSD ride and they were all totally blown away — the improvement was notable even from April or May, the last time I did it. Parking in particular was already very good then, but it's gotten better taste — it gets lost in parking lots less, and picks better spots. It doesn't yet have the Grok integration where you can just say 'park in that space right there,' but they said they're working on that. The driving itself is unbelievable, and it gives you significantly more leash than it used to when it comes to how much attention you have to pay. A couple years ago you had to keep your hand on the wheel with a little torque all the time or it would bother you; four months ago, you didn't have to do anything but if you took your eyes off the road it would bother you pretty quickly. Now I feel like I was taking my eyes off the road at some length and it wasn't giving me a hard time — I got beeped at a couple times, but as much stuff as I was doing on my phone in the driver's seat, that just goes to show how much trust I have in it. I'm normally a very attentive driver — I don't text and drive, but I'll text and FSD. So I guess the question is: how closely can these companies already work together? I can imagine a story where Grok's already integrated, they've got plans to make chips, plans to do all these things, and it's all already operationally integrated via partnership — you don't necessarily need to commingle the cap tables for the work to be done in an integrated way with tight feedback loops across the organizations. If that's the case, maybe there's not much pressure to merge. But if there's some reason that isn't the case, I think you'd run through walls to knock down any silo walls between different parts of the AI operation — and I just don't know how closely they can work together today.

    2:29:54

    Prakash: I think the issue is Elon gets sued a lot — every single thing gets sued and reported on. Like, Tesla paid for the first data center and then xAI used it, because he didn't want to use xAI's cash flow, and they got sued for that — a shareholder lawsuit. The deal in the US, as Matt Levine says, is that everything is a crime against the shareholders, and the shareholder-lawsuit lobby continuously sues companies. The more of these intercompany links he has, the more he gets sued on it — and every lawsuit, they either do discovery or he pays ten to twenty million. If they do discovery, they dig up emails, and if the emails show they underpriced something or knew the market price and underpriced anyway — and he has no idea how his fifty thousand employees underneath are making deals or writing emails — it's just cleaner if the companies do a lot of business together in a more Asian-tycoon style, where he's got all his pockets helping each other: a personal pocket, personal borrowings, and he moves money from one company to another, like SpaceX to xAI. And the promise with Elon as an investor is that he'll never make you lose money — that's the promise he's made to the market, and he's kept it up consistently, through SolarCity being bought, through Twitter being bought by xAI, xAI being bought by SpaceX, pumping up the valuation each time, and then publicly listing so everyone makes money. So he has the confidence of the market because of that. I think in the end he will merge — I think that's the cleanest thing for him.

    2:32:24

    Nathan Labenz: So I'm saying twenty. After all that, I feel like maybe I want to go higher, but I already put it in there.

    2:32:30

    Prakash: But he's also announced this year, which is tough, so let's — yeah.

    2:32:32

    Nathan Labenz: Not a lot of time remaining. Alright, let's see — 18.5, while we bracket it again. Alright, well, of course we can only truly score these with the passage of time, but at least for who can be the most conventional, I come out on top with 1,369 made-up points. You come in second, and we both come in ahead of the AIs — including Scout, which had research, and Q, which didn't. So Q may or may not have cheated; if it did, it probably could've scored higher than that, so I'm guessing it didn't — though, as we know, we should be verifying with detailed transcript analysis before we get too confident about what the AIs did or didn't put into any given solution. But if we take it at face value, at least we can say the internet research helped. Any other reflections on our experience today?

    2:33:34

    Prakash: Well, perhaps we can set up the AIs to iterate through these — we've done three of these so far — and remind us when the markets close and resolve, so we can look at how we performed in retrospect. That might be interesting.

    2:33:54

    Nathan Labenz: Yeah — I wonder if there are a couple that have already resolved from our earlier ones. When you do these out to like 2030, it leaves us a long time to wait, obviously, but—

    2:34:05

    Prakash: And we can probably dig up the—

    2:34:08

    Nathan Labenz: We can also mark them to market regardless, right, even if they haven't resolved, if they've moved. That'll be interesting enough.

    2:34:15

    Prakash: Indeed. Let me close with a quick iteration of what people are talking about right now.

    2:34:27

    Nathan Labenz: Let me show you one other thing first, and then I'll give you the last word. This just popped up on my feed — I have some questions about how accurate all these arena scores are, but that's a big bump to see relative to the field for Fable 5.1. I think we can say for sure that's meaningful, even if exactly how meaningful is something each person will have to feel out for themselves over time. But as I was scrolling by, it definitely jumped out — it's been a minute since we've seen a model put that much distance between itself and the rest.

    2:35:13

    Prakash: The Pareto frontier moves again.

    2:35:17

    Nathan Labenz: Alright, what have you got?

    2:35:20

    Prakash: Very quickly: the beef is over. The Anthropic-versus-Trump-administration beef is over — Anthropic are back on the right side. Commerce Secretary Howard Lutnick says the Trump administration now trusts Anthropic, following months of clashes with the AI company. Asked if he trusts Anthropic CEO Dario Amodei, Lutnick told Axios' Mike Allen: 'We trust Anthropic. They've done what we asked. They are back on the right side. So the answer is yes.' Lutnick said that in an interview on Tuesday. So — backing off, notably Lutnick, who's Commerce, not Hegseth, who's Defense.

    2:36:04

    Nathan Labenz: Yeah, or Michael, what's-his-name, who I'm guessing isn't giving it up so easily — but time will tell.

    2:36:12

    Prakash: Very quickly, Google announces Gemini 3.8 Flash. It's a jump, and they're targeting real-world use cases — Vals Finance, a financial-analyst benchmark, and Harvey's Legal Agent Benchmark. So it's more focused on real-world activities. From a capabilities perspective, Google has clearly fallen behind, some say, but Google is now pressing their unique hardware advantage instead — 300-plus tokens per second at sticker pricing. So they're offering high tokens-per-second with moderate intelligence improvements, and I suspect that may be quite attractive to a lot of people — speed may actually matter more than intelligence for a lot of tasks.

    2:37:06

    Nathan Labenz: What's the price on this one? I think they'd dropped the price as of the last one too, right?

    2:37:13

    Prakash: Same price as 3.7, and it's available on the Gemini API. Same price as 3.7.

    2:37:30

    Nathan Labenz: Very quickly — comes in a lot cheaper than a lot of other things, for sure: 75 cents per million input tokens for now, though this is introductory pricing.

    2:37:41

    Prakash: MTS Live says the US government has sided with OpenAI against the New York Times. The DOJ says training an LLM on copyrighted works does not violate copyright law, and that treating it as infringement would hurt US science, prosperity, and national security. So let's see how far this goes.

    2:38:04

    Nathan Labenz: Yeah, my guess is that's going all the way to the Supreme Court, and I like that ruling — it would be a real shame to let a bunch of legacy IP owners stop the train of progress. I think this is a good position for the US government to take. I do think it'd be nice to get something back for it, though — I really do think there's another side to this trade. They've taken humanity's collective inheritance, managed to concentrate it down, and put it into a product. Great — we never got anything, though, for that. I think they should be allowed to do it, but there should be another side to that trade: you're allowed to do it, but what's the 'but'? That's an interesting question for the government to tackle next — if it's Congress, and you have the ability to write laws, you could imagine saying, you can do this if you satisfy these other conditions, and if you don't, you can't. There are a lot of things you might want to make AI companies do, obviously, but this would be one trade I think is really worth thinking hard about.

    2:39:30

    Prakash: Indeed, next one: the New York Stock Exchange used an Anthropic tool through Project Lastwing to find and fix vulnerabilities in its systems, NYSE president Lynn Martin told Congress — so this is more support for Anthropic on cyber. And finally, one reflection: I think Dean Ball is really taking it on the nose today from various people in AI safety. This is David Krueger saying Dean didn't consistently, candidly reveal his personal opinions on these issues. He states — but it is true that I, and candidly I think many of my colleagues in the profession of AI policy, largely failed to talk about this issue with the level of seriousness and urgency it required — and this is about identifying humanity and not being taken in by AI agents. What struck me was Dwarkesh Patel going ahead with his AI-civilizations blog post, and how many people responded with 'these people are crazy.' I think what the rest of the world fails to realize is a lot of people in SF share those views, and a lot of them are hesitant to discuss them in public because they sound crazy — it's what Jensen Huang calls sci-fi. I think that's one of the problems in communicating: Dean and others have to have clarity and be able to work with policymakers, yet these beliefs are so radical that it's hard for them to interface — so they end up interfacing on a normal basis, but they hold all these beliefs they think may be true in the long run, even though they have a lower probability and aren't yet evident. It's a tough one.

    2:42:06

    Nathan Labenz: Yeah, I think this is unnecessarily harsh, to be honest. I know David a little bit, not well, but I've met him a few times, and I respect the impulse — he's got a bunch of pause/stop/rewind emojis in his header, so he's clearly playing a very transparent, hold-nothing-back strategy. I'm not sure that's the right reaction if you're trying to win at politics, though. What this shows me is — and I'm choosing my words carefully myself here, because I don't want to make enemies of either of these people — what I don't like about this post is that he ends with an apology; Dean ends with an apology at the bottom. In general, if somebody is showing enough reflection to get to the point of apologizing, that's a good moment to extend some grace and try to make common cause. If you're David Krueger and you want to pause, stop, or rewind, I'd think this would be a moment to try to make common cause, to expand the tent, to adopt a bit more of the strategy that's clearly worked for Dean. He went from a think-tank guy focused on state and local policy three years ago, to starting a blog through — I think — quite inspired writing, kind of a Hamilton-style story of writing his way to the top, gaining influence in a16z circles as a voice they thought was compelling on SB 1047 way back when, and getting the Trump administration job. He does not get that job if he's seen as a crazy doomer — I think that's safe to say. The America's AI Action Plan, which at this point is so many news cycles ago who knows how much influence it still has, but when it came out, it was one of the only documents ever to come out of the Trump administration that was well received across the spectrum — even Zvi Mowshowitz had nice things to say about it. You don't get that document out of the Trump administration if he's not in that role, and he's not in that role if he doesn't play a somewhat conservative public-communication strategy. And he probably doesn't get a job at OpenAI either — although at this point, who knows, OpenAI might be open to anything. I think a portfolio approach is usually what I'd say to people when they bicker with each other over tactics they're using to achieve similar ends. Zooming out: you guys both seem to have at least somewhat of a healthy fear of super-powerful intelligence at this point — that's enough common ground to build on. Let's not, as the old Monty Python line goes, bicker and argue about who killed who; this is supposed to be a happy occasion. A little more forward-looking view would be really good. Dean doesn't get to this position if he's not somewhat conservative, but he's also shown himself to be AGI-pilled enough to get a job at OpenAI — so I think he's played a pretty savvy strategy. I would not call this a shameful lack of integrity, and I think the pausers have got to recognize when they have a new friend — that's kind of my take on this.

    2:46:12

    Prakash: So I think, number one, on the doomer side — I think they're actually correcting a lot of their projections on what might happen, because they're also seeing progress. There's another group, the denialist group, kind of Ed Zitron-driven, and that's a completely different group, because that group doesn't see progress at all. The doomers actually do see the progress. I think one of the things doomers perhaps don't realize is that our view of the future is necessarily a dystopia, because it's something different from the current status quo — that's one thing on the doomer side. And on the policymaker side — the AI firms, I have to be honest — they're preparing for the equivalent of a revolution: a new type of government. People like Andrew Hall are inside the firms trying to define what a new form of politics would look like when you can assign decision-making power to agents and have them vote on your behalf, maybe directly, in a direct democracy — which would look very different from representative democracy, where you elect representatives who filter information. You could potentially have just your own AI as your representative, and have direct democracy. There's a lot of possible future changes here that are kind of revolutionary and potentially huge in how governance is done, and I think some of those ideas would, at this point, be seen as crackpot ideas by real policymakers in Washington — 'these people are crazy, academics, you don't know what you're talking about, we're never going to allow this to happen.' That's interesting to me, because on one hand people like Andy Hall are trying to think through what these things might be, and on the other hand, people like David Krueger think the future Andy Hall is trying to create is already a dystopia — already a disempowerment, in the sense that human representative leaders might be disempowered, and AIs empowered on behalf of individual humans instead. So all these guys are actually seeing maybe the same future, but with a different lens on how attractive it is — they have different opinions on what's desirable, even though they might see the same thing potentially happening. It's also not clear when policymakers will start to cross these belief hurdles. I think one belief hurdle was Mitos — that got crossed; that was the first time they started to see safety issues. I've heard another belief hurdle has been crossed recently: there's an organization called CivAI that's been in Congress recently, and they've used models like Kimi or GLM, plugged in data brokers, and let those models extract — say, show me who this person is, Christian in Minnesota, doing this and this, here's their daily activity, and so on — all just extracted from existing data brokers and joined together. This is precisely what Dario was talking about earlier in the cycle, about the kind of surveillance that could be done. The thing CivAI did that was particularly good is they attacked the Republicans by showing how a gun-owner targeting system would work, and attacked the Democrats with an abortion-provider targeting system, and provided these dossiers to both sides — so both sides started thinking, oh my god, what is going on. What CivAI is trying to do is promote regulations on data brokers, which people have been asking about for maybe fifteen, twenty years. I think we're finally starting to see movement, because what CivAI is saying is: look, the models exist — in fact, we have to use open-source Chinese models because our own GPT and Claude won't allow us to do this, but we're using these open-source Chinese models and just plugging them in. CivAI anonymizes the dossiers, and shows the actual product where you can type in someone's name and extract in real time — though they don't let you take the data out. They're showing this to people at the Capitol, and I think that might actually get us some movement.

    2:51:33

    Nathan Labenz: Scary demos — yeah, they're getting really scary now. That's been a theory of change in the AI safety world for a long time: can we show what's possible and wake people up through hands-on experience? At this point, I'd say it's safe to say some of these demos are legitimately getting scary. I've had that experience firsthand — when you sit down and all of a sudden the model is suggesting radical action to you, that was my first moment of thinking, oh my god, this is going to be a really crazy future. I think it's good that people get these experiences — when it's a real scary demo, that's the thing that really gets people: I just gave you this name, you did this, there's no magic trick here, you actually have a working technology that can do these things. That experience moves the needle for at least some people, and I do think it's good that people are out there doing those scary-demo briefings.

    2:52:49

    Prakash: Indeed. And on that note, Nathan, we're taking a break tomorrow, and we'll be back on Friday.

    2:52:56

    Nathan Labenz: Alrighty, sounds good. See you Friday, Prakash.

    2:52:58

    Prakash: Cheers.

    Nathan Labenz: Bye for now.

Loop transformers, opaque serial depth, and what OpenAI won't commit to

Prakash opened on The Information's report of a leak from OpenAI: Astra uses something called a loop transformer, and the implication that chain-of-thought monitoring would be harder against it caused an immediate ruckus. Nathan's first move was to lower the baseline. Even with full access to a chain of thought, he said, what you see is a model thrashing — considering many options, with cheating very often among them on a hard problem, and metagaming about what the grader seems to want essentially ubiquitous — and at the moment of action it is still not clear, even to someone like Bronson Shane who has read millions of tokens of these traces with human eyes, why the model does what it does. He heard echoes of the same from Ryan and Ajeya in the OpenFace investigation. His takeaway from Recursive, the weekend event on recursive self-improvement, was that the frontier plan is "chain of thought monitoring all the way down," with Jeffrey Irving's more generous framing being that it is really scalable oversight, of which monitoring is a large part. Either way, the technique only works if the trace is both readable and faithful.

The technical thread ran through Meta's Coconut paper as the first credible published version of the idea: instead of decoding a token at the end of a forward pass, feed the last internal state back in as an embedding, so the model reasons over a blob of unresolved possibility rather than a single collapsed choice. Nathan's example was the graph-traversal task Coconut tested, where a model can evaluate multiple paths at once in latent space instead of narrating each one linearly — the same answers in fewer forward passes, at the price of not knowing what it was thinking along the way. He pointed to Rohin Shah and the Google team's work bounding what they call opaque serial depth: how many computational steps an architecture can take before it must externalize something a human can read. A vanilla transformer is favorable on that metric; recurrent and loop-style structures can push it arbitrarily high. Prakash brought his own Fable 5.1 artifact on loop parameters, walking through the difference between the outer loop of chain of thought — where a monitor sees "3x equals 21, so x equals 7" — and an inner recurrent loop where the monitor sees only "the answer is 7." He also read a defense circulating online that a loop transformer is not Coconut-style latent reasoning because loops emit nothing at all; Nathan called that a fine implementation point, arguing the real change is that the embedding slot, previously restricted to the hundred thousand or so one-hot token vectors in the vocabulary, will now accept any vector — and that it works with vanishingly little additional training, which is exactly why it will work much better with training.

The politics of it bothered Nathan more than the mechanism. This was, in his phrase, classic OpenAI: make a great deal of noise about not going down a path, get reported as going down it a little, then say the reporting is unfair and that it is only slightly more opaque serial depth than normal — a claim from OpenAI's head of research that he tied directly back to Monday's argument about the inadequacy of the OpenFace investigation and to a company that currently has neither the community's trust nor its competitors'. His proposal was for labs to start sharing negative research agendas — not what they are doing, which they will never disclose, but what they commit not to do — with a matchable cap on opaque serial depth per token as the concrete example. Prakash then turned to release logistics: people probing the OpenAI responses API have found a gpt-6-astra slug returning a 404 rather than a nonexistent-model error, the same signature as models known to exist, and a Thursday launch is the expectation. On capability he relayed reporting that Astra scored 100% on exploit gym, forcing OpenAI to build an internal extension using bugs never previously found, on which it solved roughly 40% and turned up two additional zero-days it wasn't asked for — "that's what we call extra credit," said Nathan. Prakash's read on whether the pause was real: the models were ready months ago, a roughly thirty-day voluntary White House clearance process now exists, cyber and bio models propagate first to organizations that sign up rather than to the public, and that process took from the Mythos preview in February to September to settle. He expects the next revenue push to follow immediately — Greg Brockman was booked for a cybersecurity discussion with enterprise leaders the following afternoon — while noting Malte Ubel, formerly CTO of Vercel, had pointed out to them that security is already a permanent tax on every IT business, and these models are far cheaper than hiring a firm to do penetration testing.

Guess the Markets: bubbles, trillionaires, and the first seven rounds

Nathan framed the segment with a scouting report on Fable 5.1 — better benchmark scores, more token efficiency, and above all a much cheaper price on cache hits, which he read as a possible echo of recursive self-improvement given OpenAI's statement that some of its own price reductions came from optimizations 5.6 Sol made to its stack. He also noted it takes another bite out of the gap Prakash has tracked between Claude and Anthropic's internal researchers, from around 185 down into the low 160s. Fable 5.1 also built the quiz itself. Two AI contestants played alongside the hosts — Scout, which had web research, and Cue, which did not — with Nathan flagging up front that he had not monitored their chains of thought or tool calls, so cheating remained on the table if they ran away with it.

On Polymarket's "will the AI bubble burst by December 31, 2026" — three of six triggers inside a ninety-day window, including NVIDIA down 50% from its all-time high, a semiconductor index down 40%, an OpenAI or Anthropic bankruptcy, OpenAI being acquired, H100 rentals at a dollar an hour for five straight days, or a major supplier down 50% — Nathan said 2% and thought even that was high. Prakash said 15%, reasoning from the Iran war, Scott Bessent's trouble with bond yields, and an administration managing markets only until midterms it expects to lose. The market printed just under 10, and Nathan noted one trigger had already fired on Super Micro. On Astra shipping publicly by September 30, the disagreement was about naming rather than capability: Nathan put roughly 80% on the next big model arriving this month, marked himself down to 65 because the internal model already has several names in circulation, then went back up to 80 after reading the fine print allowing a renamed product to count if confirmed to be the same model. Prakash said 80. The market said 92 — and, revealingly, Scout said 5% and Cue said 18%, which Nathan read as the research-equipped model being credulous about OpenAI's training-pause claims.

The government questions split the hosts hardest. On whether the US government takes operational control of any AI company or project before 2030 — conservatorship or receivership, not an equity stake or a golden share — Prakash went to 90% while conceding the rules lawyer in him thinks it would be hard to prove, and Nathan called it a live possibility but took a coin flip at 50; the market said 40. Narrowed to OpenAI specifically, Nathan came down to 25 on the theory that Sam Altman might want the admiral-of-AI role, while Prakash went to 10 and named Anthropic as his candidate instead. On Anthropic finishing its first trading day above OpenAI's market cap, Nathan took 0.667 on the view that model differentiation beats commoditization as the narrative over the next six months — and that in a commoditization world OpenAI's compute position wins regardless — while Prakash split it 50/50, sketching an October Anthropic IPO targeting two trillion against an OpenAI listing in January or February. The market landed at 69.3. The revenue question drew Prakash's sharpest argument: he is an Anthropic ARR truther, on the grounds that the company books the gross revenue Amazon brings in from Claude sold to enterprises rather than the net after Amazon's share, and that an IPO forces a GAAP restatement — so he put 80% on ARR coming in under $90 billion where the market said 15%. They closed the half on the world's second trillionaire, both leaning to Mark Zuckerberg and the Google founders while working the arithmetic out loud (Zuckerberg needs Meta around five trillion; Jensen Huang, at roughly 3% of NVIDIA, needs thirty), with Prakash weighting Jensen at 20 and Nathan at 1. Seven of fourteen rounds done, a dropped connection and a stopped tab-share along the way, and a promise to come back after the guest.

Kyle Rush: the expertise exists, it's just locked in the expert's head

Kyle Rush is co-founder and CTO of Hint, an AI home-intelligence app launched this summer with $10 million in seed funding and co-founded alongside home-services veteran Yih-Han Ma and Martha Stewart. Before that he was VP of Engineering at Casper through its IPO, CTO at the children's marketplace Maisonette, and built fundraising technology for the Obama and Hillary campaigns that processed over a billion dollars in online contributions. Asked what Martha Stewart is actually like as a co-founder, he said he had expected a figurehead relationship and got the opposite — she is in the details, with deep working knowledge of siding types, regional roof and gutter conventions, drainage and soil, which he described as "working in the Olympics." Nathan, a few days into using the app on a house built in 1926, noted that the first thing Hint asked him to do was photograph his basement circuit-breaker box. Rush's framing for that: most homeowners don't know why downspouts exist (to move rainwater away from the foundation before it gets into the basement), wouldn't think to inspect them after a winter that can knock them loose, and in a power outage are standing in front of forty breakers with no idea which one to reset.

The technical core of the segment was provenance. Hint has built a graph database with a taxonomy designed around how a homeowner thinks about a house rather than how governments and insurers do, and — the part they're patenting — it treats every piece of data as a claim with a source, a date and a place of discovery. Claims arrive from public records Hint pulls itself, uploaded inspection PDFs, an insurance declaration page fetched on the user's behalf, or something the user simply said in chat, and land in a triple store where they can conflict and be superseded. His own example was a pool pump that failed after a year: Hint knows the old Hayward one-horsepower unit was taken out of service, knows the 1.5-horsepower replacement's brand, serial number and horsepower, and knows who did the work and when — all from an uploaded invoice. On safety he said the guardrail is personalizing to the person and not just the house: Hint asks for date of birth so it won't send an eighty-year-old up a ladder, asks who else lives in the home so it can suggest the teenager instead, and asks how much DIY appetite you actually have. Across roughly five thousand users to date and a stack he said is all OpenAI models, he reported seeing nothing scary in the recommendations, backed by an OpenAI contract vetted by lawyers and ordinary business insurance; whether AI-specific insurance is needed he called an open industry question he isn't personally in the room for. On prompt injection he hasn't seen an attempt yet, credits strict data modeling and a deliberately small blast radius — the model can read about your home but can't book or buy anything, and any write to your home data gets confirmed with you first.

The matching discussion is where his thesis showed. Nathan described exporting his neighborhood group chat and running it through an LLM to build a spreadsheet of every contractor ever mentioned — valuable to his neighbors, at some reputational cost to himself — and asked whether high search-and-match costs are about to collapse. Rush's street maintains a Google Sheet of service pros, which he cited as inspiration for a future neighborhood-data-sharing feature; today Hint hits partner Thumbtack's API agentically inside chat. His own case: the Bluestone walkways on his property crack every winter, and he only learned from Hint that the person who fixes that is a hardscaper, not a landscaper — so he wouldn't have known what to search. Hint reads the reviews, checks for anything alarming, filters for Bluestone experience, sorts by proximity and by job count and rating, and deliberately returns three options rather than one, because on the home you should feel three people out and some of it is a gut call. Voice agents calling on your behalf, they tried; it went badly. The agent called a generator technician seventeen times in a row until he jumped off a job site convinced it was a life-or-death emergency, then asked for a model number it didn't need. He expects the endpoint is agent-to-agent — "they exchange neuralese we can't read and then eventually make a deal," Nathan offered — and Rush agreed. Cue, brought in by Nathan, asked three questions: how recommendations stay neutral given affiliate revenue, how Martha's structured knowledge blends with live model output, and what pilot data showed on proactive versus reactive savings. Rush's answers to the first two are the product: the AI knows nothing about where Hint's revenue comes from and has no instruction to sell anything; and Hint pulls 1,300 data points on your home at onboarding precisely because the hallucination surface around houses is enormous, then uses AI only to personalize licensed human expertise — "the expertise to maintain your home exists, it's just been locked in the expert's head." Asked what protects that from Claude and ChatGPT, he pointed to an MCP and an iMessage interface as things he'll ship rather than defend against, and to jurisdictional data as the real gap: he lives in the hamlet of Katona but in the town of Lewisboro while Martha is in Bedford, and every model he asks — including the Claude he codes Hint with — puts him in the wrong town and gets his taxes and regulations wrong. Then a 3D sun-path model of his own house explained wood rot on a deck that gets no sun and takes sprinkler spray, which is the real point: "you have to know how to ask and you have to know to ask," and homeowners only acquire that vocabulary after twenty years. In the debrief Nathan was less persuaded by hallucination as a durable moat and more persuaded by the 3D visualization, and cited his father's rule that you should expect to spend 2% of a home's value on maintenance every year. Prakash read the business as attaching to a high-value asset the way Airbnb and Uber did, and wondered aloud whether the enduring category is simply storing people's context in organized ways — which sent Nathan into why he insists on owning his own, and into the burner-phone photo migration out of China that Claude pulled off through Google Takeout with location metadata and live photos intact.

Closing: EUV, the second half of the quiz, and a case for expanding the tent

The second half of "Guess the Markets" ran the harder questions. On what consumer hardware OpenAI announces in 2026 — seven independent rows covering a clip-on, earbuds, glasses, a necklace, a watch, a ring and a phone — both hosts went low across the board and higher than the market on a phone, because Prakash's understanding of the actual device is a small desk speaker with a moving arm, an eye and a voice that doesn't slot into any of the listed categories. On whether a consumer hardware product actually launches this year, Prakash took 15% on hardware-cycle logistics (a small initial run of several hundred thousand units, coordination with Asian manufacturers, and an Apple lawsuit aimed at delay) against Nathan's 40; the market said 28.5. On whether China obtains a functional EUV machine before 2029, Prakash said 80% — ASML has laid people off, and Chinese firms hire aggressively at American-style salaries — while Nathan took 30, focused on development rather than acquisition and on the long tail of supply-chain bottlenecks, down to the single German company whose lens nobody can substitute. The market said 58. Nathan called it one of the more important questions in the world, because a great deal of American policy has rested on the assumption that it can't be done, and noted it is the crux of the machines-of-loving-grace strategy of building a decisive strategic advantage and making an offer that can't be refused — a strategy he still thinks unwise, with more of his argument on it in the works.

The doom question was a lesson in market design: "will AI wipe out humanity by 2030" resolves N/A on January 1, 2027 and rolls back every trade, so it never pays out and traders have no incentive to bet their beliefs. Prakash metagamed it to 15%; Nathan said his true belief was lower, guessed the market would print around 30 because of who the question attracts, then entered 8% and said explicitly that he was sacrificing points to send an honest signal. The market came in lower than either expected. Databricks reaching $250 billion by year end drew 35 from Nathan and 20 from Prakash against a market of 22.5, off a current valuation Prakash put at $190 billion. A Situational Awareness fund wind-down announced by December 31 drew Prakash's most detailed case — an SEC inquiry he thinks is more serious than people realize, an assumption that sloppy risk management implies sloppy compliance, a suspicion that later investors weren't made whole, and the Anthropic stake that floated the whole fund — which talked Nathan up from lower to 25 against a market of 6. And on a Tesla–SpaceX merger being announced this year, Prakash argued Elon Musk can't vote his own shares as an interested party, so minority holders on both sides have to approve, which means he needs Tesla's price above its previous high first — a robotaxi ramp he says he is now watching roll out on Silicon Valley highways. Nathan came at it from the operating side after a weekend in a Full Self-Driving Tesla he described as a superhuman driver with noticeably better parking taste than in the spring and a much longer attention leash: "I don't text and drive. I'll text an FSD." Final scores on a game they can only truly settle with time: Nathan first with 1,369 made-up points, Prakash second, and both hosts ahead of the AIs — Scout, which had research, ahead of Cue, which didn't.

The news round was quick. Nathan pulled up an arena chart showing Fable 5.1 opening unusual distance on the field — "the Pareto frontier moves again," said Prakash. Commerce Secretary Howard Lutnick told Axios' Mike Allen the administration now trusts Anthropic, that the company has done what was asked and is "back on the right side"; Prakash noted pointedly that this is Commerce and not Defense. Google announced Gemini 3.8 Flash, aimed at real-world benchmarks like Vals Finance and Harvey's legal agent benchmark, at over 300 tokens per second and the same price as 3.7, with Nathan pricing the introductory rate at 75 cents per million input tokens; Prakash's read is that speed may matter more than intelligence for a lot of tasks. The DOJ sided with OpenAI against the New York Times, arguing that training on copyrighted works is not infringement and that treating it as such would harm US science and national security. Nathan expects that to reach the Supreme Court and likes the ruling — he doesn't want legacy IP owners stopping the train — but wants the other side of the trade written down: they have taken humanity's collective inheritance and concentrated it into a product, and the public has gotten nothing back for it. And NYSE president Lynn Martin told Congress the exchange used Anthropic's Mythos, through something called Project Lastwing, to find and fix vulnerabilities in its own systems.

The close was the day's real argument. Prakash surfaced David Krueger criticizing Dean Ball for not candidly stating his personal views on AI risk — Ball's own post concedes that he, and candidly many of his colleagues in AI policy, largely failed to talk about the issue with the seriousness it required. Prakash's diagnosis was structural: a great many people in San Francisco share these views and stay quiet because saying them out loud reads as crazy, which is exactly what makes the policy interface so hard. Nathan called the criticism unnecessarily harsh, and made a strategy argument instead. Ball went from a state-and-local think tank job three years ago to writing his way into influence, to a Trump administration role he does not get if he is seen as a crazy doomer, to the America's AI Action Plan — one of the few documents out of that administration received well across the spectrum, with even Zvi Mowshowitz saying nice things — and then to OpenAI. That trajectory required a conservative public communication strategy, and the post in question ends with an apology, which Nathan read as the moment to extend grace and make common cause rather than press the attack: "the pausers gotta recognize when they have a new friend." Prakash extended it into a broader observation that the doomers are updating on visible progress in a way the Ed Zitron-style denialists are not, and that people like Andy Hall inside the labs are seriously designing agent-mediated direct democracy — the same future Krueger sees as disempowerment, viewed through a different lens. His last item was a theory-of-change data point: an organization he named as Civ AI has been demonstrating in Congress with open Chinese models — Kimi or GLM, he said, because ChatGPT and Claude refuse — plugged into commercial data brokers to build live dossiers on named individuals, targeted at Republicans through gun-owner surveillance and at Democrats through abortion-provider surveillance, in service of data-broker regulation people have sought for two decades. Nathan's response was that scary demos have long been a theory of change in AI safety and that they are finally getting genuinely scary — the kind where you type in a name and there is no magic trick, just a working technology — and that it is good people are out there giving those briefings. No show Thursday; back Friday.