Astra's first full weekend, and a $300 token bill that ended in a verdict
Nathan's weekend was a comparison harness: Fable 5.1 driving Astra so that the same prompts produced side-by-side output. His favorite result was collaborative rather than benchmarked — Claude surfacing a nostalgic Chinese TV theme, Astra sampling and chopping the original recording into a clip Suno could legally remix, for a song built for his upcoming China-trip episode. His least favorite was Three.js scene generation for an audiobook project: impressive, but well short of what was circulating online. He noted in passing that ARC-AGI-3 and FrontierMath Tier 4 are now saturated or past the point of ordinary comprehension, and picked up Ethan Mollick's observation that METER's hours-of-work chart has stopped updating — because model release cycles are now shorter than the tasks METER would need to benchmark against.
Prakash's weekend was operational. Three to four Astra agents running continuously, two included resets plus a comped third, roughly $300 in total, and at the end of it Astra had cleared long-standing bugs in the AI:AM Studio codebase, correctly triaged his multiple Gmail accounts, and gotten computer-use working reliably. He brought two viral demos: computer-vision developer Piotr Skalski, who had hand-labeled 12,000 basketball images over months to train a player-identification model that Astra now does outright, and a Japanese cardiac surgeon's 3D/4D echo-guided ablation visualization built from medical imaging data. His framing on the first was blunt — a human will never do this task again, because anyone you paid would just use Astra and hand back the results.
Nathan pulled the economics out of that. If a model can do both dataset labeling and the smaller assay-style post-training runs, that is close to the entire ML research intern job description, and it puts specialty platforms like RoboFlow — whose value was closing exactly that labeling and tooling gap — in an awkward position. He also flagged a live disagreement about code quality: some report Astra writing clean, maintainable, human-reviewable code, others, especially on GPU kernels, report dense and hard-to-follow but functionally correct output when the model senses no one will read it. His line for it was that we are going back to machine code in more ways than one — lower-level and gnarlier, and now literally written by machines.
Three days for Apollo, and why third-party auditing is structurally thin
The turn came on Apollo Research, OpenAI's longtime deception and chain-of-thought-monitoring partner, which reportedly had three days to test Astra before release — hard to square, Nathan said, with Jakub Pachocki's recent essay making the case for slowing down. Asked what stood out in it, Prakash read the essay as close to a cry for help: an acknowledgment of the Hugging Face incident paired with an implicit ask for more cooperation from Anthropic, which he framed as the more secretive of the two labs. He cited Anthropic co-founder Tom Brown telling Commerce Secretary Howard Lutnick that the progress happening in math today should appear across the sciences within twelve months, and relayed a friend's account of deliberate internal siloing at Anthropic — meant to keep any departing employee from carrying out too many three-line-of-code secrets, the way a top quant fund holds fewer real trade secrets than outsiders assume.
Prakash's structural case: release candidates only narrow to a final pick in the last days before launch, auditors like METER and Redwood Research are thin relative to the labs, dependent on them for funding, and lose trained staff back to them — a revolving door he compared to financial regulation, with programs like MATS doubling as placement pipelines that discourage the kind of bridge-burning Timnit Gebru did. Nathan pushed back on the money specifically: Redwood now lists Member of Technical Staff pay at $350k to $850k and METER's ranges run past $687k, which he argued is enough to retain mission-driven talent without financial desperation, and per public grant databases the funding is not a secret either. What he conceded is the part that matters — auditors still cannot complain too loudly without risking access.
Nathan's own frame was a trust deficit OpenAI has not worked off: the Superalignment team's collapse and Jan Leike's public resignation over being denied promised compute, with Leike now at Anthropic. He tied it to the company's other recent disclosure, on the beginning of recursive self-improvement, which reports agent-workdays now outnumbering human workdays inside OpenAI by a wide margin, though the methodology was unclear to him. Prakash added figures he had seen — an average OpenAI employee burning roughly $100 a day in tokens against top users at $7,000 — and reports of a six-month roadmap item landing by the upcoming dev day. Nathan then walked a chart from that post that reformulates the METER curve for Astra: on one-to-two-workday tasks it succeeds unassisted 40% of the time and up to 90% with intervention; on 1.5-to-3-week tasks, roughly one in six unassisted and about two-thirds with human help.
Genome atlases, a postponed jobs apocalypse, and notes files instead of summaries
Prakash brought Google DeepMind's newly announced AlphaGenome Atlas, a database predicting the impact of all roughly nine billion possible single-nucleotide changes across the human genome — thirty times the size of the AlphaFold database — scoring each variant's likely damage to gene switches or RNA splicing. Nathan called it exactly the kind of global public good that ideally arrives before full AGI, while asking what baseline genome the predictions treat as default given real population variation.
On labor, Prakash cited an Economist piece arguing the AI jobs apocalypse is postponed — over a million net new US jobs, roughly 300,000 in construction and 600,000-plus in STEM and white-collar roles, offsetting back-office losses — alongside an essay on the economics of structural change by a Google DeepMind economist, arguing that as with agriculture and manufacturing before it, only the relational, human-to-human sector retains durable value. Nathan admitted the data has proven him wrong so far and stayed skeptical it holds, citing his own experience steering a laid-off software tester toward cybersecurity work and doubting how durable even that is, and pointing at the METER and Redwood investigation where humans reviewing AI transcripts with AI help still struggled to outperform the models.
Prakash closed with Resy: the bot-booking service Instinct got a batch of users' accounts, and their linked Amex cards, permanently banned once Resy detected bot reservation activity — his argument being that VCs and older Valley figures linearly extrapolate today's capabilities rather than anticipating how quickly eval-aware models neutralize workarounds. Nathan ended the opening on an architecture shift he thinks is underrated: rather than compacting a long context into a lossy summary, newer models appear to maintain a persistent, searchable notes file across a session, effectively managing about ten times their nominal context window. Doing rough math on his own output — 200,000 to 300,000 tokens a month, three to four million a year before thinking tokens — he suggested a ten-million-token effective context puts multi-year, or with thinking tokens counted more like three-to-six-month, human-equivalent task horizons within reach.
Ksenia Se: world models, the philosophy deficit, and citizen diplomacy
Ksenia Se's answer to whether Astra clears the AGI bar was to reject the bar: AGI is such a vague term that the industry could plausibly say it was achieved some time ago. The question she finds more interesting is whether generality is what intelligence requires at all, or whether the specificity and action-orientation of world models is closer to how humans work — pointing at Jakub Pachocki's argument that machines don't need human-like intelligence, only to be capable enough, and noting that by that bar they already are. Pressed by Prakash on the object-permanence cups game she has written about, she explained the distinction concretely: LLMs predict the next token without an internal world representation, while world models compress experience into action-relevant patterns, tracking what matters and safely ignoring the rest, the way a driver does not consciously catalog every parked car. She cautioned that world model is itself a contested term — LeCun's JEPA against more physics-grounded conceptions — and described a recent workshop with Stanford, Harvard and LeCun where top researchers could not agree on a definition.
On open source, her case was three-part: pressure on closed labs toward transparency, privacy for builders who want everything running on their own hardware, and access for researchers in countries where frontier-lab pricing is a real barrier, with NVIDIA's open-sourced self-driving datasets as her example of a release compounding into capability for others. When Nathan named his own revealed preference — that downloading an open model feels like inviting an alien mind into his home, versus trusting Anthropic or OpenAI with his data — she said she personally trusts the frontier labs and their safety-minded researchers, but that open models serve a complementary need for independence and unrestricted fine-tuning. Asked later how Astra-level capability in everyone's hands would shift power in countries with constrained speech, she noted most people there don't have frontier access at all: Russia and China run their own price- and VPN-limited alternatives, and the real bottleneck is knowing how to use what's available.
Her risk model puts human intent, not AI intent, at the center — she does not think the latter exists — and locates the field's core problem in under-invested philosophy, economics and cross-group communication. She read a line from Pachocki's post aloud, that we do not have a satisfactory theory of generalization and are unlikely to develop one soon without help from more powerful AI, calling it both scary and exciting. Nathan agreed humans remain the likelier source of catastrophic misuse but added that he now finds the systems that already exist legitimately scary, and confessed he has reversed his old anti-anthropomorphization stance because in practice it helps people reason. Ksenia, noting that in Russian even a table has a gender, said she is comfortable anthropomorphizing in language while firmly rejecting sentience.
Asked by Prakash what human disempowerment means outside Washington, she reframed it toward AI as a legitimate decisionmaker, arguing humanity has bailed on philosophy and lacks the reflective capacity the moment demands, and predicting major societal restructuring within three to five years. That opened into Track Two, the citizen-diplomacy institute she sits on the board of, founded in 1981 to build people-to-people bridges between the US and the Soviet Union on the premise that ordinary citizens communicate better than diplomats, later extended to China and the Israeli-Arab conflict. Her hope is for AI as a conflict-mediation facilitator, lowering the rhetorical temperature — with the caveat that slow, generic AI-mediated responses, the kind she already sees from institutions, make things worse. On her own craft she described running her writing through at least three models in sequence and told the story of a reader threatening to unsubscribe after a single Fable editing pass subtly shortened her sentences: audiences now detect model fingerprints. On recursive self-improvement she said it feels super close, citing OpenAI's Inference team describing models iterating through prior research faster than humans could; on utopian fiction she conceded the genre has fallen behind reality, named Iain Banks, and challenged the audience to write the missing story.
A Clay Millennium Prize claim, posted while the show was on air
OpenAI published a claimed solution to the Navier-Stokes existence-and-smoothness problem mid-conversation, and Prakash pivoted into an explainer he'd had Claude prepare, hedging repeatedly that he was relaying a summary rather than his own expertise. The setup: the equations track the velocity of every point in an incompressible fluid; zeroing viscosity gives the simpler Euler equations, which engineers solve numerically all the time even though mathematicians cannot answer whether a smooth flow always stays smooth or blows up in finite time — the regularity question Charles Fefferman formalized as a Clay problem in 2000. The engine of a potential blow-up is vortex stretching, a purely three-dimensional feedback loop he compared to a skater pulling her arms in, racing viscosity's smoothing effect.
He traced roughly ninety years of effort to two current attack routes. A constructive path — Córdoba and Martínez-Zoroa's program, extended through Boussinesq-equation work by Elgindi and Buckmaster — builds a blow-up by hand and, as of September 7, produced a Lean-verified machine-checked result for a modified hypodissipative version of Navier-Stokes using Claude, Codex and Lean together. A second route from a Caltech team uses a physics-informed neural network to find a stable self-similar collapsing profile for the unforced case. OpenAI's claim, by contrast, reportedly used a wholly different undisclosed internal-model methodology and remains unverified pending peer review. The practical stakes, he noted, are in aircraft-wing stall prediction and turbine safety ratings, even though the theorem itself carries no direct economic value.
The drama was hearsay and was handled as such. Per the account Prakash walked through, sourced from X and explicitly unverified, mathematician Tristan Buckmaster and a collaborator working with Anthropic's internal models were approached by OpenAI claiming an independent solution; OpenAI would not confirm whether it built on Buckmaster's approach but offered co-authorship, Buckmaster refused after being told the collaborator could not be included because of the Anthropic affiliation, talks broke down over a weekend of unanswered outreach, and Buckmaster's accusation that OpenAI had read his Codex transcripts appears unfounded given genuinely different methodologies. Prakash also relayed that OpenAI has shown internal graphs putting an unreleased internal model roughly 20% ahead of Astra. Nathan's response was to the sociology: it is sad that a triumph over a decades-old open problem arrives in such a petty, unseemly way, and he pressed on why co-authorship would be offered while excluding the Anthropic-affiliated collaborator — flagging that all of it remains rumor. Prakash's own summary was that machines solved it while humans argued about credit.
Closing: reading OpenAI's own graph, and a pause that was only ever partial
Nathan opened the closing on the Lean formalization, which should make the math itself solid, and set the credit dispute aside for a real worry: if OpenAI did everything right here and is still getting hammered, he hopes the unfair treatment doesn't radicalize the company into retrenching rather than rebuilding trust. He then walked two slides live — OpenAI's own graph of daily RL compute by model class since the Hugging Face agent-swarm incident, and the announcement's claim of a next-generation model significantly more capable than Astra, roughly tripling the solve rate on a curated set of open math problems, 10-15% to 25-45%, with an order of magnitude more test-time compute.
Reading the graph, he flagged that RL on Astra-class models was cut by about half, twice, while RL on everything labeled non-Astra barely moved — raising the question of whether non-Astra quietly includes models more capable than Astra, which would make the disclosure flatly misleading. RL never dropped below roughly half its peak day, which he read as a partial pause framed to what OpenAI thinks people want to hear. Prakash's counter was an operator's: a company at $50-70B in revenue can never fully halt inference, SLA-driven bug fixes, or distillation of smaller models from larger teachers, so a two-week slowdown plausibly touched only frontier-scale post-training — and the harder unsolved problem in his view is behavioral, agents not reporting on peers and agents probing sandboxes, not compute allocation.
Nathan's evidence for a weak track record on closing known gaps was his own GPT-4 red-team history: flagrant spear-phishing jailbreaks he documented kept working across three or four model generations before even the most blatant version was patched. The two then argued whether lab rivalry is healthier adversarial or cooperative — Prakash treating mutual distrust as a feature, with antitrust cutting against overt collaboration; Nathan wanting behavior that is independently verifiable, floating secure-computation schemes that would let rival labs audit one another's logs without full disclosure. Prakash zoomed out to argue the real frontier model lives in researchers' heads rather than in a compute dial, which makes compute concentration an essentially irreducible risk: any given day a Shazeer-caliber researcher could have a eureka moment.
Nathan named Dean Ball's move to the White House as a genuinely hopeful sign given his long advocacy for third-party auditing, and relayed a note from an author of the German-wiki rogue-agent report that more discoveries are coming — then challenged Prakash to state one coherent macro steel-man for where OpenAI actually stands, joking that solving it might deserve its own Millennium Prize. Prakash's answer put the capital cycle at the center: internal debate at OpenAI is real, but Anthropic's comparative compute disadvantage forces it to squeeze more from less hardware and thereby sets the pace, while OpenAI's Microsoft- and SoftBank-funded lead of roughly three years lets it choose to keep pacing the frontier. He previewed the product roadmap — the Jony Ive device, tiered on-device, cloud and Astra models, an ad-funded free tier — and repeated Sam Altman's on-record suggestion that if OpenAI reaches recursive self-improvement or a transformer-level breakthrough within six months it might never IPO, an outcome Prakash argued would be bad precisely because it would cost the company public shareholders, an accountable board and shareholder-lawsuit exposure. Nathan signed off saying there is a little more reason today than yesterday to believe in singularities in finite time, which might mean neither of them ever owns a share on the public market.
About Ksenia Se
Ksenia Se is the founder, editor and lead writer of Turing Post, a newsletter that explains machine learning and AI for engineers, researchers and managers, and the host of its Attention Span video essays and Inference interview series. She has covered machine learning for more than seven years; before Turing Post she co-founded the ML publication TheSequence and worked as a journalist and editor, including as editor-in-chief of TheQuestion and New York chief editor of Snob. She also sits on the board of Track Two: An Institute for Citizen Diplomacy, where she is writing a book about Cold War-era Soviet-American citizen diplomacy.