Thomas Sohmers and swyx on agent-driven engineering, memory-efficient AI hardware, and why accountability matters when software generation gets cheaper.
EPISODE 2026-10-06
AI Coding Agents, Chip Design, and the Value of Human Judgment
Thomas Sohmers and swyx discuss AI coding agents, chip-design testing, memory-efficient AI hardware, hiring, verification, and human judgment.
The rundown
- --:--OpeningOpening
- --:--InterviewThomas Sohmers
Thomas SohmersAI chip design with autonomous agents: why Positron spent more on tokens than payroll, and why it kept buying the best models.
Thomas Sohmers, co-founder of Positron AI, which ships inference servers to Oracle, explains how token spending overtook salaries as his team pushed new models through parallel engineering experiments. He says agents can now learn Cadence's verification tools from documentation and run testing loops themselves, with humans still reviewing failure reports. Spending fell after optimization and a cheaper strong model became available, but his purchasing rule remains capability first—not a tenfold token discount for a weaker engineering model.
How agents built their own testing harnesses for Cadence Palladium—and where they still miss Positron's design ingenuity.
Why more concurrent agents increase memory demand, and how Asimov's planned commodity-memory architecture aims to keep models on fewer devices.
Why looped transformers can reduce memory capacity requirements without reducing computation or data movement.Watch
Clips
AI token spending overtook payroll at Positron
0:00 / 0:00AI chip memory: Positron's eight-GPU capacity plan
0:00 / 0:00AI model looping: Why less memory isn't less compute
0:00 / 0:00AI chip verification: Agents learn Cadence's toolchain
0:00 / 0:00TSMC capacity: Why Positron avoids the newest nodes
0:00 / 0:00
- --:--Interviewswyx
Shawn Wang (swyx)AI coding agents are changing hiring: what makes an employee valuable when generating software gets cheaper?
Shawn “swyx” Wang runs AI Engineer and uses coding agents to build and operate its conference software. He says two or three employees are under performance review for delivering model output without adding useful insight, while his SaaS-replacement bounty exposed how cheap generation can create expensive verification work. His argument lands on accountability: workers must contribute judgment and testing, and team leads must deliver useful results at an acceptable cost.
Why a report of 14.7% growth failed his test for expertise: the employee couldn’t explain what caused it.
How two coding agents created competing message-loading paths and a race condition—and why module-level oversight matters.
How a $10,000 bounty to replace a $40,000 SaaS subscription made workflow testing the bottleneck.
Why skeptical event staff became willing to switch when they could request software changes and get them in one to two hours.Watch
Clips
AI slop at work: Why swyx put staff on performance review
0:00 / 0:00AI agents face CPU shortages: Why pause-and-resume matters
0:00 / 0:00AI coding agents vs SaaS: Changes in hours, not quarters
0:00 / 0:00AI coding demand: Spreadsheets become custom software
0:00 / 0:00AI coding agents caused a race condition in his app
0:00 / 0:00
- --:--ClosingClosing
In this episode
Thomas Sohmers, co-founder of Positron AI, describes how coding agents changed chip-design testing, drove token spending above payroll, and shaped the company’s plans for commodity memory and looped transformers. Shawn “swyx” Wang of AI Engineer discusses hiring and performance in the age of Claude: employees still need to bring expertise, judgment, testing, and accountability to AI-generated work.
- AI chip design with autonomous agents: why Positron spent more on tokens than payroll, and why it kept buying the best models. Thomas Sohmers, co-founder of Positron AI, which ships inference servers to Oracle, explains how token spending overtook salaries as his team pushed new models through parallel engineering experiments. He says agents can now learn Cadence's verification tools from documentation and run testing loops themselves, with humans still reviewing failure reports. Spending fell after optimization and a cheaper strong model became available, but his purchasing rule remains capability first—not a tenfold token discount for a weaker engineering model. How agents built their own testing harnesses for Cadence Palladium—and where they still miss Positron's design ingenuity. Why more concurrent agents increase memory demand, and how Asimov's planned commodity-memory architecture aims to keep models on fewer devices. Why looped transformers can reduce memory capacity requirements without reducing computation or data movement.
- AI coding agents are changing hiring: what makes an employee valuable when generating software gets cheaper? Shawn “swyx” Wang runs AI Engineer and uses coding agents to build and operate its conference software. He says two or three employees are under performance review for delivering model output without adding useful insight, while his SaaS-replacement bounty exposed how cheap generation can create expensive verification work. His argument lands on accountability: workers must contribute judgment and testing, and team leads must deliver useful results at an acceptable cost. Why a report of 14.7% growth failed his test for expertise: the employee couldn’t explain what caused it. How two coding agents created competing message-loading paths and a race condition—and why module-level oversight matters. How a $10,000 bounty to replace a $40,000 SaaS subscription made workflow testing the bottleneck. Why skeptical event staff became willing to switch when they could request software changes and get them in one to two hours.