Two Moonshots episodes this week argue that the interesting part of the frontier is no longer “did it pass the test.” GPT-6 Astra, discussed as OpenAI’s first full pre-train since GPT-4o, saturates ARC-AGI-3 and FrontierMath Tier 4 and was built to drive a computer, not just write text. Anthropic’s generally available Fable 5.1 is the model the panel would actually hand a stranger, and it spent 13 million lines of formal code walking Fermat’s Last Theorem. A Palo Alto lab, meanwhile, claims an AI designed a chip in two weeks. What a curious person can now see is a split field: closed models that one-shot hard exams, open conversation about whether those exams still mean anything, and the first serious hint that the next scaling law is depth and silicon, not just more tokens.
GPT-6 Astra saturates the hard benches — and was built to use a computer
Source: Moonshots EP #286 — GPT-6 Astra Saturates ARC-AGI-3, Tesla's $30K Cybercab Floods Austin, Anthropic Proves Fermat's Last Theorem (recorded 4 Sep 2026, published 5 Sep 2026), with Emad Mostaque.
Peter reads OpenAI’s own release on air: Astra is “state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.” The numbers the panel sits with are 99.9% on ARC-AGI-3, 100% on ExploitBench, and 98% on FrontierMath Tier 4. Hallucinations are described as falling from 92% to 51% while accuracy rose. Training is framed as a ~$1 billion pre-train on about 100,000 next-generation chips — the first full pre-train since GPT-4o — with computer-use assistance designed in from the start, so the model can read a screen, click, and loop without a human glue layer. Alex Wissner-Gross’s resolution of the confusing leaderboard is the useful one: Astra is third on the broad Artificial Analysis index, behind Anthropic’s Fable 5.1 and Meta’s Muse Spark, because it was optimized for intelligence per output token, not for winning every suite. The inner story, if the hosts have it right, is recurrence via looped transformers — weight-tying and depth instead of just width — which Alex calls the start of a scaling law the field has not run before. Rollout is not a public chatbot drop. The panel says trusted-access partners first, with OpenAI itself rating the system a top-tier cyber risk and talking about a kill switch the hosts treat as more marketing than physics.
Why it matters: For a decade the public test of “how smart is this?” was a text box. Astra is the first OpenAI model the Moonshots table treats as a native operator of software: it is supposed to sit on the desktop, see the screen, and finish the job with fewer tokens. That is a different object than a chatbot that saturates a quiz. It also means the next argument will be about control and proliferation, not about whether the model can pass ARC.
Horizon: NOW for partner-access computer-use; NEXT for a general public that can actually point it at their own machine.
Evidence grade: shipped (partner access) / demo
Read or watch: FULL for the Astra and benchmark stretch (from ~07:11); SKIM the rest if you only want the numbers.
Caveat: Saturation on named benches is not the same as a reliable coworker. Astra is not #1 on the broad economic-task index, the kill switch is not a law of nature, and the depth-scaling claim is the panel’s reading of OpenAI’s architecture, not an independent paper.
Fable 5.1 is the model you can actually use — and it formalized Fermat
Source: Moonshots EP #286 — GPT-6 Astra Saturates ARC-AGI-3, Tesla's $30K Cybercab Floods Austin, Anthropic Proves Fermat's Last Theorem (5 Sep 2026).
Alex’s ranking, stated without hedging: Fable 5.1 is “broadly the strongest generally available model that we have today.” It leads the Artificial Analysis Intelligence Index, scores 60.9% on Humanity’s Last Exam without tools and 65% with tools, and doubles to 52.6% on Terminal Bench Science — the suite that asks a model to drive scientific computing on its own. Cache reads are 75% cheaper than Fable 5, which the table treats as the real product change: cheap enough long context that a whole business, codebase, or research program can sit in memory. The headline that is not a benchmark is the formalization: Anthropic walked Fermat’s Last Theorem into 13 million lines of machine-checkable code and proved about 29,000 theorems on the way. Peter’s line on the pod is that math is being “incinerated”; Alex expects Clay Millennium-class problems to start falling in months, not decades. A separate, more restricted line — Mythos 5.1 — is reserved on the show for cybersecurity and life-science work with harder guardrails. Context is the industry-standard million tokens, stretched in practice by agents passing messages rather than by a single infinite window.
Why it matters: Most people will not get Astra this week. They can get Fable 5.1. That is the distinction that actually changes what a non-specialist can try: a generally available model that is now the panel’s default, plus a public demonstration that formal mathematics — long treated as the last human redoubt — is a software problem with a mounting line count.
Horizon: NOW
Evidence grade: shipped (Fable 5.1) / demo (Fermat formalization)
Read or watch: FULL for the math and “which model should I use” stretch (~33:44 and the Fable ranking); SKIM if you only need the HLE number.
Caveat: A formalization of a theorem that humans already proved is not a new theorem. It is a new kind of artifact — machine-checked, enormous, and a preview of what happens when the same stack is pointed at problems that are still open. Benchmarks here are also decaying in half-life; today’s 65% will not be a landmark by winter.
An AI designed a chip in two weeks and beat Jetson on the projection sheet
Source: Moonshots EP #285 — Humanity's First Star Probe, Architect Labs Beats NVIDIA 3.4x, Musk Wants Satellites to Cool Earth (recorded 1 Sep 2026, published 2 Sep 2026), with Philip Johnston and Matt Pines.
Architect Labs, a Palo Alto startup, is introduced on the show as having produced Redwood: a frontier accelerator designed end-to-end by an AI system from a high-level spec written by two humans. The claimed loop covers the performance model, RTL, verification, firmware, drivers, and custom kernels in under two weeks, with “zero bugs on first silicon” in the hosts’ telling. Live inference is on FPGA hardware, running models in the Llama and Qwen families. Projected onto an 8 nm class comparable to NVIDIA’s Jetson Orin Nano, the design is described as 3.4× performance per watt. The panel’s leap — and they flag it as a leap — is recursive self-improvement at the chip layer: the model that runs on the chip proposing the next chip. NVIDIA is not treated as finished; the hosts note the company already trains internal design models on Verilog. What changed this week is not that NVIDIA vanished. What changed is that a two-person spec plus an AI stack is now a credible story for a working accelerator prototype.
Why it matters: Model releases every few days do not matter if the hardware cycle stays measured in years. A two-week design loop, even on an FPGA with projected rather than taped-out silicon, is the first concrete picture the Moonshots table has offered of software and chips closing on each other. If that loop holds, the constraint on the next year is energy and foundry time, not the availability of human RTL teams.
Horizon: NEXT for real silicon; SPECULATIVE for a million-x stack collapse.
Evidence grade: demo
Read or watch: FULL for the chip segment (~01:03:28); SKIP the geoengineering and Mars-ship stretches if you are here for AI.
Caveat: Redwood has not been manufactured as an ASIC. The 3.4× figure is an FPGA-plus-process projection against Jetson, not a shipped part you can buy. “Zero bugs on first silicon” is the company’s and the hosts’ language, not a foundry tape-out report.
World models leave the slide deck: Atlas treats 3D as a native modality
Source: Moonshots EP #286 — GPT-6 Astra Saturates ARC-AGI-3… (5 Sep 2026).
Later in the same episode the table turns to World Labs’ Atlas, described as a multimodal autoregressive diffusion transformer that generates image and video with pixel-perfect camera control and reconstructs a scene in 3D from a single photo. Gaussian splats are treated as a first-class training modality, not a rendering afterthought. The use the hosts care about is not a prettier demo reel. It is training robots inside a high-fidelity world, building digital twins of places and organizations, and giving VFX and simulation a model that already understands camera motion. Emad’s metaphor is the holodeck; the constraint he names is compute — memory prices and scarce GPUs delaying the thing you can actually run at home. Astra is separately framed on the same show as a candidate desktop: a model that can stand up software, not just talk about it.
Why it matters: Language models reason about the world in words. A world model that can be queried like a place — move the camera, reconstruct the room, run the physics — is how robots, games, and planning tools stop being separate industries. You cannot buy a holodeck this week. You can watch the modality shift from pixel patches to 3D primitives, and that is new.
Horizon: NEXT
Evidence grade: demo
Read or watch: SKIM unless you care about robotics or simulation (~35:08).
Caveat: Atlas is compute-hungry and not a consumer product. Reconstruction from one photo is a capability claim on the pod, not a guarantee that your living room will come back metrically perfect.
Altman puts AGI on a four-month internal clock; buyers start paying for outcomes
Source: Moonshots EP #285 — Humanity's First Star Probe, Architect Labs Beats NVIDIA 3.4x… (2 Sep 2026).
Sam Altman is quoted as expecting an internal system he is willing to call AGI by the end of 2026 — four months from the taping. Mark Chen is cited at roughly 80% on OpenAI’s internal AGI benchmarks. Altman’s color, as read on the show: this is the first model that “invents new things in a way that matters.” The same episode notes a quieter commercial shift. Salesforce’s Agentforce is already priced on customer revenue generated rather than tokens burned; OpenAI, the hosts say, is letting some large customers pay when the job completes. That is not a new model. It is a new contract: CPM to CPC to CPA, now applied to intelligence.
Why it matters: Timeline talk from lab CEOs is always partly theater. Pair it with outcome pricing and the story gets more concrete. If the buyer only pays when the task is done, the model is being treated as labor, not as a meter. Whether that labor is “AGI” by December is a definition fight. That enterprises are willing to write the contract that way is not.
Horizon: NEXT for outcome-priced work; SPECULATIVE for Altman’s year-end AGI label.
Evidence grade: rumor (AGI clock) / shipped (outcome pricing, as described on the show)
Read or watch: SKIM (~46:18 and ~52:32).
Caveat: “AGI” here is OpenAI’s internal yardstick, not a public eval you can rerun. Previous Altman clocks have moved. Outcome pricing only works if both sides can agree what “done” means — easy for a closed ticket, hard for a research program.
What to take from it
The week does not give you one winner. It gives you a division of labor. Astra is the closed, computer-native system that lights up the hardest named exams and is being withheld from a full public drop because the same skills that pass ExploitBench are dual-use. Fable 5.1 is the thing a non-specialist can open this week, and the Fermat formalization is the cleanest picture yet of scientific work turning into verifiable code at industrial scale. Architect Labs is not a replacement for NVIDIA; it is a demonstration that the design cycle for accelerators can be collapsed by the same class of model that runs on those accelerators. Atlas is the reminder that the next interface may not be a chat box at all. Hold two thoughts at once. Benchmarks are dying of success — ARC-AGI-3 at 99.9% is a trophy and a warning that the trophy case is too small. And the binding constraints named on both episodes are not “smarter weights.” They are energy, memory, foundry time, and whether a government tries to criminalize the next jump. The attentive read is not that superintelligence arrived on Friday. It is that computer use, formal math, chip design, and world models all moved in the same six days, and they are no longer separate stories.