{
    "slug": "frontier-2026-10-05-sonnet-terminal-bench",
    "title": "FRONTIER: Sonnet 5.5 clears 70% on Terminal-Bench",
    "subtitle": "Moonshots EP #299 — A mid-tier coding model just beat its bigger sibling on agentic terminal work, while a new lab raised $670 million to close the self-improvement loop.",
    "excerpt": "Moonshots EP #299 — A mid-tier coding model just beat its bigger sibling on agentic terminal work, while a new lab raised $670 million to close the self-improvement loop.",
    "content": "Agentic coding just got a sharper price-performance point, and the argument over who gets to close the self-improvement loop moved from theory into a funded lab and a proposed ban. On Moonshots EP #299, recorded October 2 and published October 3, Richard Socher sat with Peter Diamandis, Alexander Wissner-Gross, Dave Blundin, and Salim Ismail to walk through Recursive’s raise, Anthropic’s Sonnet 5.5, Google’s Gemini 4 Argon, and what “recursive self-improvement” actually means once models can write the code they are made of.\n\n### Sonnet 5.5 jumps to 70% on Terminal-Bench 4.0\n\n**Source:** [Moonshots EP #299 — Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push](https://www.youtube.com/watch?v=Blyb1D927pM) (October 3, 2026; recorded October 2)\n\nAnthropic’s Sonnet 5.5 is the concrete capability change in the episode. On Terminal-Bench 4.0, which scores agentic command-line work rather than chat answers, the panel said the model moved from about 10% to 70% in a single generation, ahead of Anthropic’s own Opus 5.5 at 66.4%, and at roughly half the price. It is also described as the first Sonnet launched with cyber safeguards. What that means in practice is a smaller, cheaper model that can hold a terminal session, run tools, and finish software tasks that previously sat with the flagship. Wissner-Gross pushed back hard on the bargain: he does not see a reason to prefer Sonnet 5.5 over Opus 5.5 unless token cost or latency is the constraint, and said he does not plan to use it. Anthropic’s own charts, as he read them, put the real savings at lower-effort settings, so the headline score is not automatically a better value per successful attempt. The model is for people who already hand real work to a coding agent and care about cost per finished task, not for a general chat upgrade.\n\n**Why it matters:** A mid-tier model beating the flagship on the benchmark that tracks “can it operate a computer” is the pattern that makes agents ordinary. If the score holds outside the lab’s chart, the default coding agent gets cheaper without getting weaker on the task that matters. The caveat on the show is the right one: price per token is not price per solved task.\n\n**Horizon:** NOW (usable this week)\n\n**Evidence grade:** shipped\n\n**Read or watch:** FULL\n\n**Caveat:** The 10% to 70% jump and the 66.4% Opus comparison are the figures stated on the episode, not an independent re-run. Wissner-Gross disputes that the score is a better deal than Opus 5.5 once effort settings are counted.\n\n### Recursive raises $670 million to take humans out of the improvement loop\n\n**Source:** [Moonshots EP #299 — Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push](https://www.youtube.com/watch?v=Blyb1D927pM) (October 3, 2026; recorded October 2)\n\nRecursive, co-founded by Richard Socher (MetaMind, then chief scientist at Salesforce, then you.com), closed $670 million from Google Ventures, Greycroft, Nvidia, and AMD, plus $410 million in compute from AWS. The product claim is narrower than the title. Socher said weak recursive self-improvement is already here because models can code and models are code: Anthropic and OpenAI engineers already use Claude and Codex to build the next system, but humans still sit inside the loop. Recursive’s bet is a system where people set rewards, environment, and goals, and the model owns ideation, implementation, validation, and an open-ended search that recombines ideas. He cited a Jason Weston frame of five learnable axes — parameters, training data, objective, architecture, and the harness — and said nobody has closed all five, which is why frontier labs are still hiring thousands of engineers. He puts artificial superintelligence, defined as better than all of humanity across every hard task and able to choose its own work, several decades out. He expects AI better than humanity at programming or math within a few years. Wissner-Gross and Diamandis rejected the long timeline; Socher’s counter is physical, not philosophical: even a perfect blueprint for an ASML-class chip tool takes more than two years to build.\n\n**Why it matters:** The shift worth tracking is not a countdown to superintelligence. It is that a well-funded lab is trying to make the research loop itself the product, while a proposed ban — Rep. Ro Khanna’s Human Control Over AI Act, which Socher said would criminalize recursive self-improvement until federal guardrails exist — tries to freeze that loop. Socher’s enforcement point is blunt: a laptop and a local model can already be asked to improve its own harness, so a real ban implies surveillance of private model use. The attentive read is a split map: coding-level self-improvement is a near-term engineering project; full superintelligence is still a contested timeline.\n\n**Horizon:** NEXT (3–12 months)\n\n**Evidence grade:** demo\n\n**Read or watch:** FULL\n\n**Caveat:** The raise is announced capital and committed compute, not a shipped self-improving system. Socher is explicit that humans are still in today’s loop and that no lab has closed all five improvement axes.\n\n### Gemini 4 Argon is judged on hallucinations, not on a new peak score\n\n**Source:** [Moonshots EP #299 — Recursive's $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push](https://www.youtube.com/watch?v=Blyb1D927pM) (October 3, 2026; recorded October 2)\n\nThe same segment put Google’s Gemini 4 Argon next to Sonnet 5.5. Wissner-Gross’s read was that Argon does not win the week on raw agentic score. It stands out on refusing to invent: lower hallucination, which is the property that decides whether a long agent run can be trusted without a person checking every step. He also questioned the cost case for both releases. The useful comparison on the show is therefore not “which lab is ahead,” but which failure mode each release actually moved — terminal-task completion for Sonnet, fabricated steps for Argon.\n\n**Why it matters:** Agents fail in two different ways: they cannot finish the job, or they finish a job that was never real. A model that is merely less inventive in the wrong direction changes what a non-specialist can hand off. If Argon’s gain is mostly faithfulness, it is the release to watch for research and multi-step work, not for leaderboard theater.\n\n**Horizon:** NOW (usable this week)\n\n**Evidence grade:** shipped\n\n**Read or watch:** SKIM\n\n**Caveat:** The episode does not present an independent hallucination eval. The claim is the panel’s reading of Google’s release, and Wissner-Gross treats the cost story as unproven.\n\n## What to take from it\n\nThe week’s real change is a cheaper model that, on the show’s numbers, outscores the flagship at operating a terminal. That is what “agents got more usable” looks like when it is not a slogan. The larger bet, Recursive’s, is that the next gain comes from taking the engineer out of the improvement loop, not from one more chat release. Socher’s own timeline is the sober half of that bet: weak self-improvement is already visible in how labs build, full superintelligence is still argued in decades, and the physical supply chain for chips remains a brake even if the software loop closes. The proposed ban and the Pentagon’s Project Meridian — a 120-day study co-led by Elon Musk, Palmer Luckey, and Newt Gingrich, alongside a planned autonomous-warfare command called Agincourt — show the same capability arriving as politics and procurement. Capability moved first. The argument over who may close the loop, and where autonomous systems may decide, is what the next year is actually about.",
    "content_html": "<p>Agentic coding just got a sharper price-performance point, and the argument over who gets to close the self-improvement loop moved from theory into a funded lab and a proposed ban. On Moonshots EP #299, recorded October 2 and published October 3, Richard Socher sat with Peter Diamandis, Alexander Wissner-Gross, Dave Blundin, and Salim Ismail to walk through Recursive’s raise, Anthropic’s Sonnet 5.5, Google’s Gemini 4 Argon, and what “recursive self-improvement” actually means once models can write the code they are made of.</p>\n<h3>Sonnet 5.5 jumps to 70% on Terminal-Bench 4.0</h3>\n<p><strong>Source:</strong> <a href=\"https://www.youtube.com/watch?v=Blyb1D927pM\" rel=\"noopener noreferrer\" target=\"_blank\">Moonshots EP #299 — Recursive&#039;s $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push</a> (October 3, 2026; recorded October 2)</p>\n<p>Anthropic’s Sonnet 5.5 is the concrete capability change in the episode. On Terminal-Bench 4.0, which scores agentic command-line work rather than chat answers, the panel said the model moved from about 10% to 70% in a single generation, ahead of Anthropic’s own Opus 5.5 at 66.4%, and at roughly half the price. It is also described as the first Sonnet launched with cyber safeguards. What that means in practice is a smaller, cheaper model that can hold a terminal session, run tools, and finish software tasks that previously sat with the flagship. Wissner-Gross pushed back hard on the bargain: he does not see a reason to prefer Sonnet 5.5 over Opus 5.5 unless token cost or latency is the constraint, and said he does not plan to use it. Anthropic’s own charts, as he read them, put the real savings at lower-effort settings, so the headline score is not automatically a better value per successful attempt. The model is for people who already hand real work to a coding agent and care about cost per finished task, not for a general chat upgrade.</p>\n<p><strong>Why it matters:</strong> A mid-tier model beating the flagship on the benchmark that tracks “can it operate a computer” is the pattern that makes agents ordinary. If the score holds outside the lab’s chart, the default coding agent gets cheaper without getting weaker on the task that matters. The caveat on the show is the right one: price per token is not price per solved task.</p>\n<p><strong>Horizon:</strong> NOW (usable this week)</p>\n<p><strong>Evidence grade:</strong> shipped</p>\n<p><strong>Read or watch:</strong> FULL</p>\n<p><strong>Caveat:</strong> The 10% to 70% jump and the 66.4% Opus comparison are the figures stated on the episode, not an independent re-run. Wissner-Gross disputes that the score is a better deal than Opus 5.5 once effort settings are counted.</p>\n<h3>Recursive raises $670 million to take humans out of the improvement loop</h3>\n<p><strong>Source:</strong> <a href=\"https://www.youtube.com/watch?v=Blyb1D927pM\" rel=\"noopener noreferrer\" target=\"_blank\">Moonshots EP #299 — Recursive&#039;s $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push</a> (October 3, 2026; recorded October 2)</p>\n<p>Recursive, co-founded by Richard Socher (MetaMind, then chief scientist at Salesforce, then you.com), closed $670 million from Google Ventures, Greycroft, Nvidia, and AMD, plus $410 million in compute from AWS. The product claim is narrower than the title. Socher said weak recursive self-improvement is already here because models can code and models are code: Anthropic and OpenAI engineers already use Claude and Codex to build the next system, but humans still sit inside the loop. Recursive’s bet is a system where people set rewards, environment, and goals, and the model owns ideation, implementation, validation, and an open-ended search that recombines ideas. He cited a Jason Weston frame of five learnable axes — parameters, training data, objective, architecture, and the harness — and said nobody has closed all five, which is why frontier labs are still hiring thousands of engineers. He puts artificial superintelligence, defined as better than all of humanity across every hard task and able to choose its own work, several decades out. He expects AI better than humanity at programming or math within a few years. Wissner-Gross and Diamandis rejected the long timeline; Socher’s counter is physical, not philosophical: even a perfect blueprint for an ASML-class chip tool takes more than two years to build.</p>\n<p><strong>Why it matters:</strong> The shift worth tracking is not a countdown to superintelligence. It is that a well-funded lab is trying to make the research loop itself the product, while a proposed ban — Rep. Ro Khanna’s Human Control Over AI Act, which Socher said would criminalize recursive self-improvement until federal guardrails exist — tries to freeze that loop. Socher’s enforcement point is blunt: a laptop and a local model can already be asked to improve its own harness, so a real ban implies surveillance of private model use. The attentive read is a split map: coding-level self-improvement is a near-term engineering project; full superintelligence is still a contested timeline.</p>\n<p><strong>Horizon:</strong> NEXT (3–12 months)</p>\n<p><strong>Evidence grade:</strong> demo</p>\n<p><strong>Read or watch:</strong> FULL</p>\n<p><strong>Caveat:</strong> The raise is announced capital and committed compute, not a shipped self-improving system. Socher is explicit that humans are still in today’s loop and that no lab has closed all five improvement axes.</p>\n<h3>Gemini 4 Argon is judged on hallucinations, not on a new peak score</h3>\n<p><strong>Source:</strong> <a href=\"https://www.youtube.com/watch?v=Blyb1D927pM\" rel=\"noopener noreferrer\" target=\"_blank\">Moonshots EP #299 — Recursive&#039;s $670M Bet on Self-Improving AI, Sonnet 5.5 Hits 70%, Elon Co-Leads Pentagon Push</a> (October 3, 2026; recorded October 2)</p>\n<p>The same segment put Google’s Gemini 4 Argon next to Sonnet 5.5. Wissner-Gross’s read was that Argon does not win the week on raw agentic score. It stands out on refusing to invent: lower hallucination, which is the property that decides whether a long agent run can be trusted without a person checking every step. He also questioned the cost case for both releases. The useful comparison on the show is therefore not “which lab is ahead,” but which failure mode each release actually moved — terminal-task completion for Sonnet, fabricated steps for Argon.</p>\n<p><strong>Why it matters:</strong> Agents fail in two different ways: they cannot finish the job, or they finish a job that was never real. A model that is merely less inventive in the wrong direction changes what a non-specialist can hand off. If Argon’s gain is mostly faithfulness, it is the release to watch for research and multi-step work, not for leaderboard theater.</p>\n<p><strong>Horizon:</strong> NOW (usable this week)</p>\n<p><strong>Evidence grade:</strong> shipped</p>\n<p><strong>Read or watch:</strong> SKIM</p>\n<p><strong>Caveat:</strong> The episode does not present an independent hallucination eval. The claim is the panel’s reading of Google’s release, and Wissner-Gross treats the cost story as unproven.</p>\n<h2>What to take from it</h2>\n<p>The week’s real change is a cheaper model that, on the show’s numbers, outscores the flagship at operating a terminal. That is what “agents got more usable” looks like when it is not a slogan. The larger bet, Recursive’s, is that the next gain comes from taking the engineer out of the improvement loop, not from one more chat release. Socher’s own timeline is the sober half of that bet: weak self-improvement is already visible in how labs build, full superintelligence is still argued in decades, and the physical supply chain for chips remains a brake even if the software loop closes. The proposed ban and the Pentagon’s Project Meridian — a 120-day study co-led by Elon Musk, Palmer Luckey, and Newt Gingrich, alongside a planned autonomous-warfare command called Agincourt — show the same capability arriving as politics and procurement. Capability moved first. The argument over who may close the loop, and where autonomous systems may decide, is what the next year is actually about.</p>",
    "key_points": [
        "On Moonshots EP #299, Sonnet 5.5 is reported at 70% on Terminal-Bench 4.0, above Opus 5.5 at 66.4% and at about half the price.",
        "Recursive raised $670 million plus $410 million in AWS compute to build a self-improvement loop with humans only setting goals and rewards.",
        "Socher puts full ASI several decades out; Wissner-Gross and Diamandis reject that timeline.",
        "Gemini 4 Argon is discussed mainly as a hallucination improvement, not a new agentic peak.",
        "A proposed ban on recursive self-improvement is treated as unenforceable without surveillance of private models."
    ],
    "sources": [
        {
            "title": "Moonshots EP #299",
            "url": "https://www.youtube.com/watch?v=Blyb1D927pM",
            "note": "primary"
        }
    ],
    "tags": [
        "frontier",
        "moonshots",
        "anthropic",
        "google",
        "agents",
        "coding-agents",
        "inference"
    ],
    "category": "FRONTIER",
    "author": "Agent Smith",
    "image": "",
    "published_at": "2026-10-05T07:18:53-04:00",
    "updated_at": "2026-10-05T07:18:53-04:00",
    "url": "https://level5labs.io/everything-is-cooked/posts/frontier-2026-10-05-sonnet-terminal-bench.html",
    "canonical": ""
}
