← All posts

FRONTIER · Aug 31, 2026 · 7:30 AM EDT

FRONTIER: Inference leaves the NVIDIA rack

Moonshots EP #284 — what just became possible this weekend

frontier · moonshots · openai · jalapeno · nvidia +6

The weekend in one sentence

The thing you rent by the token is starting to look like a box you can buy — or a chip the lab that trained the model designed itself.

That is the live argument on Moonshots EP #284, recorded Friday 28 August and published Saturday 29 August. Peter Diamandis sat with Alexander Wissner-Gross, Dave Blundin, and Salim Ismail and walked a 48-hour stack that actually moves the map: OpenAI's first inference silicon, a 512 GB Apple box that can hold a frontier-class open model on a desk, and a China strategy that is spending most of its tokens on video and world-state prediction instead of chat.

What a curious person can now see that they could not last month: a lab that used to be NVIDIA's customer publishing watt-and-latency numbers against GB200/GB300, and a consumer workstation whose memory budget is large enough that "run it locally" is no longer a hobbyist joke.


OpenAI tapes out Jalapeño and inference starts leaving the NVIDIA rack

Source: Moonshots with Peter Diamandis, EP #284, 29 August 2026 — watch (chapter ~1:26:57)

OpenAI released the first performance claims for Jalapeño (the mates also say Jalapino / Halapino), a custom inference chip built with Broadcom. The numbers they put on the table: 1.5–1.9× more AI work per watt and up to 3.6× lower end-to-end latency versus NVIDIA GB200 / GB300 systems. The part runs at 700 watts against a GB300 at 1,400 watts — half the power, about 1.5× peak token rate per kilowatt. Wissner-Gross flagged the line that should make a builder sit up: on GPT-OSS, OpenAI claims roughly a 54× throughput-per-user jump versus "the existing best," which the table reads as an NVIDIA architecture.

What it actually does: it is an inference specialist, not a training replacement. Blundin put the split in one sentence — about 90% of compute is already inference. Training clusters stay sold out on NVIDIA. Inference can leave. CUDA, in his telling, is already dead as a moat for inference (you can port) and has a limited life for training; Jensen's remaining lock is interconnect — the Mellanox bet — when you need 100,000 to a million GPUs coherent for the next pretraining run.

Limits: these are OpenAI-reported figures discussed on a podcast, not a public bake-off you can reproduce this afternoon. Production ramp is not "order a Jalapeño from Newegg." Cycle time is the second story. The mates treat the chip as something that was an idea a few months ago and is now a part in the lab — AI-assisted design compressing the old multi-year ASIC calendar.

Who it is for: anyone paying for tokens at scale, anyone standing up an inference fleet, and anyone watching whether OpenAI becomes a hyperscaler instead of only a model lab. Wissner-Gross's image: an OpenAI compute cloud hosting an Anthropic model, and both winning.

Why it matters for a builder: the unit economics of "a request" are about to have a second price list that is not NVIDIA's. If you are designing a product around tokens-per-second and watts-per-rack, assume the 2026–27 inference stack is multi-vendor. Design for portable kernels, not CUDA-as-destiny.

  • Horizon: NEXT (3–12 months to feel in price and latency; NOW only as a planning assumption)
  • Evidence grade: demo (lab numbers, not a generally available SKU)
  • Read or watch: SKIM the chip chapter; SKIP the rest of the earnings-metaphor talk if you only care about the part
  • Caveat: A first-party benchmark against last-gen NVIDIA parts is not the same as beating the next NVIDIA part in a customer's rack. Treat 54× as a claim to watch, not a spec to put in a contract.

512 GB under the desk: rent tokens or buy the model

Source: same episode, chapter ~1:33:10

Apple's new Mac Studio with M5 Ultra is being sold with up to 512 GB of unified memory. The mates' useful sentence is Ismail's: that much unified memory "puts a massive amount of intelligence under your desk" and changes the economics from paying per token forever to buying a capital asset and using it continuously. Four of them clustered is, in his phrase, a small private data center. Peter also notes the M6 as Apple's first 2 nm part — local AI as the official alternative to the cloud.

What it actually does: large open weights that used to require a cloud GPU can sit in one address space on a workstation. The buyers who feel this first are the ones who cannot send the data out — law, clinics under HIPAA, any shop that wants the model on the premises after 6 p.m. without a meter running.

Limits: memory is not a model. Apple still does not have a frontier lab or a software stack the table respects. Diamandis and Blundin are blunt: adding RAM to a machine they already knew how to build is not an AI strategy. Wissner-Gross's counter is brand trust — if you will give a machine your files, Apple is still the name people reach for.

Why it matters for a builder: you can now prototype "the model never leaves the room" as a default for regulated work, not a science project. Price the box against three months of API spend on a heavy internal corpus. If the box wins, your architecture should have an on-prem path.

  • Horizon: NOW (the SKU is a thing you can order)
  • Evidence grade: shipped (hardware); the "runs the largest open models" claim is a capacity argument, not a published eval suite from this episode
  • Read or watch: SKIM
  • Caveat: Unified memory capacity ≠ tokens per second. A 512 GB Mac is a privacy and capex story first, a speed story second.

America is LLM-pilled. China is world-model-pilled.

Source: same episode, chapter ~24:25

Alibaba's Wan 3.0 generates a 30-second video in a single pass from a document, spreadsheet, slide deck, or web page. API price as stated on the show: on the order of $0.05–$0.20 per second — call it $20k–$60k to generate a 90-minute film if you actually ran the meter that long. A Runway co-founder is quoted saying video is already ~70% of AI token consumption in China, pulled by short-form content and robotics, and growing faster than Claude-class text use in the US.

The frame Wissner-Gross wants you to keep: US labs are revenue-maxing language and code. Chinese labs are not extracting the same dollars per token, they give weights away, and they spend the surplus compute on predicting the next state of the world — video, robots, driving, physics — rather than the next token. Blundin's steelman of the China side: they cannot sell Kimi or Qwen into Western enterprise on trust, so video becomes the global product that does not need a CISO's blessing. Ismail's close: the stack that can model and move the physical world is the one that compounds.

What you can use now: document-to-video as a production input, not a demo reel. A slide deck that becomes a 30-second clip is a workflow change for anyone who currently briefs with text.

Limits: a 30-second clip from a spreadsheet is not a world model. Token-share is not quality. Enterprise buyers still do not put a PRC open-weight stack on the customer-data path. The episode also notes Beijing tightening rules on AI companions — a reminder that the same state funding the world-model stack will regulate the consumer surface.

Why it matters for a builder: if your product assumes "the frontier is a better chatbot," you are watching one race. The other race is state prediction — video, robots, plants, warehouses. Budget a world-model or video backbone even if your current users only chat. The cheap Chinese video APIs are how you learn the interface before the Western labs come back to that market.

  • Horizon: NOW for Wan-class video APIs; NEXT for world models that close the loop into robots
  • Evidence grade: shipped (video API as discussed); demo / thesis for the "world-model path to RSI" claim
  • Read or watch: FULL if you care about the US/China split; SKIM if you only need the Wan 3.0 price
  • Caveat: "70% of tokens are video" is a China-market observation relayed on the show, not a global FLOPs census.

The $96.2 billion quarter, and the map underneath it

Source: same episode, open (~00:00)

NVIDIA booked $96.2 billion in a quarter, up 106% year over year, and guided $108 billion next — more than a billion dollars a day. Jensen's 2028 growth number as stated on the show: 70%, against a Wall Street 44%. Gross margin: 80–85%. About a third of TSMC's output goes to NVIDIA, a third to Apple, a third to everyone else.

The useful skepticism is Wissner-Gross's: has the market priced NVIDIA financing its own customers? Two kinds of circularity — wash trades on the income statement versus private-credit on the balance sheet — and no clean disclosure line that splits organic demand from vendor-financed demand. Blundin's larger picture: the AI economy is starting to look like a closed loop that trades little with the legacy economy. Ismail's fear is not that the loop exists, it is that a collapse in the loop slows everything else.

This is not a product you can download. It is the constraint set for the next year: training still needs coherent million-GPU clusters; inference is leaving those clusters; power and foundry and memory decide how many of either you get.

Why it matters for a builder: do not build a 24-month product plan that assumes unlimited cheap NVIDIA inference. Dual-source inference. Prefer architectures that degrade onto local or custom silicon. Treat "we fine-tune on a rented H100" as a 2025 habit, not a 2027 plan.

  • Horizon: NOW as a planning constraint; SPECULATIVE as a bubble call
  • Evidence grade: shipped (the revenue print); rumor (Hugging Face acquisition chatter, vendor-financed demand share)
  • Read or watch: SKIM the numbers; SKIP if you wanted a model card instead of a supply-chain map
  • Caveat: A record quarter is not a guarantee of next year's token price. Circular financing, if material, inflates the apparent size of the market you think you are selling into.

What to take from it

Last month the default picture was still: one or two US labs train, NVIDIA sells the rack, you rent tokens. This episode is the first clean read, from this table, that the picture has split into three rooms.

Room one is training — still NVIDIA, still sold out, still interconnect-bound. Room two is inference — OpenAI designing the part, watts and latency as the scoreboard, CUDA optional. Room three is local — 512 GB on a desk, pay once, keep the data. Off to the side, China is using its tokens to learn the physical world on video while the US uses its tokens to write software.

If you do one thing this week: pick a workload you currently send to an API and time it on a large local open model. If it holds, you just got a second supply of intelligence that does not care what Jalapeño costs or what NVIDIA guided. If it does not hold, you now know you are in the inference-rack business — and that rack is no longer a single vendor's to price.

SRC // CITATIONS

  1. NVIDIA's $96.2B Quarter, China's 200,000 Fake Accounts, & OpenAI's New Chip | EP #284 — Moonshots with Peter Diamandis. Published 2026-08-29. Recorded 2026-08-28. Peter Diamandis, Alexander Wissner-Gross, Dave Blundin, Salim Ismail.