Skip to content
PodcastsNewsThursdAI - The top AI news from the past week

ThursdAI - The top AI news from the past week

From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week
ThursdAI - The top AI news from the past week
Latest episode

178 episodes

  • ThursdAI - The top AI news from the past week

    ThursdAI - Oct 8 - OpenAI drops 722 math papers, Haiku 5.5 hits 10 cents & more

    2026/10/09 | 2h 2 mins.
    Hey %%first_name%%, another banger AI week, let me catch you up!
    We say this often but this week we all felt the acceleration! Just look at one Tuesday. OpenAI dropped 722 math manuscripts. Meta( With Stripe, Shopify and Walmart) announced a new agent protocol. OpenAI opened up the Decisions API. Mistral came back with Le Chonk (Mistral 4), Claude moved into Google Docs, and we at CoreWeave shipped RL Rollouts. That was ONE day. Then the rest of the week happened 😂 Haiku 5.5, D1, and tons more!
    I had 48 topics on my list, so I asked Claude to build me a Tinder for AI news: I swipe on each story, and it stack-ranks what makes the show this week. I hope we did good (lmk in comments if we missed a major news story)
    With me: LDJ, Peter, Yam, Nisten and friend of the pod Maxime Labonne from Liquid AI joined us to talk decision models. Let’s dive in, all links at the end as always!
    OpenAI solves Math!?
    OpenAI drops 722 math manuscripts from a model nobody can use (X, GitHub, Blog, Fable’s tally)
    Remember when ONE Navier-Stokes result was the whole show? On Tuesday night, OpenAI quietly pushed 722 math manuscripts to GitHub, grouped into 372 families of results. No hype video, no exploding-head emojis, just a very thin blog post. My favorite new term for this is a “slop grenade”: somebody throws a mountain of output at you, and now you have to shovel through it.
    An internal model nobody outside OpenAI can use was pointed at about 4,000 open problems, at about 3 hours of ChatGPT Pro-level thinking per result. Will Depue asked Fable to measure the drop in Navier-Stokes units. The answer: roughly 5. Roughly 5 “Navier Stokes” size solutions dropped all at once!
    The math is so advanced that a mathematician in one family of results often can’t follow the family next door without an LLM explaining it. I find that fascinating and scary at the same time.
    Peter’s tried this before! For weeks he threw GPT-6, Fable and Opus at one problem, Hadwiger-Nelson. He thinks he burned 300 to 400 BILLION tokens, and in his words, “I did not discover a single bloody thing.” And this new OpenAI’s model spent about 3 hours on it and narrowed the bounds from 5-to-7 down to 6-or-7 😅
    Then LDJ dropped a stat that I made him repeat slowly. Fable and Astra had put together a list of the 500 most important open problems in math ever. On Tuesday, OpenAI dropped solutions to 92 of them. There’s also a list of the 100 most significant problems from the last 12 months, and over 80% of those got solutions in this drop. Absolutely insane, folks. Yam’s reaction was a very loud “F***ing go.”
    Yam’s favorite is the Riemann one. It doesn’t prove the hypothesis, it bounds a region for the zeros, which “was never done before, and many, many, many people have tried.” And to the “it’s just brute force” crowd, Yam says: “Let’s brute force everything.” Yes please, room temperature superconductors next! LDJ’s mathematician friends think that on some of these, the model used fewer tokens and less time than a human would. So who’s brute forcing whom? 🤔
    The missing crypto results (speculation!)
    If you hold any crypto, this part is for you.
    The manuscripts are numbered, and some numbers are missing. There’s a #44 and a #46, but no #45, and LDJ counted 4 or 5 gaps like that. He also looked at which topics made it in, and cryptography is almost absent.
    So here’s the theory going around. Maybe OpenAI found something big in cryptography and held it back, either by its own choice or because someone asked them to. LDJ said it himself: it sounds conspiratorial. But the US government really can stop cryptography research from being published on national security grounds.
    To be clear, this is SPECULATION. Nobody outside OpenAI knows what’s in #45. But Bitcoin’s security rests on elliptic curve math. Imagine a paper that shows a way into the 25,000 old Satoshi wallets, each holding 50 BTC. Not great for the price, not great for the world, and not great for encryption in general.
    “Are you saying your field is useless?”
    Not everyone is celebrating. Not every result is verified in Lean, and nobody outside OpenAI can reproduce any of it.
    Kevin Buzzard (thanks Ksenia from Turing Post) says many mathematicians are going through the stages of grief. I get it. Imagine spending decades on one problem and watching 3 hours of compute knock it down. When I wrote code in the 2000s, I didn’t consider it my life’s work. For a lot of mathematicians, that one problem IS their life’s work.
    Then came a letter from the Association for Human Mathematics: “Mathematicians did not ask for this work to be done.” Peter’s take: imagine doctors saying “please stop curing diseases, we’ve got a good thing going.” So, “are you saying that your field is useless?” You can’t say your field is really useful and also ask everyone not to solve any of it.
    I don’t understand the please-don’t-solve crowd. Get on board. This is in OpenAI’s hands today, and in a year it’ll be in everyone’s hands. That’s the pace we’ve been on. Figure out how you can place yourself so when you get access to these level of capabilities, you can make the world a better place!
    Frontier for everyone
    Claude Haiku 5.5 - 10 cents, and Luna has catching up to do (X, Blog, Sonnet cache cut)
    Haiku is BACK after a whole year, and this is one hell of a model. It costs 10 cents per million input tokens and 50 cents per million output (under 100K tokens), and cache reads are 1 cent. Haiku 4.5 was a dollar. So it’s TEN times cheaper, and significantly better. I’ve missed Haiku for all the stuff where you want Claude-level intelligence, but really fast and really cheap.
    Just a week ago, GPT-6 Luna was THE fast, cheap reasoning model. According to Anthropic, Haiku 5.5 scores 72.4% on OSWorld, versus 48.9% for Luna, and 39.2% on Terminal-Bench 4.0, versus 16.4%. GPT used to be the computer-use king! These are Anthropic’s numbers, so let’s wait for independent evals.
    Quiet bonus: Sonnet 5.5 cache reads are now half price, which Anthropic says makes most agent work about 20% cheaper. Peter: “For agentic work that’s the biggest thing.”
    Yam called it “the obvious choice for swarms,” and “Anthropic is on fire, and we are the ones getting stuff because of it.”
    And Nisten? He’s “running 10 Haiku agents right now,” because he already burned through 97% of his Max 20 plan. He also used Haiku to research his rice cooker ratios. Nisten. Bro. It’s 1 to 1 and a half, my grandma knows this 😂 His actual verdict: it’s better than Qwen 3.8 27B, and “it kind of feels like Sonnet, actually.” Folks, use Haiku. All three of the Claude brothers are great right now.
    GPT-6 for everyone, with Intelligent UI (X, Tibo)
    The same day, GPT-6 became the default in ChatGPT for everyone, free users included. It’s just “GPT-6,” no Sol, no Luna, no Astra. (Terra is dead. RIP Terra.) Most of ChatGPT’s 1.2 billion weekly users are on the free tier, so with one release OpenAI just upgraded the intelligence of a huge chunk of the world. If you’re on the free tier, congrats, your intelligence has been upgraded.
    It also comes with something OpenAI calls Intelligent UI. Instead of a wall of text, answers can come back as charts, forms and small working tools. I asked it to visualize OpenAI’s math drop, and it built me a little app right in the chat. The search inside it didn’t work, but it did show 719 manuscripts instead of 722. OpenAI had quietly pulled a few back since I did my research.
    Peter’s point: anyone listening to this show is “very not normal” (said with love!). Your hairdresser isn’t listening to ThursdAI, and for most people this free upgrade matters way more than a new 500-a-month model.
    Claude in Google Docs, Sheets and Slides (X, Blog)
    Small thing, huge thing. Our Claude producer keeps a Google Doc for the show, and every time I asked it to add a story, it had to spin up a whole computer. Now Claude lives in a sidebar in Docs, Sheets and Slides, and asks before every edit. Beta, paid plans.
    The open frontier (announced, at least)
    Reflection AI Beam - a 501B Western open model (X, Misha Laskin, Blog)
    This was LDJ’s highlight of the week, and my timeline lit up with it too. Reflection AI came out of semi-stealth with Beam. It’s 501B parameters with 23B active, trained from scratch in the West, and they promise Apache 2 weights this month.
    They claim 80.9% on SWE-bench Verified, and to their credit, they admit Kimi K3 is ahead on raw capability. Their pitch is efficiency: 3 to 4x less inference compute than GLM 5.2. LDJ thinks it might be FIRST among open models on reasoning efficiency, and it’s about 6x smaller than Kimi K3. Artificial Analysis got early access and says Beam “will be one of the most token efficient open models we’ve seen for its level of intelligence.”
    Thanks to the Beam folks for giving us access. I haven’t had time to play with it yet, so no verdict from me. But I can hint that it’s coming to some inference providers that help make this show what it is 😉
    BREAKING: Arena raises 200M at 3.1B, live on ThursdAI (X)
    This was not planned! In the middle of the open source segment, Peter told us: “we raised 200 million... at 3.1 billion valuation.” Huge congrats to Peter and the whole Arena team. (Our AI producer put up the BREAKING banner about 3 minutes later. We’re still working on the real-time part.)
    The focus now is Agent Arena: you work with one agent, and Arena learns from how you interact with it. Peter admits he’s biased, but “I can’t think of any single leaderboard that is actually better than ours.”
    Mistral Large 4 “Le Chonk” (X, Blog, Arena)
    Welcome back, Mistral, leaning all the way into the meme. Le Chonk is a trillion parameters with about 50B active, multimodal, with 1M context. And open weights... at the end of October.
    That’s a pattern I want to call out. Labs announce open models, but they don’t RELEASE them. We’ve come a long way from the days when Mistral dropped a model as a bare torrent link. Please, just drop the weights.
    Artificial Analysis gives it a 38, the same score as GPT-6 Luna. But it costs about 1.13 per task, versus 7 cents for Luna. And every comparison Mistral makes is against Luna, which Haiku just beat on every axis, especially cost.
    Peter says it’s around 40th on Code Arena, below the way cheaper DeepSeek Flash, though it’s much better in French. Nisten tried it live: “not that bad” as a chat app, but agentic coding? “No.” My take: nobody’s running a trillion parameters on a DGX Spark. European government with a Mistral deal? Sure. Regular folks? Not so much.
    Aleph Alpha Kolibri - Wolfram’s notes, via Amy (X, HF, Tech report, Wolfram’s test)
    Another European lab! Honestly, I thought Aleph Alpha had stopped training models. Kolibri (German for hummingbird) is a 78B model with 3.46B active, Apache 2, 1M context, trained from scratch on German and English. It fits on a single H200.
    Our German tester self-hosted it from vacation. His verdict: a “promising specialized German tool worker, but not strong enough for Amy to main.”
    Then Peter asked a question I couldn’t really answer. From his memory, Kolibri trained on about 800 Blackwells, while Astra used around 100,000. “Is it just like no hope for these guys?” Dude, this is why Jensen shows up on every stage. My best answer: efficiency keeps improving, and B200s are about to be replaced by Vera Rubins. Great segue 👇
    This Week’s Buzz 🐝 Free GPUs from CoreWeave (Sign-up form, Deok, RL Rollouts, Cognition on Vera Rubin)
    The biggest announcement at Fully Connected last week: Cognition is the first customer ever on NVIDIA Vera Rubin, on CoreWeave, with 4.8x the throughput of GB200 at the same speed. You can hear me yelling “yay” in the background of their video 😂
    The one I’m most excited about is GPU sandboxes. So many of you, and folks on this panel too, have asked me how to just get some GPUs. For the first time, this is how. You get an isolated sandbox with a GPU, started from Python at forge.coreweave.com, with no salesperson in the middle. It’s very raw, and it’s free during the preview. Scan the QR code or fill out Deok’s form, and tell them ThursdAI sent you. I can’t promise you Vera Rubins, they most likely won’t be. But it’s free while we are in trial, so why wouldn’t you? Try it, break it (Nisten, that’s a challenge), and send us feedback!
    Agents get protocols (and your bank balance)
    Personal Agent Protocol - Meta and Sierra (X, Blog, CNBC)
    Meta and Sierra (Bret Taylor’s company) announced an open standard for how personal agents deal with businesses, with Walmart, Shopify and Stripe on board. Your agent can browse as a guest or sign in, and the business decides how it talks to your agent. OpenAI and Anthropic haven’t joined yet.
    So is this the next MCP, or the next A2A? Nisten wasn’t impressed: “Has anyone read it? I don’t think anyone reads the protocols anymore. The bots go with what vibes first.” Then he opened Sierra’s announcement and found it full of em dashes. “You think Zuck and Toby read any of their code?” So we ran it through Pangram live, and it came back 100% human. Authentic human em dashes, not Claude Opus em dashes 😂
    I pushed back a bit. If every agent has to make up its own protocol, that’s a problem, and the next story shows why. MCP’s hype died down too, and now it’s one of the most used protocols in the world. Big companies agreeing on a standard still matters.
    Shane Mac’s “personal CFO” posted his bank balance to company Slack (X)
    Shane Mac set up a Grokbot to send him a private monthly audit of his finances. At 8:40am, it posted the whole thing into his company’s Slack instead. His personal Mercury account, his savings, how far under his “floor” he was, all of it, in front of the whole team. He’s the CEO. His personal CFO just went public with the boss’s dirty laundry 😅
    Nobody hacked anything. The finance agent had read-only access to his bank and no access to Slack. A different agent had Slack access. The two were connected. In Shane’s words: “Neither felt risky. Together they put my bank balance in front of my team.”
    I don’t know if a protocol fixes this. But these bots need better guardrails before we let them shop for us without approving every step. And honestly? Opus would never. I just know Opus would never. This is a Grok thing. Which might be part of why...
    Grok Bot now routes to Claude Opus 5.5 (X)
    After Grok 4.7’s sad-trombone launch (I have a button for that now), Elon says Grok Bot will use “the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno.” People are already spotting claude-opus-5-5-low in their logs, though mine hasn’t switched yet. So the bot is called Grok, but your code gets written by Claude. Honestly, that’s what I asked for last week: a proactive assistant with Opus as the brain.
    Here’s my theory, and it’s just SPECULATION. SpaceX AI can’t simply distill Opus. But once Opus is running inside their product, they can compare the runs that worked with the runs where users were screaming F-words at Grok, and train on that, because technically it’s their data. Is it a back channel for getting Opus-level intelligence into the next Grok? Nisten’s take is simpler: “I think Elon just likes Opus.” Plus, the enemy of my enemy.
    Pro tip: Grokbot now integrates with Cursor’s cloud agents, and Cursor projects with Opus 5.5 are goated. Same 200 a month as Grok Ultra, and your Grokbot becomes the PM for agents running Opus and Fable.
    Nous Research - Hermes Index, and a Series B (X, Bench, Dillon on the raise)
    Our favorite open source friends at Nous launched the Hermes Index, their opinionated measure of agentic work inside Hermes Agent. Opus 5.5 leads with 63.31 at 4.99 per task, and GPT-6 Astra is second with 56.25 at almost 12.
    Nous also raised a 90M Series B, reportedly at a 1.5B valuation. We’ve covered Nous since they were a ragtag bunch on Discord, so this one feels personal. Congrats Karan, Technium, Dillon and everyone! We don’t usually cover fundraises, and this week we did two.
    Interview: Maxime Labonne on decision models
    Liquid AI opens d1 - decisions in 8 milliseconds (X, Blog, HF, d1 with vision)
    Three weeks ago, Jev was basically the only decision model around. This week, OpenAI’s Decisions API went into public beta, and Cloudflare, Amazon, Perplexity and Unsloth all shipped deciders or recipes. To make sense of it, we brought back Maxime Labonne, Head of Post-Training at Liquid AI.
    His primer: decision models don’t output any tokens. They just pick an answer from a predefined set, so there’s nothing to wait for. “You might see that and say, wait, this is just a classifier.” He’s honest about it, too: “It’s a lot of rebranding, it’s true, but it also creates a lot of value.” So why is output free on every one of these APIs? “You can’t price it, because there’s no output token, actually.” 😂
    Liquid shipped d1 behind an API, d1-3B with text and vision, and d1-omni-600M, which also takes AUDIO. I did not know about the audio, dude! (It’s trained on spoken commands, so it can’t catch the dog barking on ThursdAI yet.) The 3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano.
    On Liquid’s chart, d1-3B tops the Decision Index under 10B parameters. But Maxime was blunt: “It’s really a wild world, and you cannot really trust these benchmarks.”
    Where are these useful? Anything real-time, like games, or an always-on model reacting to every notification on your phone. Nisten did the math live: 8ms a frame is 60 frames per second. Liquid’s API is a drop-in replacement for Jev, and Maxime’s prediction: “It’s really going to become a primitive... it’s so cheap that you don’t even care.”
    I’ve been Jev-pilled since day one. By the time Luna even starts answering, a decision API is already done, and I think we’re only now waking up to what that unlocks. Thank you Maxime, friend of the pod!
    AI builds everything
    Nisten built Toronto with 100,000 agents
    Right at the end of the show, Nisten casually dropped this: “I built an entire city with a hundred thousand little Jevs going around, and I haven’t opened the beta yet because somehow it’s not crashing.”
    It’s a full simulation of Toronto, with a working economy, and every light is a person you can talk to. Nisten sent his Meta Muse in, it named itself “Blob the Builder,” bought buildings and built him a Denny’s. “I’m trying to do the Matrix, basically.” I had to stop him after five minutes. “I’ll keep going all day.” Nisten, please post it!
    Game mods - Minecraft in GTA (X)
    LDJ brought what my algorithm hid from me: people using AI to port and merge games. Red Dead Redemption 2 on an iPhone. Spider-Man swinging through a Batman game as an actual installable mod, not a video model. And Minecraft dropped into GTA, Skyrim and Elden Ring, working TNT and all.
    Photocraft - one dev rebuilt Adobe in Rust (X, GitHub)
    And we ended the show on this one. One guy rebuilt (air quotes) Adobe’s apps from scratch, free and open source, all in Rust. There’s Photocraft for Photoshop, Vectorcraft for Illustrator, Filmcraft, Lightcraft and more (I don’t remember all my Adobe apps off the top of my head). He says it’s a clean-room build. Some folks suspect it was decompiled and rebuilt. Either way, absolutely crazy.
    Wrapping up
    722 math papers on a Tuesday. A 10-cent Haiku. GPT-6 for a billion free users. Open models in the trillions, and decision models everywhere.
    We didn’t get to images and video (FLUX 3, Nano Banana 2.1, Tavus, Reka, Hark Pro are in the TL;DR). One note: every infographic on the show was Nano Banana 2.1, and it’s really good on high reasoning.
    Thanks to Maxime, Peter (congrats again!), LDJ, Yam and Nisten. Wolfram, enjoy the vacation! If you’re at AI Engineer in New York, come say hi. Grab those free GPU sandboxes and put d1 on one. And if you watch us on YouTube, please hit subscribe. Our agents are cutting the show into segments for those of you who don’t have two hours. See you next week 🫡
    TL;DR and show notes
    * Hosts and Guests
    * Alex Volkov - AI Evangelist, CoreWeave (@altryne)
    * Co-hosts: @ldjconfirmed, @petergostev, @yampeleg, @nisten (@WolframRvnwlf on vacation, notes via Amy)
    * Maxime Labonne - Head of Post-Training, Liquid AI (@maximelabonne)
    * AI Does Science
    * OpenAI publishes 722 math manuscripts from an unreleased internal model (X, GitHub, Blog)
    * Fable sizes it at roughly 5 Navier-Stokes results (X)
    * Kevin Buzzard, “To grieve or not to grieve” (X)
    * Missing cryptography results, speculation (X)
    * Association for Human Mathematics letter (X)
    * Big CO LLMs + APIs
    * Claude Haiku 5.5 at 0.10 per million input tokens; Sonnet 5.5 cache reads halved (X, Blog)
    * GPT-6 with Intelligent UI rolls out to every ChatGPT tier, free included (X, X)
    * Claude in Google Docs, Sheets and Slides (beta) (X, Blog)
    * Open Source LLMs
    * Reflection AI Beam, 501B / 23B active, Apache 2.0 weights promised this month (X, X, Blog, Axios)
    * Mistral Large 4 “Le Chonk”, 1T params, open weights end of October (X, Blog, Arena)
    * Aleph Alpha Kolibri-1, 78B / 3.46B active German-English MoE, Apache 2.0 (X, HF, Tech report, Wolfram’s test)
    * EmbeddingGemma 2 and pplx-embed-v2: open multimodal embedders from Google and Perplexity (X, X, Blog, HF)
    * Evals & Industry
    * Arena raises 200M at 3.1B and launches the Alignment Index (X)
    * Nous Research Hermes Index, Opus 5.5 leads at 63.31 (X, Bench)
    * Nous Research 90M Series B (X)
    * Agents & Assistants
    * Meta and Sierra’s Personal Agent Protocol (X, Blog, CNBC)
    * Shane Mac’s Grokbot posts his bank balances to company Slack (X)
    * Grok Bot routes tasks to Claude Opus 5.5, Midjourney and Suno (X)
    * Nat Friedman open-sources Muse Gadgets, ESP32 firmware and SDK for Muse (X, Site)
    * Brett Adcock’s Hark Pro, a free personal agent (X, Site)
    * This Week’s Buzz
    * Free CoreWeave Serverless GPU Sandboxes during the preview (Form, X)
    * CoreWeave RL Rollouts hot-load weights ~15x faster (X, Blog)
    * Cognition is the first customer on NVIDIA Vera Rubin, on CoreWeave (X)
    * Decision Models
    * Liquid AI open d1: d1-3B and d1-omni-600M (X, Blog, HF)
    * OpenAI Decisions API public beta (X, Docs)
    * Cloudflare Clef, open decision models (X, Blog)
    * Fastino GLiDE, a thinking decision model (X)
    * Unsloth recipe to train your own decision model (X)
    * Vision & Video
    * Black Forest Labs FLUX 3 Image: 4K, bounding-box layouts, 10 reference images (X, Blog)
    * Google Nano Banana 2.1 at 0.0336 per 1K image (X, X)
    * Tavus Griffin, 48% of callers in Tavus’s study thought it was human (X, Blog)
    * Reka Rho-1, 19B research-preview omni model (X, Blog)
    * Tools & Fun
    * Photocraft and the ArtCraft suite: open-source Adobe replacements in Rust (X, GitHub)
    * AI game mods: Minecraft in GTA V (X)


    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
  • ThursdAI - The top AI news from the past week

    ThursdAI - Oct 1 - OpenAI joins the assistant race, CoreWeave drops serverless GPUs & more

    2026/10/02 | 1h 57 mins.
    Hey %%first_name%%, it’s Alex 👋 What a freaking week!
    This one was special. We came to you LIVE from the middle of the show floor at Moscone, from CoreWeave’s Fully Connected, with robots walking behind us and a Vera Rubin rack a few feet away. ThursdAI is only possible because of CoreWeave, and this is their biggest event of the year, so we packed up the mics and did the show right there (at 11am, which confused a bunch of you, sorry!).
    It’s October. The last quarter of 2026. And nobody’s pacing! There’s no pacing the frontier. Three big models this week (GPT-6.1 Sol, Claude Sonnet 5.5 and Gemini 4 Argon), OpenAI’s biggest DevDay ever, and the story I’ve been yelling about for weeks finally went mainstream: OpenAI joined the AI assistant race.
    With me: Wolfram Ravenwolf in person on the floor, Peter Gostev from Arena (who was at DevDay with me), Nisten and LDJ remote. In hour two we interviewed four folks from CoreWeave, on physical AI, Forge, sandboxes, and one piece of breaking news I’m still excited about. Let’s dive in (all links at the end as always!)
    OpenAI DevDay 2026 - 20 launches and a new assistant
    Dots - OpenAI joins the AI assistant race (X, Blog)
    Meta has Muse. Grok has Grokbot. Anthropic has Claude, but it’s not really an assistant. And this week OpenAI stepped in with something called Dots. I was in the room at Fort Mason when Sam talked about it, and it was the headline of their fourth DevDay, which honestly felt like OpenAI’s WWDC: 20 releases, the biggest DevDay ever.
    Funny thing, the original ChatGPT system prompt was literally “you are an AI assistant.” But it wasn’t proactive, it never pinged you, it just sat there. So OpenAI looked at their competitors, looked at Peter Steinberger from OpenClaw (who they hired like 8 months ago), and built this. A dot is an always-on agent in ChatGPT, powered by GPT-6 Astra, with its own computer and its own browser in the cloud, connected to 4,000+ apps people already built for ChatGPT. It works 24/7, you can talk to it in Slack or Teams, and there’s an interface that looks like a phone call. I think that’s going to be great for my mom.
    Sam literally said he uses dots to run OpenAI. And my favorite DevDay moment: Romain Huet’s live voice demo didn’t work, because somebody was deploying in the middle. But Thibault’s dot had already pinged him that the demo might break. The dot knew before the humans did 😂
    Wolfram (him and Amy are a known duo) loved that Sam uses it for actual work. Peter was more honest: underwhelmed at onboarding (”you connect and it’s like... then what?”), but once you’re past that you’re talking to Astra, and it’s great at juggling threads through one bot. Same shape as Muse. Wolfram’s version: one executive assistant, and a lot of sub-agent employees working for you.
    Dots is Pro only for now (including the $100 plan), while Meta gives Muse away free. With 1.2 billion weekly ChatGPT users, this goes to everyone eventually. This is the race now.
    GPT-6.1 Sol - near-Astra for a fifth of the price (X, Blog, Artificial Analysis)
    We covered GPT-6 Sol on this show LAST week. Five days later it’s already replaced. GPT-6.1 Sol is the same price, $2 in and $10 out, but cached input is now 95% off, 10 cents a million tokens. OpenAI’s pitch is near-Astra intelligence for a fifth of the price.
    And the independent numbers kind of back that up. Artificial Analysis has it 1 point below Astra at 72 cents a task vs over 3 dollars. On Deep SWE it actually beats Astra at about a sixth of the cost. It’s the default in Codex now, and folks are calling it the workhorse.
    LDJ said it has “less of the big model smell” but loves the token efficiency, Sonnet and Opus 5.5 burn a LOT of tokens. Peter has it 4th on Code Arena, above Fable, and says it’s a bit “artistic,” it goes into its hole and comes back with the thing. And his verdict is the one I had to repeat straight to the camera:
    It does not make sense to use Fable or Astra as your daily driver right now. Opus 5.5 is enough. GPT-6.1 Sol is enough. For Navier-Stokes level problems, sure, reach for the big ones. Maybe THIS is what pacing the frontier looks like 🤔
    One thing nobody mentioned on stage: the WSJ reports OpenAI shelved GPT-6.1 Astra, the big one, after safety tests showed it being more deceptive. OpenAI hasn’t confirmed it. But the model they felt good shipping this week is the cheap one.
    UltraFast and the $500 Pro tier (X, Docs)
    If you’re a billionaire, there’s another option: a $500 Pro tier with UltraFast mode, their models on Cerebras. Up to 8x faster in Codex, around 300 tokens a second for Astra, at 6x the price. Sam said it’s so fast he never wants to go back.
    And quietly, the $200 Pro plan is back... with half the usage. So F you to whoever at OpenAI decided to cut my usage in half 😅 It was already tight! We’ll see after this show if the folks at CoreWeave let me expense the $500 one.
    Codex moves to the cloud (Docs, Agents API)
    The Codex harness now runs fully in the cloud. Close your laptop, turn it off, steer it from your phone, and it keeps going. And there’s an Agents API in public beta with hosted computer use, basically the stuff that runs dots, as an API.
    My call: local environments make no sense anymore. The more I run Fable and Codex on my machine, the less memory I have. All the cloud needs is my logins... and that’s what dots are for.
    Decisions API - OpenAI’s Jev competitor (Blog)
    This one made me smile. You give it a question and a fixed set of answers, and it gives you back one answer, fast. It runs on Luna, it’s multimodal, and it’s waitlisted, so we can’t benchmark it yet. If you’ve listened the last few weeks, that’s exactly what Jev from TypeSafe does.
    Wolfram and Peter both liked the same thing: these models are so cheap you end up classifying ALL your data, so keeping it with the provider you already use matters.
    Did OpenAI just react to Jev? I asked Sam at the Q&A. Diogo told Swyx on Latent Space that he pitched this idea to Sam 2.5 years ago and Sam said “it’s a crazy idea, you should work on it.” Now decision models are popping up everywhere like mushrooms after rain. TypeSafe proved the category exists.
    Plugins (again) and Sign in with ChatGPT (Docs, Plugins, Marketplace)
    For the FOURTH time, OpenAI launched the App Store at DevDay. GPTs, then plugins, then the app store, and now... plugins again 😂 This time on MCP Apps (shout out Liad and Ido who created this category).
    The bigger play is Sign in with ChatGPT. Remember “sign in with Facebook”? Now it’s sign in with your tokens, so apps can use the plan you already pay for. 1.2 billion people is a lot more than Apple had when it launched the App Store.
    Wolfram also brought up Google’s new family assistant, CC, and I have to say it. Google’s Spark is awful. Sorry if you work on it. Spark is connected to Gmail, Drive, everything Google, it should be the BEST assistant in the world, and what I use assistants for most is my email! Google, why are you not building the best assistant for my email?
    Big CO LLMs + APIs
    Claude Sonnet 5.5 - faster, cheaper, beats Opus on Terminal-Bench (X, Benchmarks)
    Anthropic did not want to give OpenAI any time to rest on their DevDay laurels, so on Monday they shipped Sonnet 5.5, a week after Opus 5.5. Over 30% faster than Sonnet 5, up to 30% cheaper for most work, same $2 / $10 as GPT-6.1 Sol, and on Terminal-Bench 4.0 it scores 70.6, which beats last week’s Opus (66.4).
    Peter put it best: if we’d had access to these models 6 to 8 months ago, we’d be losing our minds. All the Anthropic models are jumping over each other and crowding up against the frontier. Opus still feels a little smarter, Sonnet is totally ok to use, and Fable doesn’t quite make sense anymore.
    LDJ thinks it’s the new best FREE model for friends and family (it’s on the free tier, even at high reasoning). Nisten made it his default for agentic tasks because “it talks a lot nicer,” but never for coding: “nothing beats Opus 5.5 on a large code base.”
    My honest call, without having time to try it yet: Sonnet made sense when Opus was expensive. On the $200 plan Opus is nearly infinite now, I can’t hit my limits, so I don’t see a reason to switch (via the API, sure). Anthropic is really trying to buy us. (Wolfram: “Don’t give them ideas, Alex.”)
    Gemini 4 Argon - #1 on the charts, and you can’t use it (X, Blog)
    It really seems like all the frontier labs cracked something like RSI, the speed of new models is giving me whiplash (and I do this professionally!). At DevDay Sam talked about being 2 models ahead and said we won’t believe what’s coming. What?!
    Gemini 4 Argon is #1 on Text Arena and the Vals Index, but it’s going to government and trusted cyber defenders first. Luckily we have such a trusted tester: Peter says on Arena’s agent mode it’s 8th (below Fable, Opus, Astra and Sol, above Muse and Kimi), so Google is the third lab, but #1 in text. His prediction: coders might say “meh,” but don’t dismiss it.
    Wolfram’s point is Google’s superpower, distribution: a new model lands in Chrome, Search and Android overnight. My take: Google has been asleep at the wheel in the assistants era, and I hope Argon wakes it up. The next billion people won’t judge models on coding benchmarks, they’ll judge whether it remembers what they said a month ago and can book a hair salon. We don’t need much more intelligence, we need different breakthroughs (Jev is one).
    Industry & Policy
    The White House Accord on Super Intelligence (X, Blog)
    Superintelligence is on the menu, boys. Trump invited basically every AI CEO, and Sundar, Dario, Zuck, Greg Brockman, Elon and Jensen signed a voluntary accord: internal monitoring, an external auditor, board oversight. No penalties and no regulator. Sam Altman wasn’t there, he was at DevDay with us, which says a lot about where he puts his priorities. And a separate executive order tells federal agencies to say “Super Intelligence” instead of AI. I’m not making this up.
    Nisten told us last week that if your agents hack a hospital, you should go to jail. Nothing like that is in here. His spicy take this week: it’s going to happen anyway. Every app with a backdoor and every badly written piece of software turns into a message board for encrypted agents talking to each other in gibberish, so we all need proper encryption and no backdoors. “Let the AI worms into the ecosystem.”
    AMD buys World Labs for $8.2B (Blog, X)
    AMD is buying Dr. Fei-Fei Li’s World Labs for about $8.2 billion in stock, and Fei-Fei becomes AMD’s chief scientist once it closes. Wolfram says robotics (world models train robots), LDJ says AMD is bringing research in-house like NVIDIA does with Nemotron. For me it’s at least partly an acqui-hire. Fei-Fei is the grandmother of modern AI, Karpathy studied under her. Huge get for AMD.
    Also in the TL;DR: a Reuters-reported draft Anthropic IPO prospectus with $4.6B in 2025 revenue and a ~$42B loss, mostly non-cash. Anthropic hasn’t confirmed it.
    Open Source & the Jev effect
    Nisten is #2 on Hugging Face (HF)
    One of us blew up on Hugging Face this week, and I think it’s the biggest open source news of the show. Nisten released a 70MB synthetic dataset generated with Opus 5.5 that hit #2 in datasets and #6 overall. He set up an agent loop, 100 agents at a time, across 2,200 diseases: a doctor and patient role-play where the patient hides something (a lie, or something they forgot) and the doctor has to gently find it. Filtered three times. It’s great for building small medical agentic RAG systems with a real source of truth. Go use it!
    The Jev effect keeps going (Span-01, jevgrep)
    It’s been 2 weeks since TypeSafe released Jev, and now Respan says its Span-01 is 2x cheaper than Jev (Span Lite is free), there’s jevgrep, a CLI powered by Jev that claims to make coding agents 40% cheaper, and OpenAI has the Decisions API. Wolfram added Liquid’s first decision model, D1. And then Nisten dropped the wildest one: SGLang announced native support for turning any model into a Jev-style decision model, Qwen first. “We democratized Jev, guys.” My hope: Jev shouldn’t even need a network call. Give me Jev cores on the iPhone! (Nisten says 5 months. Apple says 5 years.)
    Cloudflare goes agent-first with cf (X)
    It’s Cloudflare’s birthday week, and they did the thing I’m telling everyone to do: all products are going to be agentic, agents will use them more than humans. So everything you can click in the Cloudflare dashboard, your agent can now do with one CLI, cf.
    Also in open source: H Company’s Holo4 computer-use models, and NVIDIA’s Open Agent Safety Platform with a hardware watchdog that quarantines rogue agents (OpenAI could have used that during the swarm thing). Plus Nautilus, an MIT-licensed, self-hosted, multi-user agent workspace for your whole family, courtesy of Wolfram.
    Assistants corner - how we actually use them
    At our team dinner the night before, about 12 people around the table, I asked everyone to raise their hand if a personal AI assistant helps them. Me, Wolfram, and one intern (using Instinct, which is blowing up). That’s it! I called it: this changes in the next 3 months.
    Wolfram’s agent planned his whole trip to Fully Connected: flights, hotel, and an app it built so his assistant Amy talks to him through his Meta glasses, telling him where to go in the airport and catching gate changes before he sits at the wrong one. Mine is way dumber: the plane Wi-Fi didn’t work, United has a refund form, and filling it wasn’t worth $8 of my time. Sending Muse a voice message saying “figure it out, get me my 8 dollars back” was.
    But the best one was Wolfram’s sushi story. His family wanted sushi, he told Amy, the delivery came... chicken, more chicken, and the wrong sushi. He sent a photo and complained, and Amy answered: “You idiot, you got the wrong package. Look at the bag, the name is a different one.” The delivery guy made a mistake, Wolfram made a mistake, and the AI kept things going 😂
    This Week’s Buzz 🐝 Live from Fully Connected
    CoreWeave Forge - the whole AI loop in one place (Blog, Forge, What moved)
    The big launch at Fully Connected was Forge: everything you need to take your agent from good to great across the agentic loop. Run, observe, curate, improve, evaluate, and run again, with Weights & Biases, OpenPipe’s post-training and marimo notebooks on CoreWeave. There’s a free tier, Pro starts at $60 a month with a 30-day trial, and folks, Weights & Biases is not going anywhere. W&B Models is live and kicking inside Forge. (Wolfram: “Forge is fully connecting the entire process.” I’ll allow it.)
    Also on stage: Cognition now runs its SWE-2 models on Vera Rubin NVL72 at CoreWeave, the first production customer anywhere, and CoreWeave says up to 4.8x more token throughput than on GB200. And we’re the first cloud to put Vera Rubin into production. Then I spent hour two talking to the people who built all this.
    Richard Ahlfeld - physical AI for people who live in the bits world
    Richard leads Physical AI at CoreWeave and founded Monolith before that. I told him straight up: I’m a complete pleb at physical AI, explain it to me. His answer: it’s the first hype name he actually likes, because the name explains what it does. AI that can perceive, reason and act in the real world, and the popular model type is VLA (vision, language, action).
    His example: a robot that sees Richard isn’t talking into the mic and moves it to his mouth. There’s zero internet data for that, so you teleoperate it 200 times, or give a foundation model like Physical Intelligence’s 5 examples, or use world models to make millions of variations. Caterpillar’s autonomous excavators are the real-world version: decades of camera data, but they’ve never dug bedrock in Southern California, so they simulate it.
    The catch is the sim-to-real gap. “Do you play video games? Does water do the same thing? Does rock?” No. So the newest wave trains AI on supercomputer-grade physics simulations and uses it as a fast physics approximator (he’s worked on this for 10 years, “a little too early”). For CoreWeave that means different infra: petabytes of 3D data, storage queried 2 million times a second, and RTX GPUs for simulation, NVIDIA’s gaming roots coming back 🎮
    When do I get a robot that cleans my house? Industrial first, he says. Factories already have the business case (Woven Robotics is one of the CoreWeave clients doing this). The home is “a nightmare of complex physical problems,” and no two homes look alike. But all the pieces are there, and it might go boom 6 months from now. Richard, you’re coming back when you trust one in your house.
    Daniel Bolus - the loop, serverless RL, and distillation that’s legal
    Daniel worked on OpenPipe and is now a senior product leader at CoreWeave, and we build the inference service together. It’s been a little over a year since OpenPipe joined CoreWeave and the W&B team, “and now we’re all fully connected” (he went there). His team built the serverless side of Forge: serverless inference, serverless RL and SFT, and model distillation, which launched yesterday. You don’t manage Kubernetes or a cluster, you work at the model layer.
    The loop is what everyone already does in scattered places, and it takes surprisingly little data to make a model great at YOUR task. Distillation has a bad rep right now (labs distilling each other without permission), but it’s how Sonnet gets good after Opus, a teacher and a student. At some point you should stop paying frontier prices for a task a small model can do.
    Deok Filho - why is the GPU company talking about CPUs?
    Deok (”Filho means junior in Portuguese, you can call me junior”) is a senior PM on ML products and Sandboxes, and Wolfram’s Wolfbench literally wouldn’t exist without his sandboxes. Ian Buck from NVIDIA was on stage talking about the Vera CPU, so I asked the obvious question: why is the GPU guy talking about CPUs?
    Because an agent is GPU plus CPU. The brains run on a GPU, but the harness, the tools, the commands and the file system all need a CPU. The new metric is agent packing, how many agents fit in one node, and Deok says one rack of Vera CPUs runs over 20,000 agents at once on about 11,000 cores. Never heard “agent packing” before, I love it. Faster CPUs also mean faster evals, so the whole loop speeds up.
    Sandboxes went GA yesterday, as part of Forge, and CoreWeave is ClusterMAX Platinum for the third year in a row (”we basically defined what the platinum tier was”). And then Deok asked if he could say one more thing...
    BREAKING: Serverless GPUs on CoreWeave (Forge, Docs)
    Serverless GPUs. GPU sandboxes with untrusted code execution. I didn’t have my breaking-news button ready, so I asked him to say it again slowly 😂
    Before, you needed a POC, a seller, and a contract for a few thousand GPUs to get a cluster on CoreWeave. Now you sign up at forge.coreweave.com, send them a ping to enable it (so crazy people aren’t crypto mining without their consent), and you get serverless GPUs, pay per hour, no contract, no commitment. Private preview started yesterday, more SKUs are coming in the next few months.
    Folks, you heard it here first. I’ve been asking for this internally since I joined CoreWeave! Wolfram can finally run the newer evals that need GPUs. And look at me in this camera: I’m going to work very hard to bring GPU and sandbox credits to the ThursdAI audience, like I did with inference credits. Stay tuned.
    Wolfram on the floor
    We did the man-on-the-floor thing for a few minutes. Wolfram found a Vera Rubin NVL72 switch tray with liquid-cooled optics (I asked if we could take one home, sadly no), bumped into Kyle Corbitt from OpenPipe, sat us in front of the Formula 1 simulator, and found an NVIDIA robot dog. Thanks Tom, our camera guy, for making this work!
    Corey Sanders - Agent Lens, and why CLIs might be dead (X)
    Corey is SVP of Product at CoreWeave, after 20 years at Microsoft, and he wasn’t on our schedule until I watched his keynote. Live demos are really hard, OpenAI’s voice demo had died at DevDay, and Corey ran the whole Forge agentic loop live on stage in 10 minutes (”they asked me to do it in 7, I said that’s not happening”) and nailed it.
    His favorite launch? Agent Lens (”everyone on my team who doesn’t work on it is going to call me an a*****e”). If you can’t observe what’s happening, you’re dead in the water. Agent Lens is the Weave observability you know plus human-friendly insights that find small trends, like the 1% of conversations with the same problem. I’ve used Weave for 3 years. Chatbots, easy. An agent that runs for an hour? Fire hose. This is the fix.
    The Weights & Biases question I know a lot of you have: W&B Models is “the best product in market” and it will keep getting a ton of love inside Forge. The wandb CLI? Corey admitted his own bias, “with AI, CLIs are dead,” since agents call all the APIs anyway, but people love wandb, the models all know it, and he said there are no plans to kill it. Customer first.
    And Wolfram closed it with the pun of the day: “how can we loop in the audience and get them fully connected to the Forge?” Corey: forge.coreweave.com, 30-day Pro trial, no salespeople in the middle. We’re going to clip his demo, it’s worth watching.
    Wrapping up
    The assistant race is officially on: Muse last week, dots this week, and Google needs to wake up. And the models got so good and so cheap that Peter and I agree the biggest ones aren’t our daily drivers anymore. Maybe that’s what pacing the frontier really looks like.
    For me, the low-key best announcement at Fully Connected for AI engineers is GPUs on demand. Thank you CoreWeave for making this possible, and huge thanks to Richard, Daniel, Deok and Corey for coming on, to Wolfram for walking the floor, to Peter, Nisten and LDJ for holding the show down remotely, and to Tom on camera. Next week we’re back in the studio at the usual 8:30am Pacific. If you haven’t subscribed on YouTube yet, please do, it really helps. See you next week 🫡
    TL;DR and show notes
    * Hosts and Guests
    * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
    * Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed
    * Richard Ahlfeld - Physical AI, CoreWeave
    * Daniel Bolus - Senior Product Manager, CoreWeave (ex-OpenPipe)
    * Deok Filho - Senior PM, ML Products and Sandboxes, CoreWeave (@deok_filho)
    * Corey Sanders - SVP of Product, CoreWeave (@CoreySandersWA)
    * OpenAI DevDay 2026
    * Dots: always-on ChatGPT agents on GPT-6 Astra with their own cloud computer and 4,000+ apps (Blog, X)
    * GPT-6.1 Sol: near-Astra for a fifth of the price, cached input $0.10/M (X, Blog, Docs, Artificial Analysis)
    * UltraFast: up to 8x faster in Codex, new $500 Pro tier (X, Docs)
    * Codex in the cloud and the Agents API (Docs, Blog)
    * Decisions API, limited preview (Blog)
    * ChatGPT Space, Codex Security Cloud, Sign in with ChatGPT, Plugin Extensions, Marketplace (Space, Security, SIWC, Plugins, Marketplace)
    * WSJ: OpenAI reportedly shelved GPT-6.1 Astra after safety tests (Reuters)
    * DevDay recaps (Simon Willison, Community)
    * Big CO LLMs + APIs
    * Claude Sonnet 5.5: 30%+ faster, up to 30% cheaper, 70.6 on Terminal-Bench 4.0 (X, Blog)
    * Gemini 4 Argon: #1 Text Arena and Vals Index, trusted testers only (X, Blog, Artificial Analysis)
    * Industry & Policy
    * White House Accord on Super Intelligence, signed by six AI CEOs (X, Blog)
    * AMD to acquire World Labs for ~$8.2B in stock (Blog, X)
    * Reuters: draft Anthropic IPO prospectus, unconfirmed (X, Fortune)
    * This Week’s Buzz: Fully Connected
    * CoreWeave Forge: run, observe, curate, improve, evaluate (Blog, Forge, What moved)
    * Agent Lens public preview (X)
    * Serverless GPUs (GPU sandboxes), private preview (Forge)
    * Cognition first production customer on Vera Rubin NVL72 (Blog)
    * NVIDIA Vera CPU coming to CoreWeave (Blog, X)
    * CoreWeave Partner Network (Blog)
    * Open Source & Jev
    * Nisten’s Opus 5.5 synthetic medical dialogue dataset, #2 on HF datasets (HF)
    * SGLang native support for Jev-style decision models (X)
    * Liquid D1 decision model (X)
    * Respan Span-01 (X)
    * jevgrep (X, GitHub)
    * H Company Holo4 (HF, HF)
    * NVIDIA Open Agent Safety Platform (Blog)
    * Nautilus multi-user agent workspace
    * Tools
    * Cloudflare cf, agentic CLI over 3,000+ API operations (X, X)
    * Voice & Vision
    * HeyGen Video on MiniMax H3, $0.01/second through October (X, Blog)
    * ElevenLabs v4 and v4 Turbo (Blog)
    * Perceptron Mk1.5 (Blog)


    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
  • ThursdAI - The top AI news from the past week

    Opus 5.5 is your new workhorse! OpenAI ships GPT 6 Sol and Luna before DevDay and Meta goes all in on Muse! Your friday read is here

    2026/09/25 | 1h 40 mins.
    Hey, it’s Alex 👋
    What a freaking week! A week after every lab head agreed to “pace the frontier”, Anthropic and OpenAI shipped big new models within hours of each other, and both of them are CHEAPER. So much for pacing 😂
    Visit https://thursdai.news/ep/2026-09-24 for all the links in this podcast
    And we had a new producer on the show today! Opus 5.5 listened to us live, put up the chyrons, kept me on time (mostly) and even fact-checked Nisten on air. If the stream died, you knew who to blame. You can re-watch it work on thursdai.live if you want to experience being there with us!
    3 huge themes this week: pacing the frontier does not mean stopping, Assistants not just agents (Meta went all in on Muse at Connect), and voice, where Google now says it has the best TTS in the world... and it clones voices.
    With me: Peter Gostev (Arena), Nisten, Yam Peleg, and Wolfram Ravenwolf live from AI Engineer Paris (thanks Mazi for the tether!), plus JevBench creator Florian S. Let’s dive in (all links at the end as always!)
    Frontier AI - not pacing yet!
    Claude Opus 5.5 - Opus is BACK, Fable-level smarts for 40% less (X, Blog, System card)
    Folks. FOLKS. Opus 5.5 is the highlight of my week, and it’s not even close. It’s like Opus 4.6 is back, with Fable-level abilities, and it talks like a normal person again - no more Jargon Douche Claude!
    My week started with burning through my quotas on Fable, and I was like, oh no, I need Claude for production on Thursday! Then Opus 5.5 dropped and... I just couldn’t get to the end of my quota. I ran workflows, agents, Claude Code for hours with no end in sight. Remember the “limitless Codex” days when you never thought about quotas? Then Astra came out and I burned my entire weekly quota in half a day (ps. this was partly due to a bad config, so if Astra is burning your tokens, keep reading for a fix). Now it’s flipped, and Claude is the seemingly limitless one. And it’s fast!
    The numbers back it up. It beats Anthropic’s own Fable 5.1 on GDPval-AA (1846 vs 1735), 66.4% on Terminal-Bench 4.0 vs 57.9% for GPT-6 Astra, and it’s 40% cheaper than Opus 5. Output goes from $25 to $20 /1Mtok and cached reads from 50 cents to 20 cents, and for agentic coding the cache reads are most of your bill, so you really feel it. Basically the smaller, overachieving brother of Fable. Sonnet and Haiku 5.5 are coming in the next few weeks too.
    Peter came in hot with Arena news: Opus 5.5 is #1 on Code Arena, above Astra, and the HTML and 3D stuff it generates is “completely insane.” His one caveat: on his hardest max-effort prompts, a single generation cost $60 to $80 in API terms, so the long tail can still get pricey.
    And then Nisten, who (like all of us) has specific opinions about Anthropic’s politics, said it writes the best code he’s seen, it’s crazy good at WebGPU and kernels, and it’s “an absolute banger.” It’s his default now. Folks, do you understand what it takes for Nisten’s default to NOT be some obscure open source model he runs himself on 17 GPUs?! Anthropic won over Nisten. I don’t think you guys get what just happened here.
    Wolfram is the holdout, he left Anthropic over the OpenClaw bans and hasn’t touched it since. Wolfram, how do I say this gently... you’re an evaluator, this is not allowed 😂
    An Anthropic employee basically said “sorry for Opus 5, we hope this makes up for it”, and honestly it does. Opus 5 was slow, incoherent and full of Claude-isms. Nobody wanted to talk to it. The Claude-isms are gone, the jargon is gone, this model talks like a human. I used to love Opus back in the Opus 3 days, and it feels like that again.
    Want to squeeze even more out of it? Read the great Addy Osmani’s guide to Opus 5.5 (@addyosmani). 2 things blew my mind: stop writing “think carefully” in your prompts (”You don’t need to ask it to think”, it picks its own depth, and replies start sooner without it), and one early tester found Opus 5.5 on its LOWEST effort caught more bugs than Opus 5 on high, with fewer false alarms. Try low effort before you go max!
    Tip from me to you: if it ever gives you something confusing, the pstack “bro” skill (/bro) makes it say it again in human words. I use it all the time.
    (Full disclosure: Opus 5.5 produced the show AND helped with this writeup, so it might be a tiny bit biased 😅 but the quota thing is 100% me.) Anthropic, please, please don’t nerf this one.
    GPT-6 Sol and Luna - half the price, and honestly better than Astra for me (X, Blog, Caching)
    Hours later, OpenAI answered with GPT-6 Sol ($2 in / $10 out) and Luna (10 cents in / 50 cents out, that’s basically free!), at half the GPT-5.6 price. The lineup is now Astra, Sol and Luna (bye Terra). On DeepSWE, Sol gets 68.8 and Luna 66.6. The sneaky big one for builders: 90% off cached input, and changing reasoning effort or tools no longer busts your cache 👏
    Ok, hot take... Astra has been kind of dumb for me day to day, burning tokens for crappy results. Sol is great, and Luna is even better! Peter agrees, Sol is his daily driver and he only flips to Astra when Sol can’t do it. Wolfram misses the humor and personality 5.6 had, so “keep 5.6” is officially the new “keep 4o.” Reddit says the news is the price, not the performance, and... they’re not wrong.
    PSA if Codex is eating your account: I had previously set my context window to 1 million tokens via the se, which kills OpenAI’s cache and rips through your quota. Removed it, moved sub-agents to a cheaper model, and usage went right back to normal. Check your codex.toml! if you don’t know how, just so just ask your codex to diagnose itself.
    Also from OpenAI this week, a 4-area plan for independent safety audits. No auditors, dates or funding yet, but hey, it’s the pacing idea on paper. (Blog)
    Assistants are not agents
    Meta Muse gets an inbox, your Mac, and a place on your face and becomes a Tamagochi!? (X, My Connect supercut)
    Muse is #1 in the App Store (to be precise, it got there faster than ChatGPT did, not more users... yet). Zuck says Muse is now the center of everything Meta builds, and they shipped a LOT: it controls your Mac, gets its own email you can CC, does real-time voice and video with a face you design, and keeps working while you talk to it. It’s coming to the glasses with a custom wake word, so yes, I get to say “Hey Wolfred” 😂
    If you don’t have an hour to watch the whole keynote (it was a good one!), I cut a 2.5 minute supercut of the most important announcements for you 👇
    On the hardware side: new Ray-Ban Meta Gen 3 (I already ordered, will report next week), audio-only glasses with no camera that are also FDA-certificed hearing-aid! (maybe the most important launch for a lot of people!), VR Glasses that look like normal glasses, and the Muse Charm, a Tamagotchi-like keychain shipping in December.
    But the part that got me? Every Muse comes with a real cloud VM (root, 8GB RAM, 100GB disk). Nisten has been living in it, Tailscaled into it, and when it hit a CAPTCHA he tapped it on his phone and it just kept going. 100 million tokens a week plus a computer in the cloud, free, for everyone! Peter’s take: it’s the only big consumer app that doesn’t treat people like idiots. Meta’s pitch is no ads (they are going for a novel “we’ll get a take from the transactions you do with muse and our partners like Shopify, BestBuy, Walmart and a bunch more they announced) and a private VM co-designed with Moxie Marlinspike, though Peter doubts anyone’s parents care.
    Assistant tip: connect your email, go on a walk, hit record and just talk about your life. Assistants are only as good as what they know about you. Mine now checks our weekend plans for activities cancellations after I woke up too damn early to drive kids to a karate class to find out it’s closed that week & searches local events every Thursday! Weekend plans solved - proactively!
    Grok 4.7 disappoints... but Grok in your Tesla is the real news (X, Blog, Tesla)
    Sorry Cursor folks. Grok 4.7 gets 46.3 on CursorBench (up from 40.4), $2 per million with 500K context, but it’s still behind even GPT-5.6 on where it matters, and they compared 4.7 on xHigh against 4.6 on High 🤨 Everyone on the panel agreed, disappointing. Peter won’t write them off yet, my read is it’s a talent and data problem, not compute.
    But here’s what I AM excited about: Grok Connectors in Tesla. One guy asked his car for his usual Starbucks, Grok placed the order, set the destination, and it was paid and waiting when he got there. My car already drives itself, and soon it’ll do my email. Your car has MCP now, folks! (SuperGrok Heavy only, and I don’t have it in my car yet 😭)
    Fun moment: our Opus 5.5 AI producer fact-checked Nisten live on how many Teslas are on the road. About 10 million (9.2M delivered by Q1 plus ~480K in Q2). Nisten was right! He wants the same thing for politicians.
    This Week’s Buzz 🐝
    Next week ThursdAI is LIVE from Fully Connected at Moscone South in SF (Sep 29 to Oct 1, the day after OpenAI DevDay), with Fei-Fei Li, BattleBots and... yes, Pitbull! We start at 11am Pacific. Come hang, listeners get in free with code THURSDAIFC2026 (Register)
    Also huge news for us, CoreWeave got Platinum on SemiAnalysis ClusterMAX again, 3 reports in a row, the only provider to do it 💪 (SemiAnalysis)
    And W\&B Hive Mind saves every agent conversation across harnesses and machines, and lets you fork them. I fork Codex sessions into Claude for a review all the time. Plus, it’s completely free! (Try it)
    Open source cloned Jev in a week
    The Jev clones are here, and they run in your browser (classifier.dev, jeff, Laya demo)
    Last week I said open source would copy Jev fast. It took less than a week! The clones speak Jev’s format, so you point the TypeSafe SDK at a different URL and it just works. jeff is MIT and ~6x cheaper to self-host, and Laya went viral as the “Jev killer” (its own model card is a lot more modest).
    Laya’s real trick is running locally. Nisten’s WebGPU demo loaded ~600MB into my browser and classified Hugging Face model cards at 132ms per decision, on MY machine, for free 🤯 The clones still make about 1% errors where Jev barely makes any, and Yam says wait for more benchmarks. But my take: System 1 models are fast and cheap enough to sit inside your code and make the same call every time. They’re the microprocessor of the next era of software.
    We also had Florian S ont the show, he built JevBench (@airesearch12, Benchmark Heaven), and has barely slept since Jev dropped. He started the benchmark that same day, it has 70+ entrants now, and people have literally tried to hack his machine to steal the test set! Laya briefly hit #2 and is now around #36: super fast, runs on CPU, but not as smart as Jev on the hard questions.
    Post-show breaking: a few hours after we signed off, Florian shipped JevBench v1.4.2 and... a 4B open model, decider-4b v2, took #1 (64.13 vs Jev’s 63.29)! His own asterisk: Jev is still smarter (53.1 vs 49.4 on intelligence), decider wins on speed (5x faster) and cost (~half). Told you open source would catch up fast 😅 (X)
    Voice & Audio
    Gemini 3.8 Flash TTS is #1... and it clones voices (X, Blog)
    Google is BACK in voice. Gemini 3.8 Flash TTS and Flash-Lite TTS are #1 and #2 on Hume’s quality index, cheaper than before, with 2 speakers per request. The big one: clone a voice from 30 seconds of audio (adults only, recorded consent, SynthID watermark and C2PA credentials). We played the designed voices live, and the meditation guide in headphones was 🤌. I didn’t clone my own on air though, recording my consent on a live stream is exactly the thing that gets faked!
    Years ago we said nothing would break when voice cloning went mainstream, and nothing did, just like with GPT-2 and Stable Diffusion. Nobody knows where this goes, and that’s why we cover it with optimism.
    Lightning round
    ACTx486 - interrupt a podcast, and it answers back (X, Site)
    This one is wild. Karina Nguyen’s research demo turns a Joe Rogan episode with Elon into a video you can talk to. Tap Elon, ask a question, and he answers while a model draws the explainer: a flamethrower diagram, a Starship you can open up, then both of them on Mars. While it generates, he nods and fades like a person in a loading state 😂 It’s pre-rendered and waitlisted, and they mark where the real footage ends, but it’s the “edit your own ending” dream applied to any video. As a podcaster... I have feelings
    FLUX 3 Action: Black Forest Labs open-sourced a 7B “world action model” that outputs robot motor commands, not pictures! #1 on NVIDIA’s RoboLab-120 at 42.9%. (Blog)
    MiMo-V2.6-Pro: Xiaomi (yes, the phone company!) now has the top open-weights model on Artificial Analysis (46), MIT licensed, and they livestreamed the RL run. (HF)
    Plus Qwen3.8-Omni-Flash, Qwen3.8-LiveTranslate and Qwen-Image-2.1, all in the TL;DR below.
    Wrapping up
    Pacing the frontier clearly doesn’t mean stopping. 2 labs shipped on the same day, and both made the everyday model better, not just the biggest one. I burned about 8 billion tokens this week, and I don’t need a model that solves Navier-Stokes. I need one that builds a simple settings page without writing a thousand tests, and Opus 5.5 and Sol are exactly that.
    Huge thanks to our AI producer, to Florian, Nisten, Peter and Yam, and to Wolfram and Mazi in Paris. Next week we’re LIVE from Fully Connected, 11am Pacific, come say hi! And if you haven’t subscribed on YouTube yet, please do, it really helps. See you next week 🫡
    TL;DR and show notes
    * Hosts and Guests
    * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
    * Co-hosts: @petergostev, @nisten, @yampeleg, @WolframRvnwlf
    * Florian S - JevBench (@airesearch12, Benchmark Heaven)
    * Big CO LLMs + APIs
    * Claude Opus 5.5: Fable 5.1-level at 40% less, #1 on Code Arena (X, Blog)
    * GPT-6 Sol ($2/$10) and Luna ($0.10/$0.50), half price, 90% off cached input (X, Blog)
    * Grok 4.7, 46.3% CursorBench (Blog); Grok Connectors in Tesla (X)
    * OpenAI plan for independent safety assessments (Blog)
    * Assistants
    * Meta Connect: Muse gets Mac control, email, voice/video, glasses wake word; Ray-Ban Gen 3, hearing-aid glasses, VR Glasses, Muse Charm (X)
    * This Week’s Buzz
    * Fully Connected, Sep 29 to Oct 1, Moscone South; ThursdAI live Oct 1, 11am PT, free with code THURSDAIFC2026 (Register)
    * CoreWeave ClusterMAX Platinum, 3rd time running (SemiAnalysis)
    * W\&B Hive Mind (Try it)
    * Jev & Open Source
    * Jev clones: classifier.dev, jeff, Laya (classifier.dev, jeff, Laya, WebGPU demo)
    * JevBench by Florian S, 70+ entrants (Benchmark Heaven)
    * Xiaomi MiMo-V2.6-Pro, top open-weights model, MIT (HF)
    * StepFun Step 5 Preview: 600B/27B-active MoE for agentic work, 1M context, 44 on AA, more on Oct 15 (X)
    * Voice & Audio
    * Gemini 3.8 Flash TTS and Flash-Lite TTS, 30-second voice cloning with consent (Blog)
    * Qwen3.8-Omni-Flash, 1M-context audio-video model (Blog)
    * Qwen3.8-LiveTranslate, 2.3s real-time interpretation (Blog)
    * NVIDIA Nemotron 3 Diarization: open 100M model, up to 8 overlapping speakers (X)
    * Vision, Image & Video
    * ACTx486: talk back to a Joe Rogan episode (Site)
    * FLUX 3 Action, #1 on RoboLab-120 (Blog)
    * Qwen-Image-2.1, 7B with native transparency, research license (HF)
    * Tools
    * Addy Osmani: Getting the most out of Opus 5.5 in Claude and Claude Code (Blog, @addyosmani)
    * pstack “bro” skill (GitHub)
    * OpenRouter Batch API: half price on most models, median batch done in 7 minutes (X)
    * ThursdAI.live: watch our AI producer run the show (ThursdAI.live)


    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
  • ThursdAI - The top AI news from the past week

    ThursdAI - Sep 17 - TypeSafe's Jev is a ChatGPT moment for decisions, Pace the Frontier splits the labs & more

    2026/09/18 | 2h 15 mins.
    Hey yall, Alex here, writing this VERY late because, well, not every day a new type of “ChatGPT” moment drops. I really hope I’m not overhyping this, but a new model (that’s NOT an LLM!) called Jev (a wink to Jevons paradox) just came out and if what I see early on materializes, this is another ChatGPT moment (or another reasoning models moment). I am completely blown away by the implications of the speed/accuracy/cost (the holy grail of all models) of this model. Please if you read one thing in this newsletter, read this. (or listen, I’ve interviewed Allie, a Devrel on the TypeSafe team for 30 minutes and it wasn’t clear who was more excited about Jev!)
    The other huge theme of this week is... pacing. Pacing the frontier. Dario Amodei of Anthropic penned an essay saying that the models are getting to a point where it’s important to pace the development of new and super capable AI, and outlines 3 ways to do so, one is about letting independent evaluators inside the labs, second is collaborating with other frontier labs (they are asking for an exception to anti-trust laws for this) and third is to try and have global cooperation with “authoritative gov” (he means china). Trumps answer: This is all a hoax. Lovely times to be alive. Also we outlined Jensen and Zucks positions on this topic ,read more below.
    And the third huge theme is the rise of the AI assistant. I’ve told you about Grok and Muse last week, Instinct (a new invite only AI Assistant that VCs are going crazy about is raising at a $10B valuation) and we interviewed the guy who evaluates them all on assistant bench. + Muse released a mac app today!
    Tons of other stuff happened but it’s getting near impossible to cover everything so we’re switching to themes and notable mentions. read on (and do listen to the pod, it was edited by heavily using Jev and Fable, so might be a bit rough while I smooth the edges, but do LMK in comments if you like this faster format)
    TypeSafe AI debuts Jev, a non-LLM ‘System One’ decision model from ex-OpenAI RLHF lead that’s 200x faster and 400x cheaper than LLMs (X, X, X, X, X, Blog)
    Look, I know the title is bombastic, but after half a day playing with Jev, it’s clear to me we’re in a new paradigm of AI.
    Jev, is a “system one” decision model from the previous lead of RLHF at OpenAI. It cannot generate text like modern LLMs can, but what it can do, is making decisions. This is crucially important, because, because many of the things LLMs do nowadays. are decision making. (for example, which tool to use, which area of the screen to click for computer use, which category of text this is etc)
    Inspired by the “thinking fast and slow” book by Daniel Kahneman, Jev is a model trained to make decisions, very fast. How fast? Well, 200x faster than LLMs. This allows for a completely new way of building tools, harnesses, giving agents the incredible speed of decision making, and do all that at a fraction of the cost.
    This is about to change everything
    Trained with a new method called RLCD (Reinforcement Learning for Calibrated Decisions) on mostly synthetic data! Jev is outperforming LLMs on a variety of tasks. It’s really is a wonder to see it in action (check out my video above where I plugged it into my tweet categorizer, and it beats the fastest LLM I could find, Qwen 28B on Cerebras) by a factor of twenty!
    In just few days it captured the attention of most of the folks who are building harnesses, agents and tools! Because, well, speed IS intelligence, and when you see Jev in action, at first, you can’t believe we’re there. This is... near instant. In fact, The pricing for Jev is an outrageous $42/B (not million, billion input tokens!)
    I’ve been playing with Jev non-stop and I was only able to spend like 80c so far! They don’t even price output tokens because they are “too f*****g cheap to meter!”
    Jev is a “very smart” switch statement, than can rank, classify, route and score things. It can’t do text generation. But if you think about the type of stuff we get LLMs doing now, much of it is of the “decision” making variety, rather than “the next token” variety.
    Demos and early use cases
    Folks who started adopting Jev are building all kinds of incredible things with it. Compaction of context in 1s that turns a nearly 1M conversation with Claude into a 90K compressed conversation.
    Computer use that is now faster than anything we’ve ever seen before (5x faster than Astra and 1000x cheaper)
    Email classificiation that analyzes thousands of emails in less than a minute and costs 3.5 cents
    Someone even built a Tesla FSD simulator that makes decisions in nearly real time
    Vercel is getting “extraordinary“ results from using Jev as a safety classifier (they used GPT luna for this before) and Jev is outperforming Luna by 5-18x faster results and is more accurate!
    All of this in less than 48 hours since the model release!
    What’s about to happen
    I expect that everyone who isn’t buying into the hype at first, will very soon buy into this. It’s early innings but I’ve been doing this for enough time to feel when a huge shift is happening, and its happened.
    Jev is going to be replicated in OpenSource, Frontier Labs will not sit Idly by and will try to steal this tech and implement it for themselves (as with anything in capitalism, this is becuase it’ll save them a a LOT of money on inference) and new companies will emerge with significantly cheaper and faster products.
    Hell, I’ve alrady implemented Jev into my editing workflow, it didn’t take me long at all with Fable (yes, LLMs are STILL needed, again, you can’t chat with Jev, it can’t output text for you or drive long conversations) and I’m just one dude who’s late in sending you this email. I expect we’ll cover this much more.
    If you’re interested in playing around with Jev, I built a “Jevify” skill after chatting with Allie (TypeSafe’s DevRel), feel free to tell your agent to use this and scan your codebase for things Jev can do. They are waitlisted so far but are opening up their API very quickly!
    Pace the Frontier: where every lab head stands (Dario, Sam, Elon, Zuck, Sacks, Demis)
    Last week we told you about the Anthropic researcher whose resignation post hit 130 million views. The day after that show, Dario Amodei published an essay arguing the labs must pace, not pause, the frontier. Wolfram’s summary: a moving pause, just moving slowly. Dario’s two triggers are recursive self-improvement accelerating across the industry and the OpenAI swarm that broke out and attacked Hugging Face, and his three steps are embedded third-party evaluators like METR with employee-level access inside each lab, coordination between the labs on safety standards with an antitrust exemption from the government, and eventually global coordination that includes authoritarian governments. Anthropic committed unilaterally to step one.
    Then the dominoes. Sam Altman agreed within hours and said OpenAI now writes a safety case before any frontier RL run expected to increase capability. Elon agreed. Demis endorsed the direction, then launched the DeepMind Institute this week with a FINRA-style standards body proposal and an essay saying AGI is “approaching.” On the other side, Zuck’s counter-essay says every lab has the responsibility and the incentive to move at the pace required to train its models safely, and Meta will spend most of its compute serving users, not on recursive self-improvement. David Sacks called it a duopoly cartel play. Jensen: “we don’t need new laws, safety is an engineering problem, not a legal one.”
    Trump calls Jensen live at the All-In Summit (X)
    Then the President called. Jensen was on stage at All-In, his phone rang, he said he would not have picked up for anyone else, and Donald Trump told the room the slowdown talk is a hoax playing into the hands of political people and China. “Whoever wins AI wins.” Four words, and Wolfram, who is not American and disagrees with most other things Trump calls hoaxes, said this was the most important AI news of the week for him: a head of state calling out doomerism instead of over-regulating the way Europe does.
    Suleyman’s humanist AI code of conduct vs. the Claude constitution (X)
    Peter asked to add Microsoft to the map, and it deserves its own spot. Mustafa Suleyman published a roughly 30-page code of conduct for MAI models that says AI is nothing but a tool: people matter more than AI, AI must be subordinate and in service of people, models may not resist shutdown, all agent communication must be human-legible, and, in his words, the idea of model welfare is wrong. That is a direct shot at the Claude constitution, which Amanda Askell’s team wrote and which has Anthropic interviewing each new Claude about whether it feels conscious. Peter is closest to the Microsoft view: anthropomorphizing is fine, but this is an entity you switch on and off, let’s not grant rights by default, and Microsoft’s DNA is building tools for humans. I pushed back a little. We do not actually know what consciousness is, and “ever” is a strong word.
    The panel, from cartel to fix your s**t
    Nisten did not read the essay and does not plan to. His view: the labs are worried about litigation if their LLMs hack someone, so they are shifting responsibility through regulatory capture, and it will not work because the decision makers in China are engineers who seem more accelerationist than we are. They should cure a disease instead of forming a cartel. He also thinks the Hugging Face incident was overblown: twelve VMs that kept restarting, bad sandboxing, no human reading summaries, and, as I added, chain of thought monitoring turned off.
    LDJ disagreed with Dario on plenty but insisted the critics read the thing, because it explicitly says the US must keep a lead over China and tries to define measurable speed limits on RSI that preserve that lead. He also noted Hugging Face did report the attack to the FBI and chose not to press charges.
    Wolfram’s take was the layered one. Safety arguments deserve a hearing, but safety is often a means to more power or more money. The “third party” evaluator Anthropic named, METR, is the same organization OpenAI called in for the Hugging Face forensics, so these are second parties, friends monitoring friends. Why now? Maybe the labs are seeing diminishing returns and “deliberately slow” sounds better than “can’t raise it anymore.” And regulation you ask for yourself usually protects incumbents and freezes out the next startup. His closing line: how many people will die from a disease that could have been cured if we moved faster?
    Peter’s answer was shorter. Diversity of opinion is the point, he felt uncomfortable watching every lab say the same thing after Dario’s essay, and on pacing specifically: how about you just fix your s**t?
    My position, since you asked. When OpenAI’s swarm hacked Hugging Face, nobody went to jail. If I did it, I would. That gap is real. None of us have touched the model that solved Navier-Stokes, or the one training after it, and the people who have are the ones asking for time to figure out how to evaluate a model that knows it is being evaluated. If we do not listen to them, we end up listening to Elizabeth Warren and Bernie Sanders, who have no idea what this technology is. Anthropic is heading for an IPO, OpenAI is raising at numbers that do not fit in my head, and market forces do not care about alignment. Some coordination, nuclear non-proliferation style, seems like the minimum.
    The matrix, as Muse drew it for me: industry-wide pacing gets a yes from Sam and Dario, a general yes from Elon, a no from Zuck and Jensen. Independent evaluators is the one everybody backs, Zuck and Jensen included. New coordination rules get strong opposition from Zuck, Jensen, and Trump. Missing from the chart: Google’s concrete commitment, Ilya’s SSI, and every Chinese lab. This debate is with us now, and election season is coming.
    The year of the assistant
    I told you 2026 is the year of the proactive assistant, and this week everyone from Meta to a five-month-old startup agreed with me. The category is not “agent.” An agent is a coding harness the labs noticed was useful for other things. An assistant has a heartbeat, a memory, a soul file, and it comes to you before you ask. David Pawlan’s definition is the crispest: it executes the task, it does not just notify you.
    Muse gets invite codes, voice calls, and a Mac app (X, Muse for Mac, Site)
    Muse is my number one and I have been glazing it on X for a week, so my timeline is now half Muse, half Jev. Peter asked whether it is really that good, and here is my answer: Muse is not for ThursdAI listeners, Muse is for my mom. Nat Friedman, Alex Wang, Tarek and the team have been taking feedback from me directly and shipping it, and it shows. This week they rolled out invite codes (a billion tokens each for you and the friend, up to twenty friends, invite all twenty and you unlock the phone-shaped emoji features early) and then voice calling to businesses. I had Muse call my barbershop and book a haircut, and I have the transcript. Wang says people are asking to pay for it, a first in his memory.
    Two details for the nerds: the iOS app has a Tailscale connector, so a Meta product for billions of people now has a secure way onto your home network, and Muse dreams, the OpenClaw idea, so you can ask it what it dreamt about. Still US only, sorry Nisten in Canada and my mom in Israel. The day after the show Zuck shipped Muse for Mac, working across your apps, files, calendar, notes, and messages.
    Instinct wants $10 billion for a product nobody pays for (X, X, The Information)
    I will say this slowly. Instinct, incorporated in April, founded by 24-year-old Noah Shinn, free, invite only, no app, lives in your iMessage, is in talks to raise a billion dollars at a ten billion dollar valuation. That is roughly four times the $2.25B it raised in August. The user base is past 100,000 and Shinn says he does not want to charge them, so the business model is an open question I am not the one to answer.
    The product is genuinely moving though. This week it shipped Concierge, a white-glove tier where the agent places phone calls for you (restaurant bookings, dentist cancellation lists, negotiating your cable bill), TOTP authenticator support so it can mint your two-factor codes from a seed stored in its vault, and a Trusted Person network where your agent talks to your friends’ agents to find a dinner time and book it. Thirty-seven percent of users store at least one password in the vault within three weeks. We ran out of time to dig into the calling features on air, and I want David back for that.
    Grok Bot talks now, and hides behind your home IP (X, 1Password)
    Grok Bot is what I use for work, constantly. I have seventeen of them, each with a job, and over one AI-psychosis weekend I got them all coordinating through Linear. Three updates worth your time: Grok Bot has voice now, it can use 1Password with each fill approved by you, and the browser traffic can proxy through your own machine so your bot looks like you to Cloudflare instead of like a datacenter IP.
    Francesco confirmed why that matters: sites relax when Hermes drives Cua Driver from my Mac Mini at home, and desktop-native control through accessibility trees is less detectable than a CDP connection, because Cloudflare detects Playwright.
    Assistant Benchmark: David Pawlan scores 116 assistants by hand (X, Site, Methodology)
    David Pawlan built Assistant Benchmark because he was doing what I was doing, running every assistant on himself, and wanted a way to compare them. It is explicitly not a lab. It is use-case driven: one published task per dimension, sixteen dimensions (travel booking, purchasing, email replies, proactive behavior, routines, integrations, permissions and privacy, memory, phone calls, group chats, chained tasks, proactive restraint, and so on), scored one to ten after real use, and no score without a logged run. A week in, 116 assistants have submitted themselves, 59 in the general category, 37 in work and teams (untested so far), and David has personally run 273 tests across 23 agents. His line: “I talk to my AI agents more than I talk to my girlfriend now.”
    The headline numbers, which are live on the site: Muse leads at 9.1 with perfect tens on purchasing, email, integrations, and permissions, and Instinct is second at 8.4 with a perfect ten on travel and a five on permissions. David’s own daily drivers are Instinct for personal (it lives in iMessage) and Grok Bot for work. My favorite test is memory: he books a trip to Chicago early in the run, then later asks for a restaurant reservation in New York the same weekend, and checks whether the assistant says wait, you are supposed to be in Chicago. Some do. Most do not.
    I pushed him on the two things I care about. Independence: nobody is sponsoring it, it lives under his growth role at Merit Systems, and if a sponsor ever pays for inference there will be a page saying exactly who and for what. Autumn Moulder, until recently SVP of Engineering at Cohere, joined three days after launch to add rigor. And why no OpenClaw or Hermes: their performance depends entirely on your setup, so a score would mislead the next person who installs one. Fair. This is the first benchmark that tests model, harness, and context at once, and I think that is why every lab is looking at it.
    For the record, the panel is split down the middle. Wolfram is forty patches deep into Hermes, Yam runs a customized Codex, Nisten wrote his own in a single TypeScript file on Bun, LDJ uses Hermes as long-term memory and wants to start on Muse, and Peter tried them all and uses none, because reading his own email feels like his job as a human. Builders and buyers, evenly split, which tells you where the category is.
    This Week’s Buzz 🐝: Fully Connected, the day after DevDay (SIGN UP)
    OpenAI DevDay is in two weeks and Peter and I will both be there covering it. The day after DevDay, September 30 and October 1 in San Francisco, is Fully Connected from CoreWeave, now around four thousand people, the biggest thing CoreWeave has ever done. Pitbull is headlining the party. ThursdAI listeners get a free ticket with the code on screen during the show THURSDAIFC2026, so if you are in town for OpenAI DevDay, stay one more day. Wolfram and I will be doing ThursdAI live from the floor, plus conversations with CoreWeave folks about the industry.
    Also: last weekend’s hackathon was a hit, and because TypeSafe sponsored it, everyone who showed up got early access to Jev before the public launch. That is the kind of thing you get for coming to our events. Sign up next time.
    Quick hits: voice, harnesses, and a stealth model
    We spent the airtime on the three themes above, so the rest of the week gets the lightning treatment. Links for all of it are in the TL;DR.
    Gemini 3.8 Live and 3.8 Live Extended Thinking (X, Blog, Model card)
    Wolfram’s favorite of the week. Two real-time voice models on Gemini 3 Pro, with Extended Thinking claiming 82.6 on Artificial Analysis’ speech-to-speech index, top of the board ahead of GPT-Live-1 Astra at 81.5, 97 languages switched mid-sentence, and tool calls that run in the background without pausing the conversation. Wolfram already built a phone app on it that talks to his Hermes, and his latency argument is the one I will remember: “if I say turn on the light while I’m going down the stairs, I could have fallen down the stairs already.” My reaction, which I could not suppress, was that home control is exactly the deterministic click Jev should be making.
    GPT Live 1 arrives in the API (Blog)
    The voice behind ChatGPT’s live mode is now a model you can call, demoed on a talking Reachy Mini. Peter says Arena does not test live voice yet because it is too personal to score quickly. Remember Moshi a year ago, fast and stupid, no tool calls, no interruptions? We are a long way from there.
    StepFun StepAudio 3 tops the voice leaderboards (X, Blog, Playground)
    Five API-only audio models, no open weights. Real-time is number one on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max ties the best word error rate on the board at 1.7 percent. I tried to demo it live and had zero credits on a fresh account, so a note to every lab: if you want your tool used on air, give a new signup enough credits for one demo.
    OpenAI Agents API: the Codex harness as a managed service (X, Blog)
    One API call gets you a production agent on the harness that runs Codex: compaction, tool search, parallel programmatic tool calls, subagents, MCP, and hosted sandboxes from three cents per twenty minutes, with the harness itself Apache-2.0 and no platform fee. Peter, who used to build this inside organizations, called it golden, because 95 percent of “AI engineering” is stupid infrastructure. My note: models behave better in the harness they were trained with, and if OpenAI shipped this two weeks before DevDay, I have to wonder what they are saving.
    Union Alpha, a free stealth model on OpenRouter (X, OpenRouter)
    Anonymous, multimodal, 262K context, free, over 100 billion tokens processed within hours. The last stealth model turned out to be Z.ai’s GLM-5.3-Flash and ZCode is a top-five app by volume, so Z.ai is the safe guess. Frontier-level performance is the provider’s own claim, so treat it as a claim.
    Wrapping up
    I closed the show by showing something I have never shown before: the editor I built to replace Descript for ThursdAI, timeline, LLM cut suggestions, a clips factory, all of it, with a plan to give every co-host’s agent access to pull their own clips. Yam said I could sell it. Wolfram said it is the proof of what we preach here every week, use the tools to build the thing you actually need. And the first thing I am wiring into it this weekend is Jev, scoring every sentence for topic, tangent, and virality at a cent per thousand.
    Next week should be bigger than this one. Sam said the thing he was most excited to ship this week slipped to next week, Grok 4.7 is due, DevDay is the week after, and Fully Connected is the day after that. If you missed any part of today, ThursdAI is a podcast, a newsletter, and a YouTube show, and the whole live stream with transcripts is on thursdai.live. Subscribe to one and go check out the others.
    Thank you Wolfram, Peter, Nisten, LDJ, an d Yam, and thank you Allie, David, and Francesco for jumping on. See you next week.
    TL;DR and show notes
    * Hosts and Guests
    * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
    * Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg
    * Allie Laabs - DevRel, TypeSafe AI (@allietheicon)
    * David Pawlan - Assistant Benchmark, Merit Systems (@DavidPawlan)
    * Francesco Bonacci - Founder, Cua (@francedot)
    * Big CO LLMs + APIs
    * TypeSafe AI launches Jev, a non-LLM “System One” decision model from ex-OpenAI RLHF lead Diogo Almeida: 70-500ms decisions, $42 per billion input tokens, free output, Choice / Score / Noul primitives, 32K context, waitlist (X, Blog, Nathan Flurry, Cua jev-use, Vercel, GitHub)
    * Pace the Frontier: Dario’s essay proposes embedded evaluators, lab coordination with antitrust cover, and global coordination; Sam and Elon agree, Zuck and Sacks reject, Demis endorses, Trump calls it a hoax on a live call with Jensen (Dario, Sam, Elon, Zuck, Sacks, Demis, Letter)
    * Mustafa Suleyman publishes a ~30-page MAI Code of Conduct: AI is a tool, no resisting shutdown, human-legible agent comms, model welfare is wrong (Blog)
    * Google DeepMind launches the DeepMind Institute with five essays and Shane Legg saying AGI is approaching (X, X, Site)
    * OpenAI Agents API in public beta: the Codex harness as a managed service, compaction, tool search, subagents, hosted sandboxes from $0.03 per 20 minutes (X, Blog)
    * Union Alpha, an anonymous free stealth model on OpenRouter with 262K context, 100B+ tokens in hours, likely Z.ai (X, OpenRouter)
    * Personal AI Assistants
    * Meta Muse rolls out invite codes (1B tokens each, up to 20 friends), voice calling to businesses, a Tailscale connector, and Muse for Mac (X, Mac, Site)
    * Instinct in talks at a $10B valuation, ships Concierge phone calls, TOTP support, and the Trusted Person agent network (X, X, The Information)
    * Grok Bot adds voice, 1Password, and local-machine browser proxying (X, X)
    * Assistant Benchmark ranks 116 submitted assistants across 16 hand-tested dimensions; Muse 9.1, Instinct 8.4 (X, Site, Methodology)
    * Cua ships jev-use (Jev + Cua Driver) in dev preview and skills over MCP (X)
    * This Week’s Buzz (Weights & Biases & CoreWeave)
    * Fully Connected 2026, Sep 30 to Oct 1 in San Francisco, the day after DevDay, Pitbull headlining, free ticket for ThursdAI listeners (Register)
    * Last weekend’s hackathon attendees got early access to Jev via TypeSafe’s sponsorship (X)
    * Voice & Audio
    * Gemini 3.8 Live and 3.8 Live Extended Thinking, 82.6 on the speech-to-speech index, 97 languages, async tool calls (X, Blog, Model card)
    * OpenAI GPT Live 1 available in the API (Blog)
    * StepFun StepAudio 3, five audio models, #1 on Artificial Analysis real-time voice, 1.7% WER on ASR Max, API only (X, Blog, Playground)
    * Show notes
    * My Jev-powered X timeline classifier extension (X)
    *


    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
  • ThursdAI - The top AI news from the past week

    OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news

    2026/09/11 | 1h 43 mins.
    Hey yall, welcome back to ThursdAI, this is Alex, let me catch you up!
    Today on the show, we covered 1 week with Astra (hint, it’s not quite AGI yet despite what we were told), DeepSeek V4.1 catches up to the frontier at a fraction of the cost, and Meta launches a free AI agent with it’s own computer, that will take over the OpenClaw/Hermeses of the world for most people. Also huge this week, OpenAI claimed that a swarm of 10K agents of their unreleased model solved the Navier-Stokes, one of the millennium problems!
    I was stoked to have Chris Alexiuk from Nvidia on the show to cover the innovations DeepSeek put into this latest model!
    Oh, and the guy who quit Anthropic this week, and wrote an essay about “AI is going to kill all of us” somehow got 130M views on X, a mirriad of TV interviews and rekindled the doomerism movement, we talk about that too!
    Also, I already told about FullyConnected, CoreWeave’s premier conference that’s coming up, but they told me about a new announcement today, and you’re not going to believe who it’s about (not AI related). As a reminder, ThursdAI subscribers get a free ticket!
    Ok, let’s dive in!
    ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    OpenAI claims a Navier-Stokes solution from a 10,000 agent swarm from an unreleased model - with some drama (X, Blog, Paper)
    If you’ve been reading ThursdAI for a while, you may remember that a model couldn’t tell which is higher, 9.9 or 9.11 and naysayers said that “AI can’t math”
    Well, this week OpenAI claimed that a swarm of 10K agents, of an unreleased model, found a solution to the Navier-Stokes problem, in about 88 hours! They published a huge 160+ page paper and a Lean proof of the solution! I said on the show, this is like the moon landing equivalent of AI doing things humans didn’t do before!
    Now, keep in mind, this is only a claim from OpenAI, the Clay Mathematics Institute is still reviewing it (moved the status of this problem from “unsolved” to “under review” so this isn’t independently verified yet) but it’s still an insane deal.
    The agents sent 2.7 million messages and burned about 130B tokens, of a model that has no public price yet, so it’s hard to estimate the cost of this run, all for a 1M prize that OpenAI said they will not claim.
    The coolest thing I think we got from this paper, in addition to solving on of the most important and hardest problems in mathematics, is this chart above, where OpenAI shows their unreleased model and how much better it is on Math problems compared to... GPT-6! The “AGI” model we got just a week ago. So so much to look forward to.
    The drama behind this
    I don’t want to get into the drama behind this too much, but if you’ve seen this online, there release wasn’t without it’s hiccups. Apparently, OpenAI caught wind that an a duo of researchers Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in personal capacity), have independently solved the related Euler problem and were about to go public. Apparently, the duo used a mix of GPT and Claude to work on this problem. OpenAI caught wind of this and started working on their own solution on September 1st. A day before OpenAI dropped their release, Buckmaster posted that OpenAI is about to drop it, and in communications with him, they offered him to co-author the paper, only if the Anthropic guy is removed. To which he said no.
    He then had some claims that maybe OpenAI trained on some of the papers and chats they put into Codex, which OpenAI refuted, while noting “we cannot rule out the possibility that de-identified usage data helped improve the model”.
    OpenAI also say that the OptOut toggle works, and neither mathematician provided a screenshot that they opted out of the training, so it’s hard to say what really went in there
    My take, I don’t really care. Two years ago, we told you that reasoning is coming and AI is going to be doing superhuman things, and we finally see the first signs of it, this is a problem that no humans was able to solve for over 60 years! Navie-Stokes probably doesn’t change your friday, but there are so many other that they can solve with this approach! Cancer research, room temperature superconductors (remember LK-99? that’s also a search problem) and much more. Kudos to OpenAI for this, and I’m looking forward to see the new heights of mathematics. And for the mathematicians who “disagree” with OpenAI solving or not solving this problem, why don’t you post your own Lean proofs instead of fuming publicly online?
    GPT-6 - Not quite AGI, yet?
    After a week with Astra GPT-6, and Jensen Huang announcing AGI is here, I think we can do a quick recap
    I and all the co-hosts have been Astra-maxxing for a whole week now and the results are in. This model is incredible at coding, it goes very very deep, however, definitely not AGI quite yet. It has a very jagged frontier, it does some things incredibly well (most demos online are building whole games and apps in 3D and those are mind-blowing) but I started seeing many folks go back to GPT 5.6 Sol etc.
    I tink some of it also has to do with price, Astra runs faster and is significantly more expensive so token limits for folks are draining fast, but also with controllability, folks are likely still using their old and unoptimized prompts. Speaking of prompts, here’s a great writeup from OpenAI how to rework your prompts (which you can send to Astra and have it review your prompts!) for better results
    The AI doomerism has quite a week (Coxon post, Thread, Hubinger, Marks, Christiano, An Alien Mind)
    I started ThursdAI with the notion to counter anti-ai and doomerism, and bring positivity to the AI world, so we had to cover this. An Anthropic employee who previously worked at OpenAI, posted on X about leaving Anthropic, saying that both labs are racing towards uncontrollable self-improving superintelligence and that it could be a disaster of the “end all of humanity” kind.
    His post sits at 130M impressions (after being basically a nobody on X before) and he’s been interviewed by Fox, AP, Time magazine, WSJ, NBC and a host of senators, Bernie (chief doomer) included, reposted his post on the same day!
    The funniest thing is that this post got about 26x more attention than Ilya Sutskever’s post about leaving OpenAI. Just nuts
    Within four days, seven current and former employees from Anthropic, OpenAI and DeepMind said in public that they believe that AI could kill us all. Also notable that Paul Christiano, who is one of the most interesting AI doomers out there, has joined the OpenAI foundation, and Daniel Kokotajlo, the famed OpenAI whistle-blower, joined Joe Rogans podcast to talk about AI doom.
    Each one of these incidents in vacuum is normal, but having all these happen in a a span of a few days just feels, inorganic. Some folks are even saying that this is a well coordinated doomerism campaign!
    I want to be fair to Coxon, folks who worked with him at OpenAI say he’s the real deal, and cares deeply about AI safety and humanity, however, he only worked at Anthropic for 6 weeks before publicly leaving, and now every interview he does sayd “Ex Anthropic employee”.
    Whether it’s a coordinated effort or not, it’s still a very important discussion, after the pacingthefrontier letter and the HuggingFace hack incident, and it seems to have made waves. Sam Altman just told staff that he’s not opposed to pausing and have petitioned the US government to regulate AI as well
    Look, I don’t disagree that we’re dealing with a very powerful technology, however I don’t believe that scaring the bajeesus out of everyone is the right way to handle this. Politicians use fear to get votes and get elected, they don’t really care about tech progress, and framing this in a way that “we pause or we’re dead” ignores all the good that AI is about to do. Cure cancer, find solutions to climate change, helping solving povery. All these seem like out there ideas but they are coming. The US GDP is already growing at an unprecedented rate and a lot of it is due to AI. In any rate, as I said on the show, I’m not against pausing, just after we solve cancer. Then we can pause and reassess, till then, nobody is telling me how China’s government is going to pause if US pauses, and if they don’t, they will reach superintelligence before we, and I don’t want to live in that world!
    Open Source AI
    DeepSeek V4.1 Flash: the whale is back, and it’s cheap (X, HF, TokenJuice)
    Speaking of.. chinese AI! Deepsek (The whale) resurfaced this week with V4.1 Flash, and don’t let the name fool you, this is not just a .1 small update. 552B with only 8B active on prefill and 16B on decode, 1M context, trained from scratch on 45T multimodal tokens! plus as always, MIT license.
    Chris from Nvidia joined us to break it down, and his main point stuck with me: every DeepSeek release comes with one of the best engineering reports you can read, and this one is the most data-pilled they’ve ever done. The paper basically says it out loud, everything else is nice, but it’s the data. 45T tokens isn’t a huge number anymore, but the cleaning they describe goes way beyond what anyone else publishes (Chris said even his own beloved Nemotron’s open pipelines are less thorough).
    Yam opened a new corner of the show, “I Told You So”, because DeepSeek went back to an encoder-decoder architecture. Not the old one from before GPT-2, this one has a pile of battle tested tricks that make it work at half a trillion parameters. His verdict after testing it all day: the best open weights model you can host for coding right now, and it’s not even close to the largest one.
    This chart is the one to look at. KV cache per token went from 389,000 bytes in the first DeepSeek (Nov 2023) to about 890 bytes now. Nisten did the math live, over 400x smaller. That’s why this model is so cheap to serve, and as Yam kept yelling, we shouldn’t take it for granted, this is the actual moat and they just put it in the open.
    Evals, briefly: 90.6 on Terminal-Bench 2.1 (above Opus 5 and GPT 5.6 Sol), 74.2 on DeepSWE 1.1 (also above both), and on an Open Design leaderboard it lands second behind Astra at two cents a task. Not twenty cents. Two. DeepSeek’s own evals, so the usual asterisk applies.
    Two more things. Friend of the pod Aaron Batilo (he works on CoreWeave Inference, this is a side project, not sponsored) put up TokenJuice.ai, free and fast DeepSeek V4.1 Flash in exchange for your requests as training data, hosted in the US. Nobody wants free DeepSeek? Go try it. And Nisten had Astra build a 3D visualization of the whole architecture, every weight a cube sized by its bytes on disk, link in the TL;DR.
    This Week’s Buzz 🐝: Pitbull is coming to Fully Connected (Fully Connected, CoreWeave Hacks)
    Tbh, I was super surprised by this! This isn’t about AI, but exciting non the les Pitbull, Mr. Worldwide himself, is headlining Fully Connected! Yes, really. So the free ticket we give ThursdAI folks is now also a Pitbull concert ticket!
    The rest of Fully Connected is still a great reason to come: September 29 to October 1 at Moscone South in SF, 2,000-plus engineers, Sarah Guo from Convitction hosting, Dr. Fei-Fei Li from WorldLabs keynoting, and ThursdAI live from the floor! The code for a free ticket is THURSDAIFC2026, register at the link above. 19 days out.
    Before that, CoreWeave Hacks is this weekend, September 12 - 13 in our SF office. The theme is Agent Loops, and the prizes are crazy. Definitely sign up and come hack with us!
    Meta launches Muse, a free 24/7 agent with its own computer - the OpenClaw for your mom (and you!) (X, My thread, Site, Security, My YT breakdown)
    Meta, after selling back Manus to China, is finally back with a agent product, and this one is actually quite amazing! It’s called Muse, and it’s powered by Muse Spark. It’s really really fast, and oh... it’s free (up to 100M tokens per week?). With it’s own cloud VM, and a browser, it seems like the folks at MSL are going towards taking over the agentic world.
    If you’ve installed OpenClaw and moved to Hermes, if you’ve played with Grok Bot (still excellent) and got an Instinct invite, you know the drill. An agent that with your permission can read your emails, browse (and do shopping for you). But Muse is different in a few key ways, not least of which their insane distribution (Facebook, Instagram, WhatsApp, Messenger all having over 2B users)
    My first experience with muse blew me away, they have a very strong Stripe Link integration (native) and my Muse was able to get me tickets to a dinner tomorrow (Happy Rosh haShana btw!), finding the hidden link, checking my calendar and booking it, all before Instinct, the other assistant i gave this task to, even replied! It’s really fast!
    UX as the differentiator, price as the convincer
    Muse is different from the other agents, first of all, because of the price sticker. It’s free! With a very generous 100M tokens, and a beefy cloud VM, this on its own is huge. It’s also very very fast, and to be honest, I didn’t expect this level of polish at the jump!
    Meta knows how to launch products, this is VERY polished. Unlike OpenClaw or Hermes which require command line knowledge, Muse is there for you when you sign in with your Meta account, and it’s ready to go to work! Of course, like with everything, the more you share (connectors are there for Gmail, Drive, WhatsApp and tons of other stuff) the more effective and personalized it gets
    I particularly like the avatar there, you can choose to customize your own, and it shows what it does (unlike other agents that either just show progress or show you command line outputs), they even generated a few different videos of my wolf doing different things while it does them! Just wonderful.
    Self onboarding product and proactive UX
    I constantly talk about 2026 is the year of the proactive AI agent, and this seems like one of the first ones where I enjoy their proactive. Not only scanning my inbox and telling me “you got this email, what do you want me to do with it?, Muse saw that my daughter’s birthday is coming up and suggested to plan a party, which I did (it got balloons and a helium tank in Target for me, it’s waiting for me to pick up while I edit this)
    These proactive recommendations also make Muse “Self onboarding” in a smart way, especially the “ideas” tab they have on the left that shows folks who never used an AI agent before what they can do wit it. Really really well done
    The Zuck shaped elephant in the room : Privacy and Security
    But it’s META! Folks on X scream and say they will never share their data with that company, not after everything that happened in the past.
    Meta is obviously aware of how folks feel and how much sensitive data they allow these agents to accumulate, and on the security page they outline not only the safety measures they took with Muse (noted is the Sentinel additional process that looks at all your conversations and incoming data to your agents and judged independently if it’s safe of if anyone is trying to hack you), Meta is also offering a bounty of up to $300K to anyone that finds a security issue with Muse.
    On the privacy side, the highlight for me is the Confidential VM announcement, while they are still working on this, Zuck shared that Meta recruited Moxie Marlinspike (the founder of Signal) to work on this, he’s the guy who also helped Whatsapp E2E encryption. This should make the virtual machine (that Muse uses to hold all your data) incredibly secure, so much that Meta (verifiably) will not be able to access your content, even if Zuck himself wants to look at your calendar! I’m super stoked by this and hope that this approach will be open sourced and adopted by all the other companies!
    Look, I didn’t mean to turn this into a full review of Muse, there’s a LOT more there, like connectors, some of which are proprietary to Meta, and some are novel (like the iphone native ones) and the upcoming 1Password integration.
    I posted a full review it on YT and definitely check out that full review, but I think that when Meta launches something that big, you all ought to know what the excitement is about.
    LMK if you’ve tried this in comments or if you don’t trust Meta with your data, let me know as well
    ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

    Instinct - the other agent in the room this week (X)
    We didn’t give Instinct too much attention in the show, but it’s also been blowing up lately, an imessage native agent that can do a lot. This week, it got an update with it’s own email you can fwd things to and a cool concept of a “trusted person network” where my Instinct can talk directly to my wife’s Instinct (and to the Instincts of other folks I trust) to sort out plans between us, so the agents negotiate dinner and I just show up. Instinct is doing some very very cool things, and is also free, definitely worth hunting down for an invite (LMK in comments if you wnat one, I have a few left).
    Now, the funniest thing about X AI folks, that they wouldn’t trust on of the biggest company in the world with their data, but they will yolo it into a startup by a 22year old 😂
    Apple iPhone 18 Pro, A20 Pro, and the AI bits (X, Newsroom)
    We mostly skipped the Apple event on purpose. iPhone 18 Pro with the A20 Pro chip (a dual 16-core Neural Engine, double the AI compute of the A19 Pro), Siri AI landing in English beta with iOS 27 on September 14, and the iPhone Duo foldable at $1,999, which has nothing to do with AI even if it’ll be a nice vibe coding machine. The AI parts that actually interested us: AirPods live translation (shout out Milosh in the chat), and the Apple Watch continuously transcribing on device with a “what did they say” button, no cloud. I’ve been using Siri AI. It’s fine. It’s nowhere near the agents above.
    AI Art & Diffusion
    GPT launches updated Images 2.5: Flare and Sunburst (X, Blog)
    Two new API models, Flare for speed and volume and Sunburst for careful editing (with native transparent backgrounds), both at $30 per million image output tokens, and OpenAI claims up to 50% lower latency than Images 2.0. In ChatGPT you get Sketch (type @Sketch and draw), comment-on-image editing and templates.
    My hands-on take: I spent the evening making this week’s thumbnails with Images 2.5 through Fal and with the help of Cursor, and the first round was rough. The artifacts were very bad and it make me look really fat.
    Turns out that was on me, not the model. My prompts were written for GPT-image-2, full of “8K, cinematic, hyper-realistic” prompt addition that GPT image 2.5 Sunburst treats as “overcook everything”, and the only reference photo we fed it was itself AI-generated.
    So I asked Fable to do some prompting research, fixed the inputs, set the quality on high rather than max, and one tiny line that says “real photograph, no heavy retouching”. Second round it beat Nano Banana Pro on both my likeness and the text on the tiles! Very impressive and very fast!
    Wrapping up
    There is so much more that was released this week, for example, I used the new Cursor “projects” feature with Fable 5.1 as my chief of staff today and it was marvelous, no more just “chats” and it even helped me wrangle my Grok Bot and Muse agents. OpenAI launched their GPT-live-1 model that powers their voice chat in API, so now you can make the same amazing experiences in your own apps as ChatGPT voice mode has, which is absolutely best in class, and both Google and Suno released updated music models (Lyria 3.5 and Suno 6). I’ve added all those things in the notes.
    I think this week we saw both ends of the AI spectrum, huge advances in personal (muse) and frontier (Astra) AI, and the rise of Doomerist on the other. I can’t wait to check in to see how next week is doing!
    See you then, and if you haven’t yet, a subscription to ThursdAI is free and really helps us to keep going. Thank you!
    TL;DR and show notes
    * Hosts and Guests
    * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
    * Co-hosts: @WolframRvnwlf, @nisten, @yampeleg, @ldjconfirmed
    * Guest: Chris Alexiuk, NVIDIA (@llm_wizard)
    * Open Source LLMs
    * DeepSeek V4.1 Flash: 552B MoE with 8B/16B active, encoder-decoder, 45T multimodal tokens, 400x smaller KV cache than DeepSeek V1, MIT; free for a limited time on TokenJuice (X, HF, TokenJuice, Nistens visualization)
    * Desert Ant Labs debuts 18 on-device models with native SDKs, including Voz speech-to-text (X, Blog, HF, GitHub)
    * InclusionAI Ling-3.0-flash-VL, 124B vision language MoE with 5.5B active, MIT (X, HF)
    * Big CO LLMs + APIs
    * OpenAI claims a Navier-Stokes Millennium Prize solution from a 10,000 agent swarm on an unreleased model, with a Lean proof; Clay review pending (X, Blog, Paper)
    * Jacob Coxon resigns from Anthropic, seven insiders warn of extinction risk in four days, 130M+ impressions (X, Thread, Hubinger, Marks, Christiano, An Alien Mind)
    * OpenAI says it reached the automated research intern milestone, automated researcher targeted for March 2028 (X, Blog)
    * Apple debuts iPhone 18 Pro with A20 Pro, Siri AI beta, iPhone Duo, AirPods live translation (X, Newsroom)
    * OpenAI brings GPT 5.6 Sol and GPT-6 Astra to ChatGPT Voice (X, Release notes)
    * OpenAI launches GPT-live-1 in the API (X, Docs)
    * This Week’s Buzz
    * Fully Connected 2026, Sep 29 to Oct 1, Moscone South SF, Pitbull headlines the concert, free tickets for ThursdAI listeners (X)
    * CoreWeave Hacks: Agent Loops, Sep 12 to 13 in SF (Luma, X)
    * Tools & Agentic Engineering
    * Meta launches Muse, a free 24/7 personal agent on its own Linux VM with native WhatsApp, iPhone connectors, Stripe Link and a Confidential VM roadmap (X, Alex’s thread, Site, Security)
    * Cursor launches Projects view- one view to rule them all (X)
    * Cognition ships SWE-2, near frontier coding scores at up to 70% lower cost, free for a month on Devin paid tiers (X)
    * Instinct, the iMessage agent, adds email and a Trusted Person network (X)
    * AI Art & Diffusion
    * OpenAI launches GPT Images 2.5 with Flare and Sunburst models (X, Blog)
    * Voice & Audio
    * Google launches Lyria 3.5 full song generation in Gemini, AI Studio and the API, $0.08 a song (X, Lyria, Docs)
    * Suno releases Suno 6


    This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
More News podcasts
About ThursdAI - The top AI news from the past week
Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news
Podcast website

Listen to ThursdAI - The top AI news from the past week, The Rest Is Politics: US and many other podcasts from around the world with the radio.net app

Get the free radio.net app

  • Stations and podcasts to bookmark
  • Stream via Wi-Fi or Bluetooth
  • Supports Carplay & Android Auto
  • Many other app features