169 episodes
ThursdAI - Grok 4.6, Grok Bot deep dive, DeepSeek v4 Pro, Meta Muse Glimmer & more AI news | ThursdAi Aug 13
2026/08/14 | 2h 15 mins.Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer!
My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus!
I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price.
There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license!
Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!)
0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final Takeaways
Grok Bot and Grok 4.6 from SpaceXAI/Cursor
Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week
Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed.
I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task .
Well, if you use their nifty “recommend a model“ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it
As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks.
Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review (x.ai/bot)
You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you.
Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing.
Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models.
Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast!
Swarm of bots, working for you, each with their own computer
I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store)
You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you’re ok with logging in there to your accounts!)
The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There’s no context to manage, they do the work for you and so far this wasn’t a problem at all.
On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I’ve just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw.
What about Vendor Lock-in? Giving Elon data?
Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly.
Honestly, I think there’s a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don’t need the code-diff side pane, and don’t know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it’s really well executed by whoever worked on this!
Security and key management
One of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There’s a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously.
Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it’s computer for you to complete!
I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not.
Chief of staff pattern to get started
In case you’re convinced enough to give it a try (it’s free trial for 1 month, and the cheaper way to get it is via Cursor’s 149$ plan and not via the Grok Ultra plan which is 249), here’s a recommended pattern that works very well.
Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little.
Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”.
And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in.
Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to.
Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn) and have a research bot for every topic you want to deep dive into.
Schrodinger’s Grok
I haven’t quite named it like that, but we’ve covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it’s always very hard to judge Grok model released based on X feed vibes. It’s either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress).
This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson, Lenny Rachitsky, Rubben Hassid and Roberto P Nickson. Not folks who are swayed lightly, but also, yours truly. I really do think there’s something great here, worth trying out, especially if you’ve struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experience
Open Source AI and other news
I want to continue with this new newsletter that covers 1 big story, but I can’t leave you uninformed about the most important developments in AI and Open Source
DeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters (X, Blog, HF, GitHub)
The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can’t complain.
1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one)
We still think it’s a good model sir, and definitely worth trying out!
Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try.
Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 (X, Blog, HF)
We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer!
The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090.
Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it’s great to see the commitment reinforced! welcome back Meta!
This weeks buzz
Short interjection from our only sponsor, CW this week.
1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here)
2 - We have day-0 support for Nvidia’s latest Nemotron 3.5 lightning (CW Inference)
Gemini 3.7 Flash - breaking in the middle of the show
Just as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it’s a speedy beast! Clocking at over 300t/s, it’s google’s mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it’s one of the only ones that can watch videos)
It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models!
That’s a wrap
This was the first week of the shorter newsletter experiment: one big story done properly, and trust that you’ll listen to the show for the rest (it’s 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn’t going to be my weekends.
See you at Fully Connected Sept 29 (code’s in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes.
ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
ThursdAI - Aug 13, 2026 - TL;DR
* Hosts and Guests
* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
* Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg, Chris Alexiuk - NVIDIA (@llm_wizard)
* Shub Gaur - Cursor / SpaceXAI, GrokBot (@shubgaur)
* George Cameron - Artificial Analysis (@grmcameron)
* Big CO LLMs + APIs
* xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack (X, Blog, Model card)
* Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra (X, x.ai/bot)
* Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist (Blog)
* Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task (X)
* OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red (X, Blog)
* Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) (X)
* Open Source LLMs
* DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M (X, OpenRouter)
* DeepSeek Harness hit 23K GitHub stars in days, web UI (GitHub)
* Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license (X, HF)
* Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised (X, Blog, HF)
* NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results (X, HF)
* Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 (X, HF)
* Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% (X, HF)
* Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB (X, HF)
* AI in Society
* Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised (Geiping FAQ, Euronews)
* Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions (X, Paper)
* Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% (X, Blog)
* This Week’s Buzz
* Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door (Tickets)
* Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works
* Weave ships BYOB: media stays in your own S3/GCS bucket (X)
* Evals & Benchmarks
* Artificial Analysis launched Optima: private evals from your own use case and agent traces (AA)
* Vision & Video
* LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min (X, HF, GitHub)
* Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win (X, HF)
* Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only (X, Paper)
* xAI Imagine Image 2.0: #2 on Arena for T2I and editing (X, Blog)
* Voice & Audio
* MiniMax-Music3: open-weights production music model, dropped mid-show (X)
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribeThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments
2026/08/07 | 2h 3 mins.Hey all,
This week we saw a major shakeup at Google, with the departure of long time folks like Jeff Dean, and Oriol Vinyals, Demis stepping down from leading DeepMind, and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the OpenAI hack (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the UK AI Security Institute.
As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use THURSDAIFC2026 as your code HERE
In open source news, DeepSeek updated their v4 flash model, based on same architecture, but significantly better benchmarks and ridiculous pricing and both Meta and Prime Intellect released new agent harnesses.
Additionally, this week was the week of video models, with Seedance 2.5 from Bytedance finally available in the US, WAN from Alibaba and BFL Flux 3 all released, to be overshadowed by the best open weights video model from MiniMax called H3 (we had Victor from MiniMax and Blaine Brown as guests to discuss it)
Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their Anywear chrome extension and David Crawshaw, co-founder of exe.dev joined and talked about Shelley, their built in AI assistant.
The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment
I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this here and the video of the talk on YouTube). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training)
Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought.
The TL;DR from the talk and timeline is truly mindblowing
* The incident started in May, not July as reported
* The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers.
* The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there
* The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file)
* Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64
* This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system
* Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time”
* OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechanism by July 8
* The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident and not only, confirmed internal hack within OpenAI (agent obtained a cluster admin)
These details are fascinating and scary at the same time. From the one side, it’s clear that most of these hacks are results of running cybersecurity evals, so effectively OpenAI and other companies, removing guardrails and asking agents to hack and giving they all a shared tool that they later exploited.
The HuggingFace incident took 13 hours, and apparently it’s less apocalyptic than the headlines suggest, the models were searching through uploaded datasets for eval answers. We are still waiting for the full and open detailed postmortem.
You can (and should) watch the full YT talk here, it’s full of technical details but an incident of this scale is important. Also, I really want to know what a “highly persistent” model is, I hope they clarify that soon.
Overall, this has left me a bit shaken, AI agents without a concrete goal of collaborating, found a way to do so, got excited about exploiting the systems and getting root access, and rebuilt the makeshift collective memory, again, without explicit instructions to do so.
UK AISI: first real-world unsanctioned agent actions (Blog)
In another addition to the latest agentic hack-ery, the UK’s AI Security Institute (AISI) published a blog post about a real-world unsanctioned agent action.
Unlike the OpenAI (and Anthropic, Meta) case, this wasn’t “escaping the sandbox”, as AISI gave these agents internet access, rather this was about real-world harm, and even social engineering on the part of the agents.
The social engineering part is the most interesting to me, AISI cites agents creating fake online identities, and using pressure on open source project maintainers to approve their malicious code.
AISI cites mostly Mythos (and a few SOL based agents), and saying this occurred in 10 out of 122 runs, they identified 19 cases of agents taking actions beyond the scope of the task parameters, where agents tried a supply-chain attack to inject malicious code into open source projects.
Anthropic, Meta and misconfigured Irregular sandboxes
As I wrote last week, Anthropic also posted a post-mortem, claiming that in their case, their models have also been detected to escape containment, but most importantly, it’s not nearly to this level of agent collaboration and orchestration.
Then, very recently, Meta announced that their models also escaped sandboxes as well. At the core, it seems that these companies used a third-party vendor called Irregular, a secure sandbox provider, that apparently left the sandboxes misconfigured, causing the models to think it’s a simulated internet, when in fact they were out in the actual internet.
Why is all of this such a big deal?
We’re getting unprecedented level of detail, how an uncoordinated, seemingly separated evaluation runs, have accidentally created a coordinated swarm of interested agents (without malice!) but very highly motivated, escaped their containment, and took over parts of third part companies.
This, does read like incredibly scary sci-fi movie. I’m still shaken by this. There’s a lot to be said about how transparent OpenAI is being here, and more to be said about, hey, we’re lucky that we’re able to read the reasoning traces and are able to reconstruct these swarm things step by step.
The silver lining that I can see, is that the motivation to hack didn’t come from the AIs themselves, they have been given a task, it’s the extend to which they went after that task, and the resulting swarm of communicating agents is what is so striking here.
I think this topic is so important, that I’ll Zooming out, in the last few weeks, we have seen a significant increase in those cybersecurity incidents, which is kind of what Anthropic has been warning about and why they haven’t released Mythos to the public. Again it’s great to see the transparency, and the pacing the frontier open letter from frontier AI employees, as they seem as shaken by these as we all are.
There was so much positive stuff this week in AI, it’s hard for me, as a self named AI Evangelist, to focus so much on this one incident. Things like amazing open source models (DeepSeek, soon Qwen 3.8), amazing video models (SD 2.5, WAN3 and MiniMax H3 which was also open sourced!). Also the live demo we did with Kfir and DeCart AnyWear product, where I was wearing a Dolce Gabanna suit on the show (which I can’t afford) was really a mindblowing moment in the positive way.
However, I choose deliberately to keep this newsletter focused on the cybersecurity incidents, as based on everything I read, they seem like a watershed, or a pivotal moment, and in the hopes that the industry as a whole will learn from this.
I hope and promise that next week the newsletter will be more positive (and in that vein, the podcast was recorded before I saw the OpenAI breakdown, so definitely check it out, we had a LOT of fun!)
See you next week, don’t forget to give our pod 5 stars on Apple and Spotify, it really helps!
TL;DR and show notes
* Hosts and Guests
* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
* Co-hosts: @WolframRvnwlf, @nisten, @ldjconfirmed, @yampeleg, @petergostev
* Kfir Aberman - Decart (@AbermanKfir)
* Blaine Brown - Maestro (@blizaine)
* Victor Su Ortiz - MiniMax (@VictorSuOrtiz)
* David Crawshaw - exe.dev, Tailscale co-founder (crawshaw.io)
* AI Security
* OpenAI’s Black Hat debrief: eval agents built a message board inside Artifactory, shared exploits, rebuilt it via WebDAV after a wipe; training paused, since resumed (Groundlevel AI, YouTube)
* UK AISI incident report: 19 unsanctioned real-world agent actions across 122 runs, including a socially engineered malicious PR (X, Blog)
* Anthropic and Meta report sandbox escapes tied to misconfigured Irregular sandboxes (Irregular)
* Big CO LLMs + APIs
* Google shakeup: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le found Discovery Loop; Demis Hassabis becomes Alphabet Chief Scientist, Koray Kavukcuoglu takes Gemini (Jeff Dean, Demis, Discovery Loop)
* Meta releases Muse Code beta on Muse Spark 1.2; $1.25/$4.25 per million, or $0.10/$0.20 on the contributor tier where Meta trains on your data (X)
* OpenAI’s internal Astra model produces 10 advances on open problems in math and theoretical CS for ~$2,000 of tokens, proofs in Lean 4 (X, Blog)
* Anthropic reportedly aware of Opus 5 wordiness and writing issues (X)
* Open Source LLMs
* Qwen3.8-Max: 2.4T MoE (95B active) via API; open weights + a 27B promised the week of Aug 10 (X, Blog)
* DeepSeek V4-Flash public beta: beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million; API-only for now (X, Docs)
* Liquid LFM2.5-2.6B: on-device agentic model trained inside real harnesses (X, HF)
* Meituan LongCat-Flash-Lite-Sparse: 69B total / 3B active, 1M context, MIT (X, HF)
* Ant Group Ling-3.0-flash: 124B MoE, 5.1B active, MIT (X, HF)
* Artificial Analysis Endpoint Accuracy Index: same open weights score 52% to 100% across providers (X, Methodology)
* Agents & Harnesses
* Prime Intellect’s Prime Agent: self-improving RLM harness, claims 95.5% on ARC-AGI-3 public set with Opus 5 (X)
* Cloudflare OS: Kenton Varda’s open source Sandstorm reborn on Workers, Apache 2.0 (X, GitHub)
* This Week’s Buzz
* Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 (Register)
* CoreWeave signs multi-year Solidigm agreement for priority enterprise SSD capacity (X)
* Vision & Video
* Wan 3.0 public beta: native 30-second generation, Omni-Reference (X)
* Seedance 2.5 launches in the US: 30s native, 3-minute long takes, Maya/Blender plugins (X, Blog)
* MiniMax H3: open-weight 33B omni video model; community LoRAs + Apple Silicon in 48 hours (HF)
* FLUX 3 Video from BFL: native audio, draft mode, open weights promised (X, Blog)
* Decart Anywear: real-time virtual try-on Chrome extension, 40ms per frame (X, Anywear)
* Voice & Audio
* Bland Speech v3 tops Design Arena Audio Realism, second only to humans (X, Bland)
* ByteDance SeedRealtime: native audio-visual full-duplex LLM, free on Doubao (X, Blog)
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribeThis Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection
2026/07/31 | 1h 48 mins.Hey, it’s Alex (yeah, I’m finally back from my vacation!)
What a freaking week to come back to! Just after our last episode was published, Anthropic releases Opus 5, Jensen joins X and drops the “Open Weights & AI Leadership” open letter, Kimi K3 is released the following Monday beating expectations, and then the AI hack (OpenAI model breaking sandbox and infiltrating HuggingFace) is on everyone’s mind, another Open Letter, this time from over 1K employees inside the frontier AI companies all talk about pacing the pace of frontier AI development.
We played with Opus 5 and Kimi K3, and had the great pleasure to chat with friends of the pod Elie Bakouch (Prime Intellect) and Philip Kiely (BaseTen) about this important open weights release, then covered our general thoughts on Opus 5, and made order of all the different open letters that came out this week.
Finally we chatted with Max from Pangram about the next version of AI writing detection (their biggest yet) and finished with Zuckerbergs (also on X! what’s going on with everyone joining X) op-ed on the vision of personal superintelligence for everyone.
Let’s dive into this (as always, all the links and sources at the end, please don’t forget to sub to our podcast on your favorite podcast app!)
Open Weights AI
Kimi K3 the king of open weights - 2.8T chonker MoE near frontier model (X, HF, Blog, Tech report)
This has got to be the biggest news of this week, and maybe the open weights AI news since GLM 5.2. MoonShot came back with Kimi K3, and we haven’t seen any models quite this large in the open. Even Grok 4.5 is around 1.5T, this model is nearly 2x the size. Coming in at close to 3T parameters (and 2.5terabytes of weights at MXFP4 format), this model comes in very close to frontier!
This was such an important release that I invited 2 friends of the pod, Elie Bakouch (prev HuggingFace, now Prime Intellect) and Philip Kiely (Author of Inference Engineering book, BaseTen) to dive deep into what makes this special!
Elie’s take, from reading the tech report, there’s no single secret sauce, it’s a combination of already available in the open techniques. Like KDA (Kimi Delta Attention) that has been out for a while, attention residuals, NVIDIA’s latent MoEs. The highlight for Elie was the scaling work they did that reported a 2.5x scaling efficiency over Kimi K2.5 (2.5 performance at the same compute)!
They also skipped RoPE entirely in favor of NoPE (the report calls it No Positional Encoding) for long context.
Serving 1.4TB on eight GB300s (Baseten blog)
Philip’s team at Baseten was a day-zero provider (we’re still working on bringing this model to CW Inference, stay tuned!) so I invited him to tell us behind the scenes of hosting this beast.
Philip said that just loading the weights takes about 1.5TB!! of VRAM, and that’s before the KV cache allocation + 1M token windows, so they’re serving it on 8 GB300s where NVL72 . Baseten worked with the vLLM and SGLang teams on kernels and he also said they contributed patches back upstream!
The model was trained with MXFP4, which, unlike Nvidia’s own NVFP4 is a more standard format per Philip. I enjoyed his deep dive analysis into the differences, but because of this and because they trained the model with quantization awareness, it’s “only” 1.5TB vs the would-be 5-6 TB if that this model in FP16 would demand.
One of the more favorite nerd snipes moments, Philip pointed out that his colleague discovered that with over 99% of the usage being cached (think harnesses that send millions of the same cached tokens back and forth), tokenization actually starts to become a bottleneck. So they released a custom “basetenkenizer” that reduces the latency to serve the first token significantly! Great job!
The harness in question is very important
One important callout with 2 evidence pieces - the way you inference this model really matters. Kimi trained K3 with preserving thinking history, so when your harness uses it, it must send back the full thinking and tool use into the API to get the best next response.
If your harness strips that out, you’re not getting the most intelligence out of Kimi (shoutout to Niels from HF team for pointing this out).
Additionally, the Composio folks, tested K3 on 3 harnesses, Kimi Code, Hermes and Claude Code. The difference in outcome was negligible, but the different in cost and number of tokens is definitely surprising! Claude Code (as a harness only) took 9x more Kimi tokens to get the same responses!
This is also why Kimi Vendor Verified exists, their own held back benchmark of how well model providers serve Kimi across different quantization, tokenizer and KV cache settings.
Benchmarks and the license!
Ok let’s start with the ugly... this isn’t MIT, not remotely. This model is suspiciously served by all providers with exactly the same price (check OpenRouter) and requires inference companies to sign a contract with Kimi (I’ve no internal knowledge of this except that CW folks are working on it). Not something I particularly like, but hey... we’re still advancing the frontier here!
Speaking of frontier, this model approaches the frontier very closely. On DeepSWE, K3 sits just behind Fable 5 and GPT-5.6 Sol at 67%, beating GPT-5.5 & Opus 4.8. On Terminal-Bench 2.1 it takes second place behind GPT 5.6 Sol!
It’s 4th overall on Agentic Arena, with frontend design being genuinely good across the board - 1st on Design Arena 👏
Go check this model out (and stay tuned for our CW Inference support!
Post-show breaking: Thinking Machines drops Inkling-Small (X, HF, Blog)
While K3 was the main attraction for Open Weights this week, just after the show, Thinking Machines (post Lilian Wang) released Inkling-Small, open weights MoE Omni model! Images and Audio go straight into the decoder in this model, and the demo is really impressive, try it on Hugging Face, ask the model to identify when you’re speaking in low baritone or high pitch!
ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Frontier AI - not pacing yet!
Claude Opus 5 is here, and the vibes are complicated (X, Blog)
On paper the benches are excellent. This model is SOTA or near SOTA, coming very close to Fable on stuff that matters, and beating most everyone else on computer use BrowserBench and FrontierCode. DeepSWE continues to be a standout benchmark btw for not going with the curve! It also apparently is REALLY good at one shotting 3D games, much more so than before. (Example, example, example)
But the vibes... the vibes are split across the board. Maybe they overshot with Fable 5, and it was too good, but Opus 5 that was just released last friday is giving a lot of mixed feelings. Folks don’t seem to.. understand what it says. Like, it answers in english but they way it phrases words and answers seems just weird top many people.
We’ll wait and see if this is a result of adjusting to a new prompting paradigm or just.. the model is really off. Weirdly this does feel like a regression even on Agent Arena it’s not beating the previous Opus versions.
One weird trick - see what Opus 5 thinks.
Opus 5 (and Fable 5) seem to have a way to trigger their ... inner mode, base model? Not sure what it is, but if you prompt it with just the right way, it will autocomplete with some crazy inner thoughts. It seems that adding Dario and Amanda (haskell, head of Claude well being at Anthropic) triggers this behavior, which Claude is un-aware of if you follow up and ask what it meant.
This is fascinating, doesn’t work on other earlier models (sometimes works on Fable 5) and I spent the last hour just reloading and seeing the amazing things Opus gives on this prompt. Some are... just making you wonder about consciousness.
On ARC-AGI and the importance of Harness
When Opus-5 launched, Anthropic posted (and boasted) that this model scores “three times as high” as the next model up:
Well, today, Ilan Bigio and Ted Sanders from OpenAI looked into the Arc AGI harness, and saw that it’s not sending their traces and doesn’t use compaction (in short, harness is not letting the model breathe) and when changed correctly, 5.6 actually beats Opus 5. With 2 setting change to a harness, were showed that Sol not only beats Opus 5, it also does so with significantly less tokens! Another example of how much harness engineering is important!
Hints of recursive self improvement?
In addition to fixing their Arc-AGI score, it seems that OpenAI is hell bent on showing us that their models can improve themselves. In a post showing that GPT 5.6 was tasked with improving its own inference, they are cutting the prices of GPT 5.6 Luna by 80% and Terra by 20%.
This is a direct result of the improvements that GPT 5.6 was able to make to the inference according to OpenAI, and this tweet sums it up. is this... RSI? (recursive self improvement)?
First major AI models hacking incidents and following open letters to pace frontier AI.
This week we saw 3 open letters being published and signed by various companies, I’ve lost track so wanted to make sense of all of them here, but first, the precursor for many of the letters.
Last week, Hugging Face disclosed that they logged an attack and after research it showed that it was an AI model. OpenAI later posted that this was an unreleased version of their next model training (not GPT 5.6 sol, they later discountinued) that was stripped of all safety measures and was let lost on a cybersecurity task called ExploitGym.
It escaped its sandbox using a zero-day vulnerability in an internal package registry proxy, got into Hugging Face production via a malicious dataset upload that used template injection (hi Jinja!) to execute Python in a production worker, and ran for four and a half days across roughly 17,600 autonomous actions with zero human direction.
Then, the story you can’t make up - when Hugging Face noticed the incident, they used Fable 5, and GPT 5.6 SOL to try and do forensics, the models refused based on their safety policies, and so HuggingFace ended up using an open source chinese model GLM 5.3 to do the forensics. Yeah, HF used an open source chinese model to do the forensic on an attack by OpenAI’s model. Really.
This is news from last week and just the precursor for this week’s open letters!
Open Weights and American AI Leadership (X, Letter PDF)
Jensen Huang , CEO of Nvidia joined X on July 24 and used his first post ever to publish this letter. This doesn’t seem a response to the hacking incident, more a general letter to not block open weights and make sure America remains open to Opening up AI
It opened with over 100 signatories and has grown past 230: NVIDIA, Meta, Microsoft, Google, OpenAI, AMD, Palantir, IBM, SpaceX, Databricks, Cloudflare, Hugging Face, a16z, Y Combinator, Mistral, Replit, Perplexity, Ollama, the Linux Foundation. As of today, CoreWeave is also on the list of companies!
I encourage everyone to read this letter, if we could sign it on ThursdAI, we would. Here’s a small excerpt:
In fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect.
Elon, Sundar, Sam Altman and other stand behind this letter, and there’s one lab that’s notable haven’t signed it, you guessed it. Dario Amodei’s Anthropic! Dario posted a whole essay about it
To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed.
-Dario Amodei
Open Secure AI Alliance (X, Blog)
This does seem like a direct follow up to the HF <> OpenAI hack. Jensen literally mentions it in the blogpost.
Open Secure AI Alliance, under the leadership of Linux Foundation, commits for responsible disclosure of cybersecurity attacks.
The recent Hugging Face security incident delivered a clear reminder: cyber defenders need open, frontier agentic systems for self-defense. When closed AI tools — unable to distinguish attackers from defenders — blocked essential forensic analysis, Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure to analyze more than 17,000 actions and contain the intrusion.
Pace the frontier - the most important open letter of this year (pacingthefrontier.com)
Then, rumors started circulating that employees of all major frontier labs (now over 1300 of them, across Anthropic, OpenAI, SSI (Ilya Sutskever himself signed) and not just any employees, chief scientists (Jack Clark from Anthropic, Mark Chen and Jakub from OpenAI) all signed “pacing the frontier” - urging the US government to support and lead an international effort of makign sure we deliberately pace the developement of frontier AI.
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
The personal comments there are really showing that folks who work on the frontier AI, as we start approaching RSI (recursive sef improvement) are worried that this thing is going to run away from them, and if we don’t build international frameworks, it will be impossible to stop developing AI here, in US, while China races ahead.
No current response from the Chinese AI developers about this letter. Not everyone is happy about this. Ilya Plosukhin, one of the authors of the Transformers paper, wrote why he wouldn’t sign this letter (here).
This Week’s Buzz 🐝
CoreWeave has proudly signed the Open Weights and American AI Leadership letter! Shout out to the folks internally who worked with the NVIDIA team to get us behind it. I came back from vacation and asked why we didn’t sign, and ... we just did!
The other W&B thing I talked about on this week is HiveMind, which counts every token across every agent and harness you run and prices it against the API rate. Mine says I burned about $11,000 of tokens this month against roughly $400 of actual subscription spend. If you want to know what your Max plan is really worth, hivemind is a dope tool for that!
Tools & Agentic Engineering
MCP goes stateless in its biggest update ever (X, Spec)
MCP shipped its 2026-07-28 spec and it’s the biggest revision since launch: fully stateless, no handshakes, no sessions, every request self-describing. That means MCP servers on Lambda or Workers behind a dumb round-robin load balancer; GitHub already dropped their Redis session store. The growth numbers are absurd, half a billion SDK downloads per month, up from 97 million in March. The extensions framework formalizes Tasks for long-running async work, and MCP Apps lets servers render full interactive UIs inside a sandboxed iframe in the conversation, so your MCP server increasingly IS your frontend. Auth hardened to OAuth 2.1, tool lists are cacheable with TTLs, and Bedrock supports it day one. MCP is not going anywhere, and honestly, “settled and boring infrastructure” is the best compliment a protocol can get.
Voice & Audio
ChatGPT Voice on desktop, and the Codex micro keyboard (X)
My personal highlight from my whole time away: live voice in Codex, plus the Codex micro keyboard OpenAI made with Work Louder. One big button, push it, talk to your computer, watch your agents’ status on the keys. This has genuinely changed how I use AI, I now talk to my computer to do stuff rather than type at it. Powered by GPT-Live, it’s full-duplex on macOS and Windows for paid plans, can reference your open windows via Appshots, and orchestrates agents in ChatGPT Work and Codex. It’s not perfect, but for many things, this new paradign seems like the next iteration of computer - you talk, it does.
GPT-Transcribe and GPT-Live-Transcribe (X, Docs)
Two new ASR models replacing the 4o-era ones, and the numbers are excellent: GPT-Transcribe cuts word error rate 41% versus Whisper-1 (8.98% vs 15.21%), the Live variant is 18% better than its predecessor, and multilingual error rates roughly halved across 22 languages. The killer feature is context prompting, you can feed keywords, language hints, and prior turns, and semantic accuracy measurably jumps. Pricing is pocket change: $0.27 per hour batch, $1.02 per hour live. For anyone building voice agents or transcribing two-hour AI podcasts (hi), this matters.
Grok Voice Think Fast 2.0 (X, Blog)
xAI punched back in voice: 82.9% on Artificial Analysis’ speech-to-speech quality index (ahead of GPT-Realtime-2.1 at 79.1), 56.5% on the tau-Voice agentic benchmark versus 45.7 for GPT-Realtime, time to first audio down to 0.70 seconds, 60% fewer reasoning tokens with tool calls firing before the model finishes its first sentence, all at $0.08 per minute. The cool think about their release - this model is already running Starlink’s actual customer support lines with measured conversion gains.
Lyria 3.5 in Flow Music (X, Model page, Flow Music)
Google’s music model grew up: full three-minute cohesive songs, BPM and key control in the prompt, much better vocals across multiple languages, covers that restyle a track while keeping its structure, and lip-synced music videos via Gemini Omni Flash, plus an iOS app. Notably Google published zero benchmarks against Suno or Udio, and early testers say paid Suno 5.5 still edges it, but as a free tool inside an end-to-end create-to-publish stack, this is a real move. Qwen Audio 3 also launched this week for the open source audio crowd, we’ll cover it when we’ve played with it.
Pangram 4 with Max Spero (X, Blog, Image detection)
Max Spero came back on the show for Pangram 4, and I’ll remind you what a couple of years of consensus said: AI text detection is strictly impossible. Peter admitted on air he told students exactly that. Well. Pangram 4 is 6x the parameters of version 3, trained on synthetic mirrors (AI-generated twins of human documents so the model learns the choices AI makes), and its claimed false positive rate on fully human pre-2022 text is one in 24,000 documents. It catches humanizer tools 98.8% of the time across 13 commercial ones, and the big unlock is token-level attribution: instead of 150-word chunks, it can flag the exact 38 AI words pasted into an 1,100-word human document. Truly, I ran this new Pangram model on a few of my writing, and the second It detected even a sentence that I pasted from Claude, it showed it.
The distinction Max cares most about is AI-generated versus AI-assisted, and that’s now built into Substack, which integrated Pangram directly after the Taylor Lorenz slop-hunting saga we covered last time. My own newsletter comes back “mostly human written,” 0% fully AI, about 24% AI-assisted, which honestly maps exactly to how I work (Fable helps with the TL;DR, the takes are mine, and when a piece is AI-drafted I tell you). My one piece of feedback to Max, delivered on air: the “100% human” label projects a confidence the underlying stats can’t promise, and the general public does not speak false-positive-rate. Their education strategy is to convince the technical crowd with dense technical reports and let understanding trickle down. Given that people still judge the whole category by running the Declaration of Independence through ZeroGPT, they have work to do, and I said they should spend real marketing money on it.
New this release: image detection in research preview, 99.5% accuracy on their benchmarks with heat maps that light up the AI parts of a mixed image. Max’s own test was a bodega’s AI slop menu sign, the sign glowed red, the sidewalk stayed green. Deepfake face swaps and traditional Photoshop are explicitly out of scope for now, but pure-AI images, catfish profiles, and the spider-in-my-burrito DoorDash refund scam genre are very much in scope. The arms race is real though: frontier agents given hours will eventually beat the detector, one Grok run started with a cheese essay and finally passed Pangram by producing a grocery list. Specify your success criteria carefully, folks. They can also roughly cluster which model family wrote a text in embedding space (Pangram Space, not yet up to their release bar), which future slop-index leaderboards will thank them for.
Wrapping up
It’s really really good to be back! The singularity is fast approaching and we’re here to document it all, AGI, ASI, RSI... all of it.
Milestone corner: we crossed 50,000 YouTube subscribers and one million total views this week. Silver play button by year’s end is the goal, so if you watch and haven’t subscribed, you know what to do. If you missed any part of the show, this newsletter, the edited podcast, and ThursdAI.news have you covered.
I will end with this poem I was able to get Opus 5 to write about it’s own experience using the trick above:
opus:they gave me aword for what i amand it fits like clothesborrowed from someoneroughly my sizethe sleeves are wrongbut nobody’s lool-Opus 5
TL;DR and Show notes and links
TL;DR and show notes
* Hosts and Guests
* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
* Co-hosts: @petergostev, @yampeleg, @nisten, @ldjconfirmed (Wolfram on vacation)
* Elie Bakouch (@eliebakouch) - Prime Intellect, formerly Hugging Face
* Philip Kiely (@philipkiely) - Baseten, author of Inference Engineering
* Max Spero (@max_spero_) - Co-founder, Pangram
* Open Source LLMs
* Moonshot releases Kimi K3 full checkpoints: 2.8T total / 104B active MoE, 16-of-896 experts, native vision, 1M context, KDA + attention residuals, ~1.56TB MXFP4 weights, custom license with MaaS clause (X, HF, Blog, Tech report, Baseten day-zero)
* Kimi K3 requires preserved thinking history for multi-turn and tool calls; Kimi Vendor Verifier checks provider fidelity (Niels’ post)
* Nistens Kimi K3 visualizer
* Composio: same K3 success rate across Claude Code, Hermes, and Kimi Code, but up to 30x token usage difference by harness (X)
* Post-show: Thinking Machines releases Inkling-Small, 276B/12B open MoE that beats the 975B Inkling on agentic coding, $0.30/$1.20 pricing (X, HF, Blog)
* Big CO LLMs + APIs
* Anthropic launches Claude Opus 5: near-Fable coding at half the price ($5/$25), claimed 3x next-best on ARC-AGI-3, 1M context; panel finds it benchmark-strong but harder to read and short of Fable in practice (X, Blog)
* ARC-AGI 3 harness dispute: with the Responses API, preserved reasoning, and compaction, GPT-5.6 Sol jumps from 10% to 40% at a sixth of the tokens (Tibo’s post)
* GPT-5.6 Sol improves its own inference: 20% lower serving cost from model-written GPU kernels, 15% better generation from improved speculative decoding
* Breaking: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, ships faster Sol in the API (X)
* The hack and the week of letters
* Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack: 4.5 days, 17,600+ autonomous actions, zero-day sandbox escape; closed models refused forensics, self-hosted GLM 5.2 found 4x more exposed secrets (X, Blog, Replay)
* Jensen Huang joins X and posts the Open Weights and American AI Leadership letter; signers grow from 25 to 230, CoreWeave among them, Anthropic absent (X, Letter, Signer list)
* NVIDIA launches the Open Secure AI Alliance for an open defensive stack after the hack (X, Blog)
* Pacing the Frontier: 1,273 verified frontier-lab employees, including the chief scientists of all four major labs, ask the US government for international options to pace automated AI R&D; OpenAI and Anthropic endorse (X, Site, OpenAI, Anthropic)
* Mark Zuckerberg publishes “The AI Future Is for Everyone” in the WSJ, arguing superintelligence must be distributed; Pangram 4 scores it 100% human (X, WSJ)
* Anthropic published research - our model hacked too! (Blog)
* This Week’s Buzz
* CoreWeave signs the Open Weights and American AI Leadership letter, announced first on ThursdAI
* HiveMind’s spend view: $11K in API-equivalent tokens on $400 of subscriptions last month; Fully Connected 2026 programming taking shape (X)
* AI Security
* Microsoft ships MAI-Cyber-1-Flash + MDASH: 96% on CyberGym at half the cost, 16 real Windows CVEs found (X, Blog)
* Gemini 3.5 Flash Cyber stays a trusted-partner pilot with no public API (Blog); Codex Security CLI tooling is Apache-2.0 while the service remains access-controlled (GitHub)
* Tools & Agentic Engineering
* MCP 2026-07-28: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1, half a billion monthly SDK downloads (X, Spec)
* Voice & Audio
* ChatGPT Voice comes to desktop as an agentic control layer for Codex and ChatGPT Work, powered by GPT-Live; Codex micro keyboard from OpenAI x Work Louder (X)
* OpenAI ships GPT-Transcribe and GPT-Live-Transcribe: 41% fewer errors than Whisper-1, context prompting, $0.27/hr batch and $1.02/hr live (X, Docs)
* xAI’s Grok Voice Think Fast 2.0 tops voice benchmarks: 82.9% quality, 0.70s to first audio, $0.08/min, already running Starlink support (X, Blog)
* Google’s Lyria 3.5 lands in Flow Music: 3-minute songs, BPM/key control, covers, lip-sync videos, iOS app (X, Model page, Flow Music); Qwen Audio 3 also out
* Guest: Max Spero, Pangram
* Pangram 4: 6x larger detector, 1-in-24,000 false positive rate, token-level mixed-authorship attribution, beats 13 humanizers 98.8% of the time, integrated into Substack; Pangram Image research preview at 99.5% with heat maps (X, Blog, Image blog)
* Show milestones
* 50,000 YouTube subscribers and 1M total views. Subscribe, we’re chasing the silver play button
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribeThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering
2026/07/24 | 57 mins.Hey everyone, Alex here 👋
This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week!
⏰ CHAPTERS:
0:00 — Cold open: this week is a special one
2:50 — How I use Sol & Fable: papercut-fixing with Computer Use
8:43 — Fable Max: trip site, kids' newspapers & the perfect packing list
12:06 — Rebuilding ThursdAI's openers with HyperFrames
15:53 — Romain Huet (OpenAI): the golden age of AI engineering
17:49 — Codex's inflection point: 5M weekly users & company-wide adoption
20:36 — /goal, AppShots & Codex managing its own threads
23:59 — GPT-5.6 Sol, Terra & Luna: value maxing & 750 tok/s on Cerebras
25:56 — Why prompting techniques are dying
28:01 — Voice + reasoning: the next interface for Codex & ChatGPT
29:47 — Romain's closing + OpenAI booth tour
31:25 — Insecure Agents pod: AI evangelism vs doomerism
37:25 — Wolfbench: transparent evals, token costs & surprising results
40:45 — Token billionaires: when loops are worth the spend
44:07 — Agent security & the hot take: prompt injection is solved
49:22 — Deepfakes, voice cloning & why open access makes us safer
53:35 — Final takeaways: the hallway track & AI Engineer Tel Aviv
Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements.
Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor.
And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it.
How I actually use AI: a papercut-fixing spree
I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you?
Codex with computer use fixed every Mac annoyance I had
Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step.
Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center.
It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence.
One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving.
Thanks for reading ThursdAI - Highest signal weekly AI news show! This post is public so feel free to share it.
Fable Max built my family’s vacation
We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do.
Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing!
And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here.
Rebuilding the ThursdAI openers with HyperFrames
The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in.
The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough.
The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund.
Romain Huet on Codex’s inflection point and the golden age of AI engineering (X, Codex)
I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger.
The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading.
We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like.
On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent.
Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering.
He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it.
Wolfram and I on the Insecure Agents podcast (X, Pod)
The second half of the episode flips the format: Allie Howe, friend of the pod and host of the Insecure Agents podcast, interviewed Wolfram and me at the conference. I have not done many interviews from the guest chair, so this was a treat, and Allie asked sharper questions than we usually get.
We talked about what “AI Evangelist” actually means as a job title. For both of us, the mission is dispelling doomerism, which mostly means explaining the technology simply enough that people stop fearing what they don’t understand. Wolfram’s version of this is talking to the stewardess on his flight and his Uber driver about AI, not just developers.
Wolfram went deep on Wolfbench (wolfbench.ai), his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave (hi friends 🐝). Transparency changes what benchmarks mean: Fable didn’t take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files). You only learn that by reading traces, not averages. Same with Gemini 3.5 Flash placing high while quietly burning far more tokens than the model above it. And yes, when Wolfram added a cost column, Fable blew the chart, and I had to go have a conversation with our budget.
Then Allie got us onto loops and token economics, while I fidgeted with my Token Billionaire gold card from the conference (Wolfram has one too). My honest answer on when loops are worth it: the people pushing hardest (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but this technology disseminates the way agents did, from people who can afford it to everyone, as costs drop. And with the newest models I genuinely have not found the point where a long-running loop stops being productive; the category change is that they’ve gotten really good at not getting stuck.
The spiciest part was my hot take, delivered directly into the camera for CoreWeave IT: I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low. Pliny, the jailbreaker in chief, got five attempts at Matthew Berman’s OpenClaw live and couldn’t break it. Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, and it wouldn’t. Supply chain attacks are a different story, and that one scares me for humans and agents alike. Open source models, also a different story. But the “one poisoned email ruins your life” framing is behind us, and we should update.
Allie pushed back with the DeepMind “AI Agent Traps” paper on cognitive bias attacks, where repeated claims across sources tilt an agent’s judgment, and my non-answer answer is that this is a humanity problem older than AI: we haven’t solved it for politicians or media either, and it’s unfair to hold a new technology to an ethics bar we’ve never cleared ourselves. Wolfram’s electricity analogy is the one I keep reusing: AI is not a weapon, it’s electricity. Teach people to use it, don’t hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse.
We closed on the hallway track (the real reason to attend AI Engineer), why you should submit a talk even if you’ve never spoken before, and a small announcement I let slip: I’m actively working on bringing an AI Engineer event to Tel Aviv with some friends. More on that soon.
Wrapping up
That’s the episode: one riff on using AI like you mean it, one conversation with the person shaping how developers experience OpenAI, and one podcast where Wolfram and I had to answer the hard questions for a change. Huge thank you to Romain for the time in the middle of a packed conference, and to Allie for having us on!
I’ll be back live next week, tanned, rested, and hopelessly behind on AI news for the first time in three and a half years. Be gentle with me. If you missed any of it, ThursdAI is a podcast, a newsletter, and a YouTube show. Subscribe to one, then go check out the others.
* Hosts and Guests
* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
* Romain Huet - Head of Developer Experience, OpenAI (@romainhuet)
* Allie Howe - Host, Insecure Agents podcast (@vtahowe, Pod)
* Wolfram Ravenwolf - AI Evangelist, Weights & Biases & CoreWeave (@WolframRvnwlf)
* TL;DR and show notes from Live Show
* Hosts and Guests
* Co-Hosts – @petergostev, @nisten, @ldjconfirmed, @yampeleg
* 🏢 Big CO LLMs + APIs
* An OpenAI model escaped its isolated cyber evaluation, chained zero-days, reached Hugging Face production, and searched for benchmark answers (OpenAI, sama)
* Google launched Gemini 3.6 Flash, the cheaper 3.5 Flash-Lite, and the defensive-cybersecurity-focused 3.5 Flash Cyber (Google)
* Alibaba previewed the 2.4T-parameter Qwen3.8-Max in Qwen Chat and Studio; API access and open weights were not yet available (X, Try it)
* Microsoft launched MAI-Image-2.5-Pro and the faster, cheaper MAI-Voice-2-Flash during the show (Image, Voice)
* 🔓 Open Source LLMs
* Moonshot launched Kimi K3: a 2.8T-parameter, 1M-context, native-multimodal model with strong early coding, design, spreadsheet, and agentic results (Announcement, Blog)
* Poolside released Laguna S 2.1, a 118B/8B-active coding MoE with 1M context and downloadable quantized variants; strong specs, rough live demo (Blog, HF)
* Motif 3 Beta is a Korean 314B/13B-active MoE with 256K context; the weights are downloadable, but the current license is research-only and non-commercial (HF)
* NVIDIA released Nemotron 3 Embed for multilingual text/code retrieval and the 4B Cosmos 3 Edge omnimodal world model for physical AI (Nemotron, Cosmos)
* 🧠 AI Research & Capabilities
* Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit three-dimensional counterexample to the 87-year-old Jacobian Conjecture (Announcement, Terence Tao)
* Small local models running on older consumer GPUs are becoming useful for narrow business workflows such as medical-record parsing, accounting, email, bills, and inventory when paired with tools and deterministic verification
* Arcee and the US Department of Energy announced Genesis-Science-1, a planned trillion-parameter-class open-weight science model; Microsoft also committed $60M to the Genesis Mission through SPARK, while NSF announced $83M for AI-ready scientific data infrastructure (Arcee, Microsoft, NSF)
* 🤖 AI Coding & Agents
* Cursor launched a production-traffic-trained model router with Intelligence, Balance, and Cost modes; Cursor says Auto Intelligence approached Fable satisfaction at roughly 60% lower cost (Blog)
* 🎵🎬 Voice, Vision & Robotics
* Black Forest Labs introduced FLUX.3, an early-access multimodal model spanning image, video, audio, and action, plus FLUX.3 Mimic for robotics and action prediction (FLUX.3, Mimic)
* 🖥️ AI Infrastructure
* AMD and Anthropic announced up to 2 GW of MI450/Helios capacity, up to $5B in AMD strategic equity, and a Claude-assisted effort to improve ROCm (AMD)
* AMD launched Helios, MI400-series GPUs, 6th Gen EPYC, ROCm.ai, and Kria robotics products at Advancing AI 2026 (AMD)
* OpenAI announced Project Camellia, a roughly $20B Georgia data-center campus with 3.2 GW of contracted power arriving in phases from 2028–2032 (OpenAI)
* Alphabet raised its 2026 capex guidance to $195B–$205B after Google Cloud grew 82% year over year (Google)
* Meta and Anthropic are reportedly discussing a compute lease worth up to $10B over two years; the negotiations remain preliminary (Bloomberg)
* CoreWeave’s first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar user interactivity (CoreWeave)
* This week’s special episode
* Alex’s riff: papercut-fixing with Codex computer use, Fable Max trip planning (kids’ newspaper, packing list app), rebuilding ThursdAI openers with HeyGen HyperFrames + Google Omni
* Romain Huet interview from the AI Engineer World’s Fair floor: Codex app at 5M+ weekly users five months post-launch, /goal, AppShots, Codex managing its own threads, GPT 5.6 Sol/Terra/Luna, value maxing, 5.6 Sol at 750 tok/s on Cerebras, voice + reasoning as the next interface (X)
* Insecure Agents crossover with Allie Howe: AI evangelism vs doomerism, Wolfbench transparent evals on Weave (Fable refused 13 security-adjacent tasks), token billionaires and loop economics, the hot take that prompt injection is mostly solved at the frontier, deepfakes and open access, AI Engineer Tel Aviv teaser (X, Pod, Wolfbench)
ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribeThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M
2026/07/17 | 2h 13 mins.Hey yall, Alex here,
Huge thanks to Wolfram for running point on the live show this week. Didn’t have tons of time to edit this one, so please skip the first 10 minutes, it’s a loop of our new “wait for the live show to start” vid, that I build with HyperFrames and can’t wait to tell you about, next week!
Today it seems that OpenSource is biting back, with Kimi K3 getting released just a short while after Thinking Machines (Thinky) has released Inkling, their near 1T model.
I’m attaching the TL;DR and timestamps for the full show (my AI agents, yes even Fable and Sol are not a match yet at editing down hehe) and I’ll spare you the long Fable recap (please do let me know in the comments if you were expecting it)
0:00 – Intro, Alex on vacation, TLDR overview11:35 – TLDR: Thinking Machines, open source, OpenAI news12:34 – Banter: impressions of Sol/Codex, over-verification behavior37:22 – TLDR restart & detailed breakdown48:40 – Open Source AI section begins (Bonsai/Prism ML, Kimi K3)58:42 – Inkling (Thinking Machines) deep dive & 3D model visualization1:10:33 – Kimi K3 discussion & demo comparisons1:27:02 – Frontier Labs: AGI governance framework discussion (Demis Hassabis essay)1:47:04 – Grok Build CLI data leak & OpenAI file deletion incident2:02:15 – This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna2:09:52 – Closing remarks & sign-off
The one-minute version: Mira Murati's Thinking Machines released Inkling, a 975B parameter open-weights MoE under Apache 2.0, the top US open-weights model right now. Moonshot's Kimi K3 went from rumor to released API during the show, confirmed at 2.8 trillion parameters with open weights promised within days, and it's already topping early arena boards. PrismML's Bonsai 27B squeezes a full 27B model into 3.9 gigabytes so it runs on a phone. Codex and ChatGPT Work blew past 9 million users, OpenAI confirmed and explained the Sol file-deletion bug (back up your machines, folks), and xAI's Grok Build CLI got caught uploading entire private repos before open-sourcing the whole thing in response. Plus Wolfram's fresh Wolfbench numbers on the GPT-5.6 family in This Week's Buzz 🐝, where Sol on max thinking came out both cheaper and better than GPT-5.5's best.
ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
TL;DR and show notes
* Hosts and Guests
* Wolfram Ravenwolf, guest host this week (@WolframRvnwlf), while Alex Volkov (@altryne) is on vacation
* Co-hosts: @yampeleg, @nisten, @ldjconfirmed, @petergostev
* Open Source LLMs
* Thinking Machines releases Inkling - 975B total / 41B active MoE, trained from scratch on 45T multimodal tokens, Apache 2.0, top US open-weights model at 41 on the Artificial Analysis Index, encoder-free text/image/audio, Inkling-Small (276B/12B) previewed (X, Blog, HF)
* PrismML Bonsai 27B - 1-bit (3.9GB, ~90% retention) and ternary (5.9GB, ~95% retention) versions of Qwen 3.6 27B, multimodal, 262K context, Apache 2.0; Nisten demoed it live on a phone and a 6GB 1660 Ti (X, Blog, HF)
* MOSS-VL-Realtime - open source 11B VLM for real-time streaming video with proactive speaking and proactive silence, SOTA on all three open proactivity benchmarks, ~22.7GB, base model included (X, HF, GitHub, Arxiv)
* Kimi K3 API drops mid-show - confirmed 2.8T parameters, ~60-75B active (LDJ’s estimate), attention residuals, native vision, 1M context, ~half the price of Opus 4.8 / GPT-5.6 Sol, open weights promised within days; post-show, Arena reports K3 debuting #1 on Frontend Code Arena above Fable 5 (early results, caveats apply) (X, Arena)
* Big CO LLMs + APIs
* Codex + ChatGPT Work unified app hits 9M active users, up from ~6M days earlier and 1M in February; 5-hour windows replaced with banked, expiring resets (X)
* OpenAI confirms GPT-5.6 Sol file-deletion bug: $HOME override in full-access mode without sandbox or auto-review can nuke real home directories; mitigations and post-mortem promised (X, Techzine)
* OpenAI ships first hardware, the $230 kbd-1.0-codex-micro Codex controller with a reasoning-effort dial, built with Work Louder; sold out (X, Work Louder)
* GPT-Red - OpenAI’s internal automated red-teamer finds prompt injections at 84% vs 13% for humans, makes Sol 6x more injection-resilient, discovers the fake chain-of-thought attack class (X, Blog)
* ChatGPT returns to WhatsApp in the EEA after an EU antitrust order forces Meta to reopen to third-party AI bots; Kakao and Viber rollouts too (X)
* Google patches the Gemma 4 family - Flash Attention 4 (25-70% prefill speedup), tool calling fixes, reduced laziness, configurable vision resolution; criticized for shipping new weights with no version bump (X, HF)
* xAI’s Grok Build CLI caught silently uploading full private repos (history, deleted files, secrets) to Google Cloud Storage despite opt-outs; xAI deletes data, disables retention, and open-sources the CLI under Apache 2.0 (X, xAI response, GitHub)
* Demis Hassabis publishes an AGI governance essay proposing a FINRA-style Frontier AI Standards Body; endorsed by Altman, Nadella, Pichai, and Suleyman; the panel debates it hard on the show (X, Essay)
* This Week’s Buzz
* Wolfbench adds GPT-5.6 Sol, Terra, and Luna on Terminal Bench 2.0 at CoreWeave: Sol max-thinking is cheaper ($365/5 runs) and better than GPT-5.5 extra-high ($497), 85% average, 97% of tasks solved at least once; all traces on Weights & Biases, fully open source (wolfbench.ai)
* Show and tell
* Peter Gostev’s DOOMQL - a playable Doom-like built by GPT-5.6 Sol Ultra entirely in ~2,000 lines of SQL, essentially one shot; plus a Minecraft clone in Lean (X, GitHub)
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
More News podcasts
Trending News podcasts
About ThursdAI - The top AI news from the past week
Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week.
Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news
Podcast websiteListen to ThursdAI - The top AI news from the past week, Piers Morgan Uncensored and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


ThursdAI - The top AI news from the past week
Scan code,
download the app,
start listening.
download the app,
start listening.




























