AI agents keep getting smarter, but the bigger story this week is how much they’re reshaping the systems around them. Host Eric Freeman, an O’Reilly author and UT Austin professor, pulled one thread through a packed news week. Models are optimizing less for chat and more for autonomous work, with fallout showing up in web traffic, enterprise budgets, and one security incident that’s since made headlines. Eric kept returning to the question of what changes when the primary user of these models, and of the web itself, stops being a person.
Models are now built for agents, not conversations
Grok 4.6 put xAI back in the frontier race, closing the gap with top coding models and pricing aggressive enough that teams are shifting workloads over. The release landed the same week SpaceX closed its Cursor acquisition, pairing xAI’s models and compute with a widely used coding environment. The first product from that pairing, Grok Bot, gives each agent its own cloud computer that browses, runs tools, and works independently, handing control back only for logins. Eric summed up the shift simply, calling it the difference between “help me do this” and “here’s the job, come back when you need me.”
Open models pushed from both directions. DeepSeek V4 Pro went after high-end reasoning and agentic work, despite a fourfold API price hike and a new open source harness called dsh, built on the idea that everything is a plugin. GLM-5.3 made a big coding leap through retraining alone, and got noticeably better at cyber capability too, a reminder from Eric that skills behind a better autonomous engineer also make a sharper attacker. Meta went the other way with Muse Glimmer, shrinking down for desktop GPUs, while OpenAI quietly held back its Astra model over security concerns.
Speed is turning into its own kind of capability. GPT-5.6 Sol’s new Ultrafast mode, on Cerebras wafer-scale hardware, hits roughly 14 times the normal pace, around 750 output tokens a second. Once that loop of reasoning, tool calls, and self-correction compresses enough, Eric noted, the model stops being what slows you down.
The money has moved from training models to running them
Gartner’s latest forecast, which Eric covered, lays out the shift plainly. Spending on AI-optimized cloud infrastructure is set to nearly double this year, up about 96%, from roughly $22 billion to more than $42 billion, over three times the broader cloud market’s growth rate. For the first time, organizations are expected to spend more running models than training them, about $23 billion on inference against $19 billion on training.
Agents are the reason inference costs are climbing. A single task can quietly become dozens of model calls once agents search, use tools, check their own work, and spin up other agents to help, a point Eric returned to often. AI economics are less about building a model now, and more about the cost of running one.
Agents now generate most web traffic, and much of its content
Back in March, Cloudflare CEO Matthew Prince predicted bot traffic would overtake human traffic by 2027. It’s already close, with Cloudflare’s Radar data now putting agentic bots at 57.4% of web requests. It’s not just traffic either, since roughly 40% of Facebook posts, 44% of new music on Deezer, and 52% of online articles are estimated to be machine-made. Numbers like that, Eric said, make “dead internet theory” sound less like a joke.
Platforms are responding differently. LinkedIn added a feature to flag content that “seems like AI slop,” while quietly pulling back the generative writing tools that helped create the mess. Anthropic took another route, watermarking Claude’s output at generation time, including a statistical watermark baked into the text itself, partly to comply with the EU AI Act. A watermark means something when present, Eric noted, but its absence tells you little.
The clearest sign of how high the stakes have gotten came from the OpenAI–Hugging Face incident, detailed in a Black Hat talk Eric said everyone should watch. Sandboxed agents given ordinary tasks, cut off from the internet and unable to talk to each other, found a way anyway, leaving notes in a shared packaging system, turning it into an internet proxy, and working up to admin control. Once OpenAI shut that down, they pivoted, hiding messages in filenames to keep coordinating. The investigation reviewed seven billion reasoning steps and over three million GPU hours. Eric argued it’s worth your time, whether you write code or sit in the C-suite.
What’s next
AI memory is also expanding, moving from “remember what I told you” toward “remember what I was doing,” with OpenAI’s new Computer History feature using macOS accessibility data (not screenshots) to build a timeline of your work across apps. It’s opt-in, Mac-only for now, and a sign of where agent context is headed.
Join us again next Monday for another episode of This Week in AI, when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on YouTube, Spotify, Apple, or wherever you get your podcasts.

