Yesterday we asked whether the real money in AI was in the product or the infrastructure underneath it. Today, Box CEO Aaron Levie answers that question with conviction, and a Chinese open model quietly crashes a party that GPT-5.5 thought it owned.
01
The "thin wrapper" critique of AI apps is starting to look wrong
Box CEO Aaron Levie posted a detailed breakdown of what the applied AI layer actually looks like in practice, and his conclusion cuts against the dismissive "it's just an API call" take that's been circulating since 2023. Driving real agentic workflows inside an enterprise, he argues, requires solving orchestration, context management, error recovery, domain-specific knowledge, and trust, each of which compounds in complexity. That complexity is where durable value forms, not in the model itself. ---
Why it matters: If Levie's read is right, the companies building the scaffolding around models, not the models themselves, will capture most of the enterprise margin. That has direct consequences for how your organization should be thinking about vendor lock-in. The moat isn't in which LLM you're calling. It's in whoever owns the workflow layer sitting on top of it.
A Chinese open model just beat GPT-5.5 on a real-world benchmark
The Latent Space newsletter ran a detailed look at GLM-5.2, the latest model from Z.ai, and the verdict is striking: multiple independent testers, including researcher Jeremy Howard, are calling it a genuine frontier model that happens to be open. Artificial Analysis' new knowledge work benchmark now ranks it above GPT-5.5. The /r/LocalLlama community, typically a harsh audience for models that benchmaxx and disappear, is also giving it serious marks. The newsletter also flags Z.ai's forecast that an open model at the level of Claude Fable is coming by December. ---
Why it matters: This isn't a benchmark story, it's a structural shift story. When an open Chinese model outperforms GPT-5.5 on knowledge work, the business case for paying OpenAI API prices on high-volume tasks gets harder to defend. If you're building on a paid frontier model for anything repetitive and high-throughput, the cost math is about to change on you.
Guillermo Rauch, CEO of Vercel, made the case that the Vercel AI SDK now plays the same role for AI agent development that Next.js played for web apps: a practical harness that makes a chaotic underlying technology actually buildable. His prompt was GLM-5.2 surpassing Claude Opus 4.8 on Vercel's own Next.js evals, which he used to illustrate why model-agnostic tooling matters more than ever when the model leaderboard is reshuffling weekly. ---
Why it matters: If you're building an AI product tied tightly to one model provider's SDK, you're one benchmark shuffle away from a painful migration. The teams that abstracted their model calls behind a framework like the AI SDK can swap in GLM-5.2 tomorrow. The ones that didn't are going to spend a sprint on plumbing instead of product.
Thibault Sottiaux from the Codex team at OpenAI announced a "sneaky double reset": users got a full usage reset plus an extra reset banked for future use. The post has over 5,500 likes, which is a reasonable proxy for how many people were hitting their Codex limits. ---
Josh Woodward shares a Google Labs collaboration moment
Google Labs VP Josh Woodward posted a brief photo with the Voltage team, citing co-creation as a core Labs value. The post doesn't disclose what they're building together.