← Back to today

Wednesday, August 26, 2026

5 stories · 4 min read

Two stories today are pulling in opposite directions. Box CEO Aaron Levie is arguing that enterprise data infrastructure matters more than ever now that agents are doing the work. Meanwhile, Replit CEO Amjad Masad just publicly said he replaced one AI tool with another in his own daily workflow. One says the pipes matter most. The other says the agent on top of the pipes is already winning. Both are probably right, which makes the next 18 months genuinely hard to predict.

01

Your company's data governance problem just got 100 times more expensive

Box CEO Aaron Levie posted that systems of record, the databases and platforms where companies actually store their business data, have never mattered more than right now. His argument: AI agents will execute hundreds of tasks on these systems for every one task a human used to do. The OpenAI-Hugging Face incident he references is a preview of what happens when agents hit data they shouldn't touch, or do things at machine speed without the guardrails humans naturally apply by slowing down. ---

Why it matters: If your company's access controls and data governance were "good enough" when humans were clicking through screens, they are not good enough when an agent is making thousands of API calls a night. Whatever technical debt you have in permissions, audit logs, and business logic, agents will find every crack in it, faster and at larger scale than any employee ever could.

Source →

02

Replit CEO says Replit Agent replaced Claude in his daily work

Amjad Masad, CEO of Replit, posted that Replit Agent has fully taken over from Claude CoWork in his day-to-day workflow. His reason: it's more persistent, more thorough, and uses code more effectively to complete tasks. ---

Why it matters: When the CEO of a company that competes partly with Anthropic says he switched away from Claude for his own daily work, that's not a press release. It's a real product signal worth watching, especially for developers choosing between general-purpose AI assistants and task-specific agents built around a coding environment.

Source →

03

OpenAI teases something big at DevDay 2026

Thibault Sottiaux, who works on Codex at OpenAI, posted a bold preview: "OpenAI DevDay 2026 will be our best DevDay in the history of the company. It will not be close." That's it. No details. ---

Why it matters: Yesterday's digest covered Sottiaux announcing fixes to Codex's quota accounting bugs. Today he's hyping DevDay. The combination suggests OpenAI is planning something substantial for developers, not just a model update. If you're building on their APIs, it's worth clearing your calendar.

Source →

04

Evals series continues: your benchmarks might all be grading on the wrong curve

Madhu Guru posted part 8 of a running series on eval construction, focused on what he calls "discriminatory power." The problem: if you run five AI systems through your eval and they all score between 92 and 95, your eval is useless, even if you know for other reasons that some systems are dramatically better than others. It's like giving a fifth-grade math test to PhDs. Everyone passes and you've learned nothing. ---

Why it matters: This connects directly to yesterday's thread from the same series. Bad evals don't just fail to catch problems, they actively mislead you by creating false confidence. If your team's benchmark shows all your candidate models clustering near the top, the benchmark probably needs to be harder, not the models better.

Source →

05

Peter Yang on having too many AI chiefs of staff

Product builder Peter Yang posted that he now has three AI "chief of staff" agents running and is apparently considering adding a fourth. This one is thin on substance. Either this is early-adopter humor about agent sprawl, or Yang is genuinely stress-testing what it looks like to manage a small fleet of personal AI agents. No real takeaway until he writes up what's actually working.

Source →