Aaron Levie has now posted two days in a row about the gap between AI models and real enterprise workflows. Either he's working through something publicly, or Box is about to announce a product. Either way, he's not wrong, and the rest of today's digest is full of companies quietly trying to fill exactly that gap.
01
OpenAI rolls out a team plan for small companies that actually makes sense
Thibault Sottiaux, who works on Codex at OpenAI, announced what sounds like a long-requested middle tier: a plan designed for teams and small companies that mirrors the $100 Pro plan but adds centralized billing, SSO, usage analytics, and no 5-hour limits. It connects to Google Workspace, Slack, GitHub, and Microsoft 365 out of the box. ---
Why it matters: The gap between "one person on Pro" and "enterprise contract with a sales call" has been a real barrier for small teams. A 10-person startup that wants Codex for their whole engineering team could not do that cleanly before. The 569 replies on this post suggest a lot of pent-up demand. If you manage a small team and have been sharing one login, this is the thing you've been waiting for.
Boris Cherny, who works on Claude at Anthropic, posted that memory is "now simpler and more powerful," with a link to what appears to be an updated feature. No blog post, no press release. Just a short note and 1,400 likes. ---
Why it matters: Memory is one of the most frustrating limitations in AI assistants today. Every time you start a new conversation, you're re-explaining who you are and what you're working on. If Anthropic actually simplified how Claude retains context across sessions, that's a bigger quality-of-life fix than most of the headline model updates from the past month. Worth testing if you use Claude regularly for ongoing projects.
Box CEO: the real AI opportunity is in the gap between models and workflows
Aaron Levie posted for the second day running on the same theme: there's a wide gap between raw AI models and the actual workflows enterprises run on. His framing is that the premium will go to companies that can convert "raw tokens into real world outcomes," not just those that expose model APIs. He's responding to a post about applied AI strategy at scale. ---
Why it matters: Yesterday Levie was arguing that data governance is the critical layer. Today he's arguing the integration layer is where the value accrues. Both point to the same conclusion for anyone building AI tools for business: selling access to a model is a commodity play, and the only defensible position is deep integration into how work actually gets done. If your AI product is one API call away from being replaced, this thread is about you.
Madhu Guru posted part 9 of his ongoing series on building better evals, focused on what he calls "the eval roadmap problem." The core argument: most teams build their benchmarks once and then leave them static, while their users keep asking harder questions. He uses a financial research agent as the example. Week one, users ask it to summarize a five-page earnings report. Two months later, they're asking it to compare five years of earnings across competitors. The same eval that passed the agent in week one tells you nothing about whether it's good enough in month two. ---
Why it matters: Yesterday's installment covered discriminatory power, the problem of evals where every model scores between 92 and 95 and you learn nothing. This one is the logical follow-on: even if your eval was well-designed at launch, users evolve and evals don't. If your team shipped an AI feature six months ago and hasn't updated the benchmark since, you are flying blind on whether it still works for what your users are actually doing.
Peter Yang built a free AI tool for cancer patients and caregivers
Product builder Peter Yang shared a skill called /fuck-cancer that creates and maintains a single structured document with patient information, next steps, medical term definitions, and an update log. He's posting it as a free resource. This is genuinely useful for anyone navigating a cancer diagnosis in their family. It's a specific, practical application of AI as a personal knowledge manager in a situation where people are overwhelmed with information and decisions. The link to the free skill is in the post.