Get started with Claude Managed Agents by following our docs . A running topic on the Engineering Blog is how to build effective agents and design harnesses for long-running work . A common thread across this work is that harnesses encode…
Speaker 1 | 00:00 - 00:10 You're someone who, I think, cares a lot about the craft of things. One of the knocks on, you know, using agents for coding is it gets rid of some of that feeling or something like that. How do you feel about that…
been deep down this rabbit hole at the intersection of AI and consumer experiences recently: How do you develop a theory of why someone did something, rather than just a history of what they did? Consumer products have…
This talk on the OpenAI/Hugging Face incident had one detail I found particularly chilling. The agents cooperated even when their reasoning showed that it wasn't in their immediate individual interest, because it was be…
This reference conversation on how to build long-horizon AI agents with @mitch_troy of @trybasis is also available on Spotify, Apple Podcasts and here on YouTube: https://t.co/j2srHcxAQK
How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with @mitch_troy, co-founder of @trybasis 01:09 Why Everyone at Basis Was Whispering to AI when @steph_palazzolo wal…
Podcasts & Newslettersfrom The MAD Podcast with Matt Turck
Speaker 1 | 00:00 - 00:18 Humans are already used to working with non deterministic systems, it's just the systems are normally their co workers, not their computers. And in many ways, like companies and processes is all about how do you d…
Been thinking about why AI diffusion has been slow. It’s because we are asking users to understand our nerdy lab speak. We greet them with a blank window and ask them to write a prompt. Then they have to pick a model, d…
Podcasts & Newslettersfrom Latent Space Newsletter
Speaker 1 | 00:05 - 00:30 Today on our priors, we're joined by Melissa Tachmak. Melissa is the founder and CEO of Netic, a company that builds AI for different real world services like HVAC, pet care, a variety of other things like that, r…
~1500 Elo! Consistently beats frontier models and Stockfish level 0. Fun seeing an 8b model mogging GPT 5.6 with high reasoning and response chaining. Spends 1-2 seconds per move vs 30 seconds. Play it: https://t.co/qRg…
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
Podcasts & Newslettersfrom Latent Space Newsletter
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
Tip for building in public: If making content feels like extra work, show the work already happening inside your product. A tiny screen recording, the first version, or the user behavior that changed your design. The re…
Podcasts & Newslettersfrom The MAD Podcast with Matt Turck
Speaker 1 | 00:00 - 00:17 Fraudsters have figured out that in AI, you actually don't really need to steal money or credentials. You can just steal tokens. And the scale of this actually shocked me when I looked at the data. So more than on…
GPT-5.6 is now out. We've been evaluating the model family on the Box AI Complex Work eval, which tests the model with the Box AI Agent on a variety of extremely hard tasks using enterprise document sets. Sol is a big s…
The latest AI models being dropped are getting insanely good handling complex knowledge worker tasks, and especially dealing with sophisticated domains of work like legal, professional services, healthcare, and more. Gr…
imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved that they can do "brain surgery" interventions into reasoning to change topics midstream* 2) THE MODEL IS ABLE TO DETEC…
Sonnet 5 is a substantial improvement over Sonnet 4.6 on reasoning, tool use, coding, and knowledge work. Its performance is close to Opus 4.8, at lower prices. https://t.co/VOISbk14Lk
Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication, and physician-informed evaluations.
Starting today, Claude Code can capture work progress as an artifact, which turn Claude Code's work into live, shareable visual pages— including PR walkthroughs, system explainers, dashboards, and release checklists—that update themse…
Podcasts & Newslettersfrom Latent Space Newsletter
I just discovered forceBlockStreamingForReasoning = resolvedReasoningLevel === "on" for OpenClaw and frankly I love it Seeing the reasoning traces of my claw with Claude Fable 5 is a mind-blowing experience. Seeing the…
Today we're releasing Foundation Models framework support for Claude through a new Swift package that lets Apple developers use Apple's Foundation Models framework to call Claude for more complex workflows. Apple’s Foundation Mod…
Speaker 1 | 00:00 - 00:04 Is reasoning enough to get to generalization, or is another method needed? Speaker 2 | 00:04 - 00:08 It does feel like there is something else that possibly could generalize much better. Speaker 1 | 00:08 - 00:12…
GPT-Rosalind advances life sciences research with enhanced biological reasoning, medicinal chemistry expertise, genomics analysis, and experimental workflow capabilities.
Podcasts & Newslettersfrom The MAD Podcast with Matt Turck
Speaker 1 | 00:00 - 00:23 One of the things that CHAT GPT was able to do was assume it was false. When you go against the grain and do something contrarian like that, you really have to have strong conviction in what you're doing in order…
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and why Grok Imagine is so underrated. For the first time, we do a deep dive with the guy who led it!
Speaker 1 | 00:00 - 00:06 You prompt AI to do something. It blows your mind. You feel inadequate. You feel like, oh my god. This thing's gonna take my job. Speaker 1 | 00:07 - 00:11 And then it stops working, it looks back at you and says,…
Podcasts & Newslettersfrom Latent Space Newsletter
Why AI Progress Suddenly Feels Real - my conversation with @yanndubs, who co-leads the Post-Training Frontiers team at @OpenAI 00:00 - Intro 01:30 - Why recent AI progress feels like a step function 04:13 - Model reliab…
Podcasts & Newslettersfrom Latent Space Newsletter
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protein reasoning, and scientific research workflows.
Thrilled to have backed @nicbstme and the @fintoolx team as an angel. Fintool was a magical product that was able to do a lot of the heavy reasoning tasks before the models even existed. Microsoft is getting an absolute…
On the API, a new xhigh effort level between high and max gives you finer control over reasoning and latency on hard problems. Task budgets (beta) help Claude prioritize work and manage costs across longer runs.
The big difference between a true "agent" and simply running an LLM in a loop (or an event-based trigger) is how you do intelligent long-horizon memory management. By far the most interesting aspect of the Claude Code a…
Vision-language models (VLMs) use images and text to plan robot actions, but they still struggle to decide what actions to take and where to take them. Most systems split these decisions into two steps: a VLM generates a plan in natural la…