Coding
50 items tagged with this topic
Recent
Older
Stampli cuts launch hours by 68% using ChatGPT Work
With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.
Introducing Claude Opus 5 \ Anthropic
Opus 5 is a step change improvement for the Opus tier powering long-running agents while delivering improvements in coding and professional work.
Introducing Claude Sonnet 5 \ Anthropic
Our most agentic Sonnet yet, with top-tier intelligence for coding and everyday professional work.
Broadening access to Skala creates a faster path to predictive DFT
Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational per…
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.
DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and GPT-5.6 Sol. Sol leads pass@1 by 10 points at 35x the cost; Pro wins pass@4, and a Pro-first cascade hits 83.0%.
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.
A/B test models in production
Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.
Update on rate limits in Codex. We do see that for some users the cache hit rate has been worse…
Update on rate limits in Codex. We do see that for some users the cache hit rate has been worse this week than the stable state the weeks before. This could explain that usage is draining somewhat faster for those users…
The banked reset will be there by 8pm PST. For all paid users of ChatGPT Work and Codex. Do wit…
The banked reset will be there by 8pm PST. For all paid users of ChatGPT Work and Codex. Do with this information what you may.
https://t.co/G0UyDr4xgg now supports your @Grok & Codex subs Test it out on a sandbox. Instant…
https://t.co/G0UyDr4xgg now supports your @Grok & Codex subs Test it out on a sandbox. Instant to install: https://t.co/ogzI7TeKQ4 https://t.co/DzGu67GnJF
This is still SO underrated btw.. My daughter just started kindergarten. The daily meal is on a…
This is still SO underrated btw.. My daughter just started kindergarten. The daily meal is on a random website with weird unstructured data. Asked Claude Code to find the API through network requests, which turned out t…
Suggested patches open in Claude Code on the web, using the models your team already uses. Thes…
Suggested patches open in Claude Code on the web, using the models your team already uses. These updates are part of our work to give defenders greater access to Mythos results without requiring direct access to the mod…
Point Claude Security at a GitHub repo and Mythos scans for vulnerabilities tracing data across…
Point Claude Security at a GitHub repo and Mythos scans for vulnerabilities tracing data across files and reasoning about how components interact. Each finding comes back with a CWE category, confidence and severity rat…
We've investigated a few messages about codex usage limits being different. That's not somethin…
We've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users man…
I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate cros…
I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate crossing 100K subs on YouTube. Excited to share a lot more practical interviews soon to answer your most burning AI questions: 1…
One of the more underrated aspects of the new Free Mode is how fast it is. Making coding intera…
One of the more underrated aspects of the new Free Mode is how fast it is. Making coding interactive again! https://t.co/qwCBr2vBS2
An update on recent Claude Code quality reports
Over the past month, we’ve been looking into reports that Claude’s responses have worsened for some users. We’ve traced these reports to three separate changes that affected Claude Code, the Claude Agent SDK, and Claude Cowork. The API was…
Scaling Managed Agents: Decoupling the brain from the hands
Get started with Claude Managed Agents by following our docs . A running topic on the Engineering Blog is how to build effective agents and design harnesses for long-running work . A common thread across this work is that harnesses encode…
[AINews] Memory prices up 500% in 12 months
the Memory crunch continues - Moore’s Law reversed to 2007 levels
It has not been used yet, but would you look at that. Codex for scale. https://t.co/o1pulwoifd
It has not been used yet, but would you look at that. Codex for scale. https://t.co/o1pulwoifd
Agents made software cheaper but made coding expensive. Today, together with @OpenAI, we’re cha…
Agents made software cheaper but made coding expensive. Today, together with @OpenAI, we’re changing this: https://t.co/OI3eowMt5s
There tends to be a debate between being an expert or generalist in the era of AI. So far, the…
There tends to be a debate between being an expert or generalist in the era of AI. So far, the experts appear to have the upper hand, and that’s not slowing down. AI makes it 10X easier to get started with any kind of t…
Good details on the Stripe + OpenRouter deal here. For AI to diffuse more broadly, developers a…
Good details on the Stripe + OpenRouter deal here. For AI to diffuse more broadly, developers and enterprises will want ways of being able to mix and match intelligence from a variety of provers seamlessly and better ma…
[AINews] Stripe buys OpenRouter for $7B
No GPUs, no Agents, just really, really, really good infra and distribution.
Hi! Recapping some changes we have rolled out over the last couple of weeks that have further r…
Hi! Recapping some changes we have rolled out over the last couple of weeks that have further reduced the risk associated to potentially destructive actions being performed by Codex during its work. A few weeks ago, we…
2. Non-engineers are shipping more code PMs attaching pull requests rose from 3% to 10% in two…
2. Non-engineers are shipping more code PMs attaching pull requests rose from 3% to 10% in two years. Designers went from 1% to 8%, and founders are second only to engineers at 23%. I'm actually surprised that designers…
I’ve been using https://t.co/OL0LzGtvAw as my daily driver and it’s a one-way street. It’s 10-2…
I’ve been using https://t.co/OL0LzGtvAw as my daily driver and it’s a one-way street. It’s 10-20x smaller than the major coding CLIs. It starts up instantaneously. It feels more like using 𝚣𝚜𝚑 than an IDE in your ter…
Your software factory should be a monorepo. All your company context (design, marketing, sales,…
Your software factory should be a monorepo. All your company context (design, marketing, sales, engineering, support…) in one place for agents to build upon https://t.co/MRPrmkSAPd
I don’t know why anyone would learn Claude Code by reading a book, but apparently it’s a thing…
I don’t know why anyone would learn Claude Code by reading a book, but apparently it’s a thing in Japan https://t.co/NlOttUCdD4
How we contain Claude across products
Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service. Today that level of access is routine, and Anthropic developers are more productive for it. Th…
Partnering with CodeAI to prepare the first AI generation
OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue
Flue 2 takes its inspiration from React. Creator Fred Schott, of Astro fame, tells Latent Space why he added hooks and why agents are defined by their harnesses.
[AINews] Cursor's $60B acquisition by SpaceXai closes
Congrats to the team!
Gimme, gimme, gimme Codex after midnight Won’t somebody make these failing tests all go away? G…
Gimme, gimme, gimme Codex after midnight Won’t somebody make these failing tests all go away? Gimme, gimme, gimme Codex after midnight Ship it through the darkness by the start of the day
What is an obvious thing that we should do with Codex, API or our models that we should just do…
What is an obvious thing that we should do with Codex, API or our models that we should just do but haven't yet? What is 100% within reach, but we just seem to be missing?
What AI skills or tools are people using to edit YouTube talking head intros? Stuff like zoom i…
What AI skills or tools are people using to edit YouTube talking head intros? Stuff like zoom ins and animating in captions, logos, and b-roll. I've been testing @HyperFrames_ for this but curious if there are alternati…
the nice things about code is that it is easier to edit and nudge in the directions you want an…
the nice things about code is that it is easier to edit and nudge in the directions you want and export to work with existing tools
all of the recent proc gen art, video editing and 3d game demos recently have made me update to…
all of the recent proc gen art, video editing and 3d game demos recently have made me update towards LLM coding models being better at a lot of creative work than diffusion models
It’s not enough to scan your code for vulnerabilities; it’s important to try to break them with…
It’s not enough to scan your code for vulnerabilities; it’s important to try to break them with pen testing. https://t.co/1Hc1AT4jz0
You can now host your repos in Cursor Origin and deploy to Vercel via Cursor Origin which is it…
You can now host your repos in Cursor Origin and deploy to Vercel via Cursor Origin which is itself hosted on Vercel. And unlike GitHub, it's online 😁 https://t.co/ybWprI8gm4
What do you get? A private github repo with 70 of my proven skills and the beginnings of your K…
What do you get? A private github repo with 70 of my proven skills and the beginnings of your Karpathy-style knowledge wiki. Read the full docs in the README. All of this is MIT-licensed open source and free. https://t.…
It's free to try and works with your existing Claude Code or Codex subscription right now. Just…
It's free to try and works with your existing Claude Code or Codex subscription right now. Just make a new directory and start Codex or CC (on command line OR on Desktop, slightly better on Desktop) Just paste this imag…
Remote agents in Vibe. Powered by Mistral Medium 3.5. | Mistral AI
Introducing Mistral Medium 3.5, remote coding agents in Vibe, plus new Work mode in Le Chat for complex tasks.
China's Quantum Flywheel
Six months of Party mobilization
Codex ✅ Almost 100% reliable ✅ Occasional resets ✅ Open-source ✅ (will have Astra) https://t.co…
Codex ✅ Almost 100% reliable ✅ Occasional resets ✅ Open-source ✅ (will have Astra) https://t.co/DxNdAgpag5