Speed & Cost
50 items tagged with this topic
Recent
An update on recent Claude Code quality reports
Over the past month, we’ve been looking into reports that Claude’s responses have worsened for some users. We’ve traced these reports to three separate changes that affected Claude Code, the Claude Agent SDK, and Claude Cowork. The API was…
[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law
Every lab CEO is on X now
Older
the models have no moat (OpenAI, Anthropic, XAI) the IDEs have no moat (Cursor, Windsurf) the h…
the models have no moat (OpenAI, Anthropic, XAI) the IDEs have no moat (Cursor, Windsurf) the harnesses have no moat (Cognition, Factory, LangChain) the app builders have no moat (Replit, Lovable, Bolt) the wrappers hav…
Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Want to be able to deal with RSI? Here are 23 actionable…
[AINews] How to steal a Reasoning Trace
Speculative Decoding by any other name would distil as sweet
🤖Android users love Gemini + Gemini can automate actions across 40+ popular apps + Great for b…
🤖Android users love Gemini + Gemini can automate actions across 40+ popular apps + Great for booking rides, reserving tables, and more + More updates coming tomorrow at @madebygoogle!
[AINews] Megakernels are so dead and so back
A quiet day lets us highlight a Cursor launch and an engineering debate
Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Self-sustaining and self-replicating AI viruses are here:…
a very primitive form of the near term multiagent agi future is setting up one thread to ping b…
a very primitive form of the near term multiagent agi future is setting up one thread to ping back once its done so you create an implicit kanban/waterfall graph of dependent threads but each preserving their own work a…
Autoscaling endpoints for LLM inference
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.
The best model for early product validation is almost never the best model for production. Talk…
The best model for early product validation is almost never the best model for production. Talked to a bunch of founders this weekend. There’s a standard playbook emerging : 1/ Prototype with the best frontier models. M…
How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
EvoLib: Turning experience into evolving knowledge
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience…
Another day another near frontier open weights model release. If you had gone back even 3-6 mon…
Another day another near frontier open weights model release. If you had gone back even 3-6 months and given everyone access to what we’re now seeing in open weights even as a closed model, their minds would be complete…
Configuring Dedicated Model Inference
The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.
[AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization
Distillation is all you need!
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Epoch and METR release MirrorCode, a benchmark for seeing…
Ep 92: xAI Co-Founder Unpacks the Future of Model Development
Speaker 1 | 00:00 - 00:16 Welcome back to unsupervised learning. I'm Jacob Efron. We had an awesome episode today with Igor Babushkin. Igor has been at all the right places at all the right times, leading a lot of really interesting AI wor…
verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prio…
verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prioritize pretrain data quality enough that commoncrawl isn't good enough for you, you have to build a Whole Web scraper anyway…
Thought provoking post by Dwarkesh. In general - as AI gets more powerful - we should expect on…
Thought provoking post by Dwarkesh. In general - as AI gets more powerful - we should expect on the margin that inference goes toward the most economically useful task, and everything else gets priced out. This in theor…
How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Serving large models is hard. https://t.co/nCbfTx9lp3
Serving large models is hard. https://t.co/nCbfTx9lp3
The production platform for open-weight AI inference
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
1) Competition is good for the ecosystem. 2) Serving models at scale is hard. Proud that @OpenA…
1) Competition is good for the ecosystem. 2) Serving models at scale is hard. Proud that @OpenAI signed the letter. Ant is silent on it. https://t.co/r9QVQ2R9JH
This reference conversation on fast inference, AI chips, and the next compute bottleneck with @…
This reference conversation on fast inference, AI chips, and the next compute bottleneck with @andrewdfeldman of @cerebras is also available on Spotify, Apple Podcasts and here on YouTube: https://t.co/iP21ZRHgpB
My conversation with @andrewdfeldman, CEO of @cerebras. We started from "what is a wafer?" and…
My conversation with @andrewdfeldman, CEO of @cerebras. We started from "what is a wafer?" and built up to why the entire chip industry is reorganizing around inference speed. 00:00 Cold open & Intro 01:31 Why speed bec…
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
Speaker 1 | 00:00 - 00:21 This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly, and therefore, you get answers in less time. For…
The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vuln…
The Claude Security plugin for Claude Code is now available in beta. Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Claude inference you a…
Today’s launches are all about better performance, lower latency, and a smaller bill. + 3.6 Fla…
Today’s launches are all about better performance, lower latency, and a smaller bill. + 3.6 Flash cuts token usage by up to 65% on complex coding + 3.5 Flash-Lite reaches speeds of 350 output tokens/sec Both are live in…
What does 99.9% uptime mean for inference?
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.
Four years post peak web3 and crypto tokenomics debates… turns out the tokenomics debate that m…
Four years post peak web3 and crypto tokenomics debates… turns out the tokenomics debate that matters is open vs closed weight, inference costs and model routing. https://t.co/THz3HEyx1d
The reason enterprises struggle to go beyond basic chat bots is the talent gap to build harness…
The reason enterprises struggle to go beyond basic chat bots is the talent gap to build harnesses and evals. 1. Evals: do you clearly understand your use cases and can you replicate that in the form of offline and onlin…
5.6 sol growth is insane. the inference team has done heroic work to be able to support demand.…
5.6 sol growth is insane. the inference team has done heroic work to be able to support demand. we are going to move mountains to continue to scale, but it is possible there are some hiccups soon.
Open, convenient and predictable: Introducing Provisioned Throughput
Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no infrastructure t…
Updates for Codex and ChatGPT Work users. No nerfing, only good stuff! - We have landed inferen…
Updates for Codex and ChatGPT Work users. No nerfing, only good stuff! - We have landed inference optimizations and are passing down savings to all the subscriptions for GPT-5.6 Sol. That should result in around 10% mor…
Make the model a cog in a machine you own. ◾ AI SDK → open model API ◾ https://t.co/O7y9dmUqk5…
Make the model a cog in a machine you own. ◾ AI SDK → open model API ◾ https://t.co/O7y9dmUqk5 → open Agent API ◾ AI Gateway → open ZDR inference Startups and enterprises must own their data, evals, model choices, softw…
Inside Zipline's Autonomous System: 140M Miles, Zero Incidents
Speaker 1 | 00:00 - 00:21 I remember being in Rwanda early days and going out and meeting with some of the doctors and lab techs that we were serving and asking for them, like, you know, how's it going? What what do you think? What's your…
If you’ve ever wondered why we will need 100X more AI inference in the future, and what it’s go…
If you’ve ever wondered why we will need 100X more AI inference in the future, and what it’s going to be driven by, this is another good example. Devin pushes forward an idea of agentic mapreduce, which means we’ll now…
AI is expensive to run partly because most workloads today run on generic hardware designed pre…
AI is expensive to run partly because most workloads today run on generic hardware designed pre-LLMs. Etched is the first system designed from the ground up for modern inference. https://t.co/KluYZIZATE
Inference runs on Azure infrastructure, operated by Anthropic. Prompt caching and extended thin…
Inference runs on Azure infrastructure, operated by Anthropic. Prompt caching and extended thinking are supported today, with more capabilities on the way. Read more: https://t.co/9MILqQYun4
OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.
[AINews] It's Meta-Harness Summer
Move over, Harness Engineering, it is time for the harness of harnesses!
An interesting way to take Noam at his word in regards to always keeping a constant inference b…
An interesting way to take Noam at his word in regards to always keeping a constant inference budget for any eval reporting - is that open models have a lot more dollar per token mileage than closed model APIs. So anyon…
Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
Speaker 1 | 00:00 - 00:13 With GBT three, you couldn't scale test time compute. Like, if you gave it a budget of $10,000,000 and said, okay. Well, let's see what GBT three can do. It really can't do that much. The precarious frameworks and…
Rough estimate on $ productivity lost by Fable 5 ban: $12M per hour Frontier AI-coding daily ac…
Rough estimate on $ productivity lost by Fable 5 ban: $12M per hour Frontier AI-coding daily actives, mid-2026: 5M devs Fully-loaded cost: $90/hr Work routed to Fable in 48 hours: 17.8% Fable is on average ~15% more pro…
One of the biggest questions in AI is how far behind open weights models remain from closed mod…
One of the biggest questions in AI is how far behind open weights models remain from closed models at any given time. There are huge differences in market structures depending on whether open weights models remain 3 or…
[AINews] Loopcraft: The Art of Stacking Loops
a quiet day lets us highlight a great concept from Peter Steinberger, Boris Cherny, and Andrej Karpathy