← Back to home

Speed & Cost

50 items tagged with this topic

Recent

Older

Official Sourcesfrom Google DeepMind Blog

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

Watchlistfrom Anthropic Engineering

An update on recent Claude Code quality reports

Over the past month, we’ve been looking into reports that Claude’s responses have worsened for some users. We’ve traced these reports to three separate changes that affected Claude Code, the Claude Agent SDK, and Claude Cowork. The API was…

Official Sourcesfrom Microsoft Research Blog

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI work…

Official Sourcesfrom Together AI Blog

Canary rollouts: upgrade models in production without downtime

A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.

Podcasts & Newslettersfrom Latent Space Newsletter

Foundries vs Navigators: Lowering the Cost of Science

Guest Post: In science, thinking has gotten cheap but doing has not. This asymmetry is reshaping how research companies operate, largely inconspicuously.

Podcasts & Newslettersfrom The MAD Podcast with Matt Turck

Who Feeds the GPUs? Inside AI's Hidden $30B Layer | Renen Hallak, VAST Data

Speaker 1 | 00:00 - 00:17 Sometimes it scares me. We had a customer, one of these AI clouds, they said we're probably going to need about 500 petabytes over the next three years. Last week, came back to us and said, we're gonna need an ext…

Official Sourcesfrom OpenAI News

Better prompt caching for GPT-6

Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.

Official Sourcesfrom Together AI Blog

The Open Source AI Stack

A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead of rebuilding your workflow.

Official Sourcesfrom Together AI Blog

Autoscaling endpoints for LLM inference

GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.

Official Sourcesfrom Microsoft Research Blog

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience…

Official Sourcesfrom Together AI Blog

Configuring Dedicated Model Inference

The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.

Podcasts & Newslettersfrom Unsupervised Learning

Ep 92: xAI Co-Founder Unpacks the Future of Model Development

Speaker 1 | 00:00 - 00:16 Welcome back to unsupervised learning. I'm Jacob Efron. We had an awesome episode today with Igor Babushkin. Igor has been at all the right places at all the right times, leading a lot of really interesting AI wor…