← Back to home

Speed & Cost

50 items tagged with this topic

Recent

Older

Official Sourcesfrom Together AI Blog

Autoscaling endpoints for LLM inference

GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.

Official Sourcesfrom Microsoft Research Blog

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience…

Official Sourcesfrom Together AI Blog

Configuring Dedicated Model Inference

The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.

Podcasts & Newslettersfrom Unsupervised Learning

Ep 92: xAI Co-Founder Unpacks the Future of Model Development

Speaker 1 | 00:00 - 00:16 Welcome back to unsupervised learning. I'm Jacob Efron. We had an awesome episode today with Igor Babushkin. Igor has been at all the right places at all the right times, leading a lot of really interesting AI wor…

Official Sourcesfrom Together AI Blog

What does 99.9% uptime mean for inference?

Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.

Chinese Modelsfrom Together AI Blog

Open, convenient and predictable: Introducing Provisioned Throughput

Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no infrastructure t…

Podcasts & Newslettersfrom Training Data

Inside Zipline's Autonomous System: 140M Miles, Zero Incidents

Speaker 1 | 00:00 - 00:21 I remember being in Rwanda early days and going out and meeting with some of the doctors and lab techs that we were serving and asking for them, like, you know, how's it going? What what do you think? What's your…

Podcasts & Newslettersfrom Latent Space Newsletter

[AINews] It's Meta-Harness Summer

Move over, Harness Engineering, it is time for the harness of harnesses!