← Back to home

AI Testing

50 items tagged with this topic

Recent

Older

Podcasts & Newslettersfrom ChinaTalk

North Korean Messiah

American Protestant Christianity and the DPRK cult of personality

Official Sourcesfrom Microsoft Research Blog

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial re…

Podcasts & Newslettersfrom Latent Space Newsletter

[AINews] AMD buys Taalas

The Inference Inflection is HEATING up.

Official Sourcesfrom Microsoft Research Blog

Orchard: An open framework for scalable agentic AI

Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infra…

Podcasts & Newslettersfrom The MAD Podcast with Matt Turck

“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf

Speaker 1 | 00:00 - 00:20 You kinda have to move fast. It's a matter of at least hours and even more minutes. So you don't have time to apply for cybersecurity programs. The model was not at all tasked with attacking us, but decided to do…

Official Sourcesfrom Together AI Blog

Kimi K3: The Complete Developer Guide

Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.