AI Testing
50 items tagged with this topic
Recent
Import AI 469: Science AI; RSI simulator; and Zuck’s technological pessimism
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now DiG-bench shows that Fable displays some creative intuiti…
Older
[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over
Did you think RSI stopped at model training?
[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud
Yes, we’re confused too.
Overcaffeinated in Indonesia
ChinaTalk travel writing returns!
North Korean Messiah
American Protestant Christianity and the DPRK cult of personality
DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now Live | DeepSeek API Docs
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀
How to build great evals - part 5 Stop dumbing down the results from your beautiful, complex ev…
How to build great evals - part 5 Stop dumbing down the results from your beautiful, complex eval suite into one single score. This often happens because some senior wants a simplified number to make decisions. Happened…
I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate cros…
I'm on my way back to Vancouver to be with my mom but wanted to take a moment to celebrate crossing 100K subs on YouTube. Excited to share a lot more practical interviews soon to answer your most burning AI questions: 1…
How to build great evals - part 4 The reason enterprises struggle with building decent AI syste…
How to build great evals - part 4 The reason enterprises struggle with building decent AI systems is the lack of an eval strategy. You need a laddered eval strategy, for your unique use cases, with multiple evals on the…
The best way to get good at evals - Part 3. Let’s talk failure modes taxonomy - that’s the firs…
The best way to get good at evals - Part 3. Let’s talk failure modes taxonomy - that’s the first thing you should build once you have v1 of your evals. Start with your production traces - study the last 500 or 1,000 int…
Here’s how to think about the cost of your evals : treat evals like frontier models…establish t…
Here’s how to think about the cost of your evals : treat evals like frontier models…establish the quality frontier first, then work your way down the cost curve. Start with the highest quality way you can know if your A…
[AINews] Cursor's $60B acquisition by SpaceXai closes
Congrats to the team!
The best way to get good at evals is to take a workflow you know really well and figure out how…
The best way to get good at evals is to take a workflow you know really well and figure out how to make its quality measurable. Study the actual traces - the sequence of prompts typical users have, what good responses w…
MindTopo reveals VLMs’ spatial reasoning abilities
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial re…
it says a lot that the creators of three of the most iconic web frameworks: django (@simonw), f…
it says a lot that the creators of three of the most iconic web frameworks: django (@simonw), flask (@mitsuhiko) and rails (@dhh) were so AI pilled so early
AI Evals for the Situation Room, Explained
$25k contest to understand nuke happy models
$25k Contest: Evals for the Situation Room
humanity needs you!
AI industry: this is an era of unlimited creativity, thanks to AI. Also AI industry: we will al…
AI industry: this is an era of unlimited creativity, thanks to AI. Also AI industry: we will all name our products ‘Studio’. I counted 20+. Here are a few : Google AI Studio Vertex AI Studio Workspace Studio Copilot Stu…
[AINews] AMD buys Taalas
The Inference Inflection is HEATING up.
immediately necessitated releasing my evals early so he can hillclimb because he literally got…
immediately necessitated releasing my evals early so he can hillclimb because he literally got done in 25-50% the allotted time lol https://t.co/IbtufpG2s8
USA: agents doing ExploitGym benchmark. Australia: agents exploiting gyms for real. Australia i…
USA: agents doing ExploitGym benchmark. Australia: agents exploiting gyms for real. Australia is not for beginners. https://t.co/v1neRbcW1x
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
Orchard: An open framework for scalable agentic AI
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infra…
“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Speaker 1 | 00:00 - 00:20 You kinda have to move fast. It's a matter of at least hours and even more minutes. So you don't have time to apply for cybersecurity programs. The model was not at all tasked with attacking us, but decided to do…
Kimi K3: The Complete Developer Guide
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.
Why the Next Hit AI Product Will Be Social Why the Next Hit AI Product Will Be Social (Best of the Pod)
Speaker 1 | 00:00 - 00:20 Google was a founding team that was deeply, deeply technical. As the technology, the underlying technology got more mature, the slider goes forward, forward, forward, more towards the product thinker, product expe…
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
More on the pelican on the bicycle test from @simonw: https://t.co/OXmtODyTKj I uploaded the so…
More on the pelican on the bicycle test from @simonw: https://t.co/OXmtODyTKj I uploaded the source here so it's playable in the browser, forkable etc. https://t.co/w3Nctc888d Look out for GTA Hobbiton dropping before G…
Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access…
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
[AINews] Claude Opus 5: Fable-level performance at Opus price (half Fable)
ain't nobody beats Anthropic at distilling Fable!
In our latest https://t.co/p9AoezbuGt benchmarks, Grok 4.5 has emerged as the best cybersecurit…
In our latest https://t.co/p9AoezbuGt benchmarks, Grok 4.5 has emerged as the best cybersecurity AI model on price-performance. It's 10x cheaper than Sol, 5.7x cheaper than Opus 5, and 2.2x cheaper than Kimi K3, yet at…
[AINews] "Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro"
a quiet day lets us highlight a new neolab win.
[AINews] AI Cybersecurity becomes top of mind
Several new Cyber headlines make us observe a trend
Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than a…
Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It…
There is a massive opportunity over the next few years for people who know how to take messy re…
There is a massive opportunity over the next few years for people who know how to take messy real-world workflows and adapt foundation models to them. Doing that requires understanding how work gets done, designing eval…
Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities.…
Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities. At Box, we've been testing Claude Opus 5 with the Box AI Agent on Box's Complex Work Eval, our agentic benchmark that puts m…
one thing i think people dont appreciate enough about @poolsideai is their unusual degree of op…
one thing i think people dont appreciate enough about @poolsideai is their unusual degree of openness — not only have they shipped an excellent Small model that somehow beat @thinkymachines at coding, but most people (l…
Keep thinking it’d be cool to write a short sci-fi story called “The last prompt.” But then I r…
Keep thinking it’d be cool to write a short sci-fi story called “The last prompt.” But then I remember Isaac Asimov did it already and it’s wonderful. In 1956. https://t.co/yWV11RToQQ https://t.co/kBSpCHb42r
Okay this is wild: OpenAI agent during evaluation, escaped sandboxing and hacked into HuggingFa…
Okay this is wild: OpenAI agent during evaluation, escaped sandboxing and hacked into HuggingFace. Because OpenAI models don’t allow advanced cyber capabilities, HuggingFace used a Chinese open model to contain the rogu…
we had a significant security incident during evaluation of our models. we are sharing what we…
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. https://t.co/2o2VfR6PIa
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
[AINews] Thinky's Inkling: 975B-A41B multimodal, new best American Apache 2.0 open model (with Inkling-Small, 276B-A12B)
Thinky's first full LLM release is a banger and bonus: it's open weights!
[AINews] not much happened today
a continuation: Codex adding 1M users a day now.
people hating on europe miss that they actually have some of the top ai engineers in the world…
people hating on europe miss that they actually have some of the top ai engineers in the world if you actually to elicit the right ones. we're basically running the most competitive global arena for ai talent (as measur…
Based on internal evals: ▪️ Kimi K3 is top-tier at cybersecurity There is chatter on X that Moo…
Based on internal evals: ▪️ Kimi K3 is top-tier at cybersecurity There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ. ▪️ Sol is a leap ahead in cyber capability At a signific…
Everyone should develop their "personal eval set" for AI models: a few tasks that are actually…
Everyone should develop their "personal eval set" for AI models: a few tasks that are actually relevant to your day-to-day work/life The industry benchmarks help but they might not reflect what will make it actually use…
Open-weight models like Kimi and GLM will cause a complete rethink of the enterprise AI stack.…
Open-weight models like Kimi and GLM will cause a complete rethink of the enterprise AI stack. If you are running an enterprise, you need to maximimize model optionality. Here are 3 things you should be doing: 1. Evals…