AI Testing
50 items tagged with this topic
Recent
[AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics
Our head hurts.
Data is the Hard Part
Bharat Patel is currently Accenture’s AI and data lead for its defense portfolio, a former enlisted Navy sailor, and a veteran of the Army acquisition world, with experience at MITRE and across the Department of Defense.
Older
One thing the engineering community will (re)discover is that every program can be ① hardened (…
One thing the engineering community will (re)discover is that every program can be ① hardened (edge cases squashed, inputs tightened, errors handled) and ② optimized (profiled, benchmarked, rewritten) basically ad infin…
Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to H…
Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior.
[AINews] not much happened today
a quiet day.
Ep 94: Applied Compute CEO on the Limits of RL, the New AI Hyperscaler & Why Post-Training Wins Inference
Speaker 1 | 00:00 - 00:02 How do you kind of characterize where we are today? Speaker 2 | 00:02 - 00:25 What we have with RL is we essentially have a hill climbing machine. The hardest part is actually defining the hill to climb, which is…
Advancing computer use with Ironclad
Learn how OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work.
Forecasting space weather risks on power grids
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Fore…
The most common mistake teams make is to treat evals as an additional QA step once the agent is…
The most common mistake teams make is to treat evals as an additional QA step once the agent is built. AI products are fundamentally different. Your evals are your product spec.
A decent amount of agent adoption and deployment is bottlenecked by being able to test, tune, a…
A decent amount of agent adoption and deployment is bottlenecked by being able to test, tune, and optimize agents on “real” work environments. The files agents need to access, the CRM environments, email, and more. It’s…
It’s hard to fathom https://t.co/OL0LzGtvAw getting even faster, but it’s now… much faster 😂 A…
It’s hard to fathom https://t.co/OL0LzGtvAw getting even faster, but it’s now… much faster 😂 As models speed up (as with Astra 𝚞𝚕𝚝𝚛𝚊𝚏𝚊𝚜𝚝), harness overhead matters more and more. Next release will feature even…
[AINews] Gemini 4 Argon: GDM’s answer to Astra/Fable, with 1M output
... but you can’t try it yet unless you are “government users and trusted cyber defenders in the Fairwind Program”
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate a…
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet. https://t.co/jTTAhBsMfu
The future is verification-engineering. Proofs, (e2e) tests, benchmarks, linters… Some tests wi…
The future is verification-engineering. Proofs, (e2e) tests, benchmarks, linters… Some tests will be deterministic, some agentic. This looks great. https://t.co/ZSwn1UEupp
WarTalk: Potemkin Pacific, Midterm Budget Woes
Every service for themselves + autonomy
[AINews] Opus 5.5 is good at explainer videos
a rare feature of a capability
[AINews] The Future of Latent Space
A quiet day lets us discuss the work behind the scenes - now open for business!
Partnering with Accenture on embedded evaluation \ Anthropic
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the nex…
If you want to see interactive explainers of benchmarks and demos I built, you can read it on o…
If you want to see interactive explainers of benchmarks and demos I built, you can read it on our new dev site here: https://t.co/oSTZrBggXM
What is effort really? When do you change it it and why not just use max effort for everything?…
What is effort really? When do you change it it and why not just use max effort for everything? I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results. https://…
You can’t automate what you can’t measure. This means that evals are one of the gates to diffus…
You can’t automate what you can’t measure. This means that evals are one of the gates to diffusion of AI in the enterprise. We can test our deterministic processes through software, but most enterprises have no useful w…
i asked opus 5.5 to explain why personal benchmarks are so important (this is one shot) https:/…
i asked opus 5.5 to explain why personal benchmarks are so important (this is one shot) https://t.co/GE9OFZ0CvQ
There is an extensive and ongoing review related to our agents’ use of internet access during t…
There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to. We have not been as fast as…
[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%
overshadowing more efficient GPT6 models from OpenAI
We ran fresh Next.js evals. The tally: ① Opus 5.5 [𝟿𝟽%] ② GPT 6 Sol [𝟿𝟽%] ③ Fable 5.1 [𝟿𝟽…
We ran fresh Next.js evals. The tally: ① Opus 5.5 [𝟿𝟽%] ② GPT 6 Sol [𝟿𝟽%] ③ Fable 5.1 [𝟿𝟽%] ④ Grok 4.7 [𝟿𝟺%] Notably, Grok is 2x-7x cheaper https://t.co/BAUg14981G
Safety and alignment of "models" is the raging topic today. But the original "safety" debate fo…
Safety and alignment of "models" is the raging topic today. But the original "safety" debate for AI/MachineLearning was actually in autonomous vehicles! My first Waymo ride truly felt like a religious experience. It was…
Migrating from closed to open source models, Together
Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.
6. I forgot to mention Apple and Siri. They have all the devices and a great reputation on priv…
6. I forgot to mention Apple and Siri. They have all the devices and a great reputation on privacy. They can easily win this race but what holds them back imo is their annual release cycle for major updates and not bein…
Building standards for the next phase of AI
OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.
[AINews] Reality Checks on AI News (Yegge shuts down Gas Town, Databricks’ +60% Astra cost)
A dash of cold water keeps the foomers away.
[AINews] Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs
congrats to TypeSafe!
Wife wanted Black Forest cake for her birthday.. I haven’t baked in a long long time, so not pe…
Wife wanted Black Forest cake for her birthday.. I haven’t baked in a long long time, so not perfect but think it turned out ok (thanks @ChatGPT) Happy birthday @shimoleejhaveri - I love you mostest ❤️ https://t.co/5LC3…
the main thing i was excited about launching this week will be next week instead, but imo worth…
the main thing i was excited about launching this week will be next week instead, but imo worth the wait! https://t.co/j8tcW05Fmk
There’s a massive chasm between the power of AI models and the ultimate workflows that enterpri…
There’s a massive chasm between the power of AI models and the ultimate workflows that enterprises are trying to automate. This gap is the opportunity for the applied AI layer to fill. You need to connect the intelligen…
The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Before co-founding Kepler, Vinoo Ganesh led Spark at Palantir and built Project Frontline — a pioneering program for Forward Deployed Engineers. He takes us through the best practices of FDEs.
[AINews] not much happened today
a quiet day
Prediction: we will see movement of a lot more of the brightest minds on frontier model evals t…
Prediction: we will see movement of a lot more of the brightest minds on frontier model evals towards groups like METR over the next 12 months. This skill is currently concentrated in in a few labs, data providers and i…
Worth spending a few minutes this weekend to read this essay in full. Embedded evaluators might…
Worth spending a few minutes this weekend to read this essay in full. Embedded evaluators might sound unusual in tech but it's pretty normal in other industries. Big banks have federal examiners with desks in the buildi…
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussi…
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great id…
Why most enterprise AI efforts fail. 1/ Using old-school playbooks for AI The CEO picks a trust…
Why most enterprise AI efforts fail. 1/ Using old-school playbooks for AI The CEO picks a trusted lieutenant to lead a central AI team. That leader pulls in a group of trusted people from across the company. The problem…
we heard feedback that it's hard to know if your skills are still working with new model releas…
we heard feedback that it's hard to know if your skills are still working with new model releases plugin evals are here to help run `claude plugin eval init` in your plugin folder https://t.co/Q0I1ZnugDs
it's basically impossible to interpret evals by looking at just at the pass/fail scores these d…
it's basically impossible to interpret evals by looking at just at the pass/fail scores these days many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes m…
better benchmark scores don't tell you much about how well a model does on your real work that'…
better benchmark scores don't tell you much about how well a model does on your real work that's why for the last 3 years @every we've done vibe checks on new models: long-form reviews based on hands-on testing on each…
How to build great evals - part 10 Measure the steps, not just the result. Much like high schoo…
How to build great evals - part 10 Measure the steps, not just the result. Much like high school math, it isn’t sufficient just to get the right answer, the steps to get there are critical. Two agent trajectories might…
[AINews] Collusion.wiki: A second undisclosed OpenAI agent swarm incident...
AI News for 9/2/2026-9/3/2026.
AI Benchmarks with @Benchmark & @Vercel. It's ① the most aptly named event in SF history and ②…
AI Benchmarks with @Benchmark & @Vercel. It's ① the most aptly named event in SF history and ② about one of the most important software categories of our generation. The companies that benchmark models and guide the wor…
GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour
We spent 20B+ tokens of GPT-6 Astra to explore everything. Here’s our learnings.
[AINews] Muse Spark 1.3 matches GPT-5.6-Sol, confirming Meta Superintelligence as the newest Frontier Lab, >90% discount for training
an epic comeback story for Meta
Here’s one practical thing you can do this weekend to get better at building AI products. Pick…
Here’s one practical thing you can do this weekend to get better at building AI products. Pick a workflow you know well - from your personal life or your work. Automate the whole thing with AI. Use the AI product you’re…