Posted On: September 23rd 2026, 08:40 pm
I’ve been looking into how new AI models are actually tested, because benchmark scores are usually one of the first things people check when a new model is released.
Benchmarks give different models the same tasks and score them using the same rules. This helps us compare their abilities in areas like coding, maths, and scientific reasoning.
The problem is that models are improving so quickly that some tests are struggling to keep up. Researchers recently reviewed several advanced physics benchmarks and found unclear questions, incorrect reference answers, and grading mistakes. Once corrected, the results showed that the models were far more capable than the original scores suggested.
This is creating opportunities for companies like Snorkel, which builds training data, practice environments, and better ways to measure performance.
Sometimes the biggest opportunities aren’t in the new technology itself. They’re in the tools that help us train it, test it, and understand what it can actually do.
#AI
#Benchmarking #Testing