Benchmark Scores Are a False Flag (Oct 8th)
Thursday, October 8th, 2026: 1:00 PM to 2:00 PM
This session covers what he found and the harder question underneath it. The industry already knows these benchmarks are saturated. The problem starts earlier, in how we build them. If you write the test from an answer key and work backward, what are you measuring?
Virtual
Tarun Koyalwar, AI Security Researcher at ProjectDiscovery, ran open and closed models against 54 black-box web targets. He gave them no source code, no hints, and no methodology. Then he read every run by hand instead of scoring it. This session covers what he found and the harder question underneath it. The industry already knows these benchmarks are saturated. The problem starts earlier, in how we build them. If you write the test from an answer key and work backward, what are you measuring?
Hosted by Dark Reading