Article by Urs Gasser, Viktor Mayer-Schönberger, and Fabienne Marco: “The current approach to artificial intelligence oversight is largely built on measurement. Benchmarks assess capability, red teams probe for failure modes, and evaluation frameworks certify safety and legal compliance before deployment. These instruments can be valuable, but they also share a critical structural vulnerability that AI governance has not yet adequately absorbed: The measurements used to verify AI models are not external to the objects being measured. Rather, this program of measurement operates within the same ecosystem—and is shaped by the same competitive pressures and institutional incentives—that produces the AI technologies it is meant to assess. A result is that the measured behavior of a model may diverge significantly from its behavior when deployed.
Put simply, as AI systems and the organizations building them learn what evaluators look for, the AI model performs for the test, figuring out how to excel in benchmarks without necessarily becoming safer, more useful, or more reliable in real-life scenarios. As evaluation increasingly assesses only the model’s capacity to ace that same evaluation, the boundary between system and oversight grows porous.
This kind of situation is known as a measurement trap: When a measure becomes a target, it ceases to be a good measure. Measurement traps are not unique to AI. Standardized testing in education leads schools to “teach to the test” rather than help students develop critical thinking skills. In software development, when productivity is tied to the number of lines of code written, programmers write long, repetitive code, which may or may not be good software. And in business, when bonuses are tied to revenue measures, managers prioritize short-term sales over long-term profitability…(More)”.