2. Measurement that is not free to fake

Charles Goodhart's 1975 observation was that any statistical regularity collapses once it is used for control. The familiar restatement — when a measure becomes a target, it ceases to be a good measure — is Marilyn Strathern's (1997), and it is hers that is almost always quoted as his. AI has made this acute, because it collapsed the cost of moving the usual numbers.

Donald Kirkpatrick's four levels of evaluation — reaction, learning, behavior, results — supply the diagnosis. Most evaluation stops at the first two levels because those are the two that are easy to collect.

Four levels of evaluation, with Levels 1 and 2 marked free to fake

An AI initiative reporting enthusiasm and completions is reporting Levels 1 and 2. The claim being made — that the organization is better off — is a Level 4 claim. The evidence and the claim are two levels apart, and almost nobody says so out loud.

The operator's test

If someone wanted to move this number without doing the underlying work, how hard would it be?

MetricCost to fakeWhat it actually evidences
Seats provisionedNear zeroProcurement happened
Logins, sessionsNear zeroA tab was open
CompletionsNear zeroContent was clicked through
Outputs generatedNegative — cheaper than not doing itCapacity exists
Outputs accepted after reviewHighSomething met a standard
Cycle time on a named decisionHighA process changed
Rework rateHighQuality changed
Downstream error rateHighThe organization changed

The pattern: anything counting activity is cheap. Anything counting accepted work, or work no longer needed, is expensive — because faking it requires doing it.

Two tests worth carrying. A metric that improves when the AI is switched off is measuring the wrong thing. And a metric that gets worse when governance is added is measuring throughput, not value.