Strategic Alignment — Did This Produce Value?

Site: DrBill360 Learning Portal
Course: AI Operator — Strategic Track
Book: Strategic Alignment — Did This Produce Value?
Printed by: Guest user
Date: Friday, 28 August 2026, 2:57 AM

Description

The fifth discipline: converting governed AI capability into organizational value you can prove. Five chapters.

1. Did this produce value?

An executive presenting a case to colleagues

Four disciplines can be done well and still produce nothing of value.

Modules 1 to 4 built an AI-augmented workflow that is bounded, informed, reviewed, and whose records hold.

All four can be done well and produce nothing of value. A governed process with no purpose is a well-run process with no purpose.

This module asks the question the Core track deliberately does not:

Did this produce organizational value — and how would you prove it to someone who did not want to believe you?

Why that second clause matters

You will be handed an AI initiative and asked to certify that it worked. Possibly this year. Quite possibly by someone who has already announced the result.

The instruments most of us reach for first — seats, logins, activity, completions — have all just become nearly free to produce in volume. You will be able to report something impressive. It will not be false, exactly. It will simply not be evidence of anything.

If you own evaluation where you work, this is about to become your problem whether or not you chose it.

2. Measurement that is not free to fake

Charles Goodhart's 1975 observation was that any statistical regularity collapses once it is used for control. The familiar restatement — when a measure becomes a target, it ceases to be a good measure — is Marilyn Strathern's (1997), and it is hers that is almost always quoted as his. AI has made this acute, because it collapsed the cost of moving the usual numbers.

Donald Kirkpatrick's four levels of evaluation — reaction, learning, behavior, results — supply the diagnosis. Most evaluation stops at the first two levels because those are the two that are easy to collect.

Four levels of evaluation, with Levels 1 and 2 marked free to fake

An AI initiative reporting enthusiasm and completions is reporting Levels 1 and 2. The claim being made — that the organization is better off — is a Level 4 claim. The evidence and the claim are two levels apart, and almost nobody says so out loud.

The operator's test

If someone wanted to move this number without doing the underlying work, how hard would it be?

MetricCost to fakeWhat it actually evidences
Seats provisionedNear zeroProcurement happened
Logins, sessionsNear zeroA tab was open
CompletionsNear zeroContent was clicked through
Outputs generatedNegative — cheaper than not doing itCapacity exists
Outputs accepted after reviewHighSomething met a standard
Cycle time on a named decisionHighA process changed
Rework rateHighQuality changed
Downstream error rateHighThe organization changed

The pattern: anything counting activity is cheap. Anything counting accepted work, or work no longer needed, is expensive — because faking it requires doing it.

Two tests worth carrying. A metric that improves when the AI is switched off is measuring the wrong thing. And a metric that gets worse when governance is added is measuring throughput, not value.

3. Reviewing against intent

Four questions after deployment. The order matters, and one of them is almost always skipped.

QuestionCommon failure
1Accuracy — is the output right?Sampled once at launch, never again
2Efficiency — did it cost less?Counts AI time, ignores review time
3Alignment with intent — is it doing what we meant?Never asked, because the task succeeded
4Downstream impact — what changed elsewhere?Invisible unless someone looks

Question 3 is the one that gets skipped — and it is the one Direction exists to catch. A workflow can pass accuracy, efficiency and impact while doing something nobody intended, because the task was completed correctly and the problem it was meant to solve was never restated.

The efficiency trap

Most efficiency calculations count what AI saved and omit what review cost. If verification time is not in the denominator, the number is fiction — and the verification ceiling from Module 3 is precisely the cost being left out.

Evaluating tools, not just capability

Capability is the easy half, and the half vendors demonstrate.

CriterionThe question
CapabilityCan it do the work?
BoundabilityCan its scope be constrained — and is the constraint enforceable?
AuditabilityDoes it record what it did and why?
AttributabilityCan you tell which actor did what?
Exit costWhat does leaving cost — data, workflow, skills?

Attributability is the one discovered too late. A platform where every action is logged under one shared service identity cannot support after-the-fact accountability, however complete the logs look. That is the Module 4 attribution failure arriving as a procurement decision rather than a bug.

4. What nobody budgets for

AI-augmented work does not slot into an unchanged organization. Three changes are required, and initiatives that defer all three have deferred their own results.

Role redesign

If the constraint is verification capacity, then the scarce role is not the producer — it is the reviewer. And reviewing is currently nobody's job description, which means it is being done in the margins of jobs designed for something else.

Capability development

The four Core disciplines are not intuitive. An organization deploying AI without teaching them is relying on individual judgment at exactly the point where judgment is hardest and least supported.

Policy updates

Approval tiers, retention, disclosure, escalation paths — most were written for a world in which every action had a human author. They do not fail loudly when that stops being true. They simply stop describing what happens.

The business case

The artifact this module produces. Six sections:

SectionMust contain
ProblemThe organizational problem, stated before any mention of AI
InterventionWhat the workflow does, and its scope boundaries
GovernanceApproval tiers, the record, who is accountable
MeasurementMetrics that survive the free-to-fake test, with baselines
CostIncluding review time and the three changes above
RebuttalWhat would show this was the wrong call

The last row is not optional. A business case with no stated falsification condition is advocacy. That is Toulmin's rebuttal, arriving at the executive level.

5. Two questions worth arguing about

Both of these are live disputes. This course holds a position on each, and you are expected to argue against it if the evidence in your domain supports that. A certification that teaches only settled answers produces operators who fail on the first novel case.

ISO/IEC 42001 or NIST AI RMF?

ISO/IEC 42001NIST AI RMF
FormCertifiable management-system standardVoluntary framework — Govern, Map, Measure, Manage
EnforcementExternal auditInternal discipline
Gives youA certificate procurement can demandA way of thinking, adaptable
RiskCertifying the management system rather than the outcomesHolds exactly as well as internal will does

This is Module 4 in the wild. ISO is enforceability with external backing. NIST is enforceability by internal will — espoused theory that has to be genuinely held. The real choice is about which failure your organization is more prone to.

Is governance an accelerator or a brake?

The position this course takes: structured organizations move faster, with less risk, and more durably.

The strongest case against. Governance carries real latency — review cycles, documentation, approval queues. In a fast market, first-mover advantage may exceed the cost of some bad output. Worse, over-governance produces shadow AI: people route around the process entirely, which is more dangerous than light governance. And the accelerator claim may be survivorship bias — we see the governed organizations that succeeded, not the ungoverned ones that also did.

Assignment 5.3 requires you to make that case, not the instructor's.

A limit on everything in this module

The measurement argument comes from a domain where outputs are documents and code, and quality is judgable within days. Where feedback loops run for years — clinical outcomes, safety, education — Level 4 evidence may be genuinely unavailable when the decision must be made. The honest position there is a stated proxy with its limitations named, not a confident number.

How long is your feedback loop, and what will you claim in the meantime?

A working meeting in progress

Both questions are live. You are expected to argue against the instructor on either.