Skip to main content

Readiness at Portfolio Scale

· 8 min read
Calvin Cheng
Shape what gets built and the value it creates.

Shaping, Not Just Shipping closed the series with a one-page self-assessment and three questions for a leadership team. That works when you are looking at one workflow. It breaks when you are looking at forty — each at a different point on the scope-readiness grid, each owned by a different team, each evolving at its own pace. This post asks what readiness discipline looks like when the unit of governance is not a single deployment but a portfolio.

You cannot run a four-component deep assessment on every tool every quarter. You need triage — and triage requires a theory of consequence.

The single-deployment assumption​

The Other Dimension series built its argument around individual AI-assisted workflows. Each post asked: for this deployment, who owns it, how has it been stress-tested, how confident should you be, and how fast do you learn when it fails? Those are the right questions. They are also expensive questions. A deep readiness assessment for one workflow takes a named reviewer, a failure-scenario exercise, a confidence conversation, and a feedback-loop health check. That is a morning's work for a senior team — assuming they already understand the system.

At NUS, I lead five teams — Enterprise Architecture, Data Engineering, App Development, Integration, and Inter-Reality — each adopting AI-assisted tooling at different velocities and for different purposes. The Data Engineering team uses AI for pipeline generation where a wrong output corrupts downstream analytics. The App Development team uses it for prototyping where a wrong output wastes an afternoon. The Integration team uses it for API mapping where a wrong output breaks a partner handshake. Same organisation, same quarter, categorically different consequence profiles. The single-deployment model does not scale to this without a triage layer.

Triage requires a theory of consequence​

The temptation is to treat all AI deployments equally — one governance process, one approval gate, one review cadence. That produces two failure modes. The first is under-governance: everything gets a light touch because the process cannot afford depth everywhere, and the high-consequence deployments slip through at the same cadence as the low-consequence ones. The second is over-governance: everything gets the deep treatment, the process becomes a bottleneck, teams route around it, and you end up with the same under-governance problem dressed in compliance language.

The alternative is triage by consequence. Not by capability level — a Level 2 assistant generating compliance-adjacent outputs is higher consequence than a Level 3 agent generating internal documentation. Not by team seniority — a junior team working on student-facing systems has higher consequence than a senior team automating internal tooling. Consequence is about what happens downstream when the output is wrong and nobody catches it.

Three tiers, not one process​

At portfolio scale, I have found that readiness governance works in three tiers, distinguished by the blast radius of a wrong output that escapes undetected.

The first tier is full readiness assessment — the four-component deep dive from the series. This applies to deployments where a wrong output reaches someone outside the team: students, researchers, partners, regulators, or commercial counterparts. These get a named reviewer, a pre-deployment failure scenario exercise, a documented confidence profile, and a mapped feedback loop. There are not many of these. In a university technology division with hundreds of AI-assisted workflows, fewer than fifteen belong in this tier at any given time. But those fifteen are the ones that generate institutional risk if ungoverned. The MDM hard-delete incident belonged here — and did not receive this treatment until after the failure became visible downstream.

The second tier is lightweight readiness check — a shorter protocol that asks two questions: who owns it, and what happens if it is wrong? This applies to deployments where a wrong output stays within the team or within internal systems. The check takes thirty minutes, not half a day. It does not require a full failure scenario exercise, but it does require a named human. If nobody can answer "who owns this?" without checking a wiki, the deployment has drifted out of governance regardless of how harmless it appears. Most AI-assisted workflows sit here — perhaps forty to sixty at any time. The discipline is keeping this tier from silently graduating into the first tier without anyone noticing.

The third tier is expiry-based governance — the prototype expiry mechanism from The Invisible Line, applied without additional process. These are experiments, sandboxes, and proofs of concept that have not yet touched anything outside the team's own workspace. They get no assessment beyond the thirty-day expiry. If they survive expiry and someone extends them, the extension conversation is the governance moment — it forces the question of whether this has moved into tier two.

The graduation problem​

The hardest governance problem at portfolio scale is not the initial classification. It is detecting when a deployment has graduated from one tier to the next without a formal decision. A prototype that started in tier three gets extended, then gets connected to a shared data source, then gets referenced in a partner integration spec — and suddenly it is tier one without ever having passed through a readiness assessment.

This is not hypothetical. I have seen it at every organisation I have led a technology function in. At ION Mobility, a dashboard prototype that displayed stale data — tier three, inconvenience — gradually became the source of truth for investor reporting. Nobody made a decision to promote it. The scope crept because it was useful and nobody asked whether its governance posture had kept pace with its importance. The firmware team would never have allowed that creep because the consequence of firmware failure is physical and immediate. Dashboard failure is silent and cumulative — which makes it harder to govern, not easier.

The mechanism that catches graduation is not more process. It is a periodic portfolio review — quarterly at most — that asks one question per deployment: has the blast radius of a wrong output changed since last review? If yes, reclassify and apply the appropriate tier's governance. If no, leave it. The review is fast because most deployments have not changed. The value is in the three or four that have.

What the portfolio view reveals​

When you plot all AI-assisted deployments on the scope-readiness grid from The Other Dimension, the portfolio pattern becomes visible. Most organisations I have worked in show a cluster in the bottom-left quadrant: low scope, low readiness. Many small experiments, none individually dangerous, collectively producing the accumulation risk the series named. A few sit in the top-right: high scope, high readiness — the crown jewels that received deep governance because they were obvious from the start. The danger is not in either cluster. It is in the deployments that have moved right on scope without moving up on readiness — the ones that graduated without a decision.

The portfolio view also reveals resource allocation problems. If your full-readiness-assessment capacity is consumed by deployments that should be in tier two, you do not have capacity for the ones that actually need it. Governance is a finite resource. Spending it on low-consequence deployments is as irresponsible as not spending it on high-consequence ones. The triage discipline is not about being lax — it is about being honest about where consequence concentrates.

The institutional memory layer​

At SGInnovate, I built a talent database of ten thousand professionals. The database was technically populated. It was not institutionally useful until the feedback loop from recruiter outcomes reached the curation team. The parallel at portfolio scale is a registry — not of talent, but of AI-assisted deployments. A registry that is technically populated (every deployment has an entry) but has no feedback loop from consequence events back to classification is the same filing cabinet. It decays. Deployments change scope, teams change ownership, and the registry becomes a compliance artifact rather than a governance tool.

The registry that works is one where consequence events — incidents, near-misses, scope changes, ownership transitions — route back to the classification. Each event is a signal that asks: has this deployment's tier changed? The registry is alive when it responds to those signals. It is dead when it requires a human to remember to update it.

Three questions for a technology leadership team​

If you are leading a division with multiple teams adopting AI-assisted tooling:

  1. How many AI-assisted workflows do you have in production today — not the ones on a roadmap, the ones already running? If you cannot answer within an order of magnitude, you have a visibility problem before you have a governance problem.

  2. Which of those have a blast radius beyond the team that built them — outputs that reach students, customers, partners, regulators, or downstream systems? Those are your tier one. Name them.

  3. When did you last check whether any tier-two or tier-three deployment has graduated — not by design, but by drift? If the answer is never, schedule the first portfolio review this quarter.

The readiness framework from the series gives you the depth. Triage gives you the coverage. Neither alone is sufficient at portfolio scale.


Beyond The Other Dimension · Follow-up 1 of 3 · Next: The Governance Speed Problem