Research · Perspective

If AI Does the Work, What Must the Institution Still Be Able to Do?

Task capability is not role substitutability. Reliability is not authority. And keeping the outputs is not keeping the capability.

AI can increasingly perform the observable tasks of a role while the institution quietly loses the capability the role carried. The more useful question is not which jobs AI will take, but what the institution must still be able to do for itself.

Somewhere in your organisation, a role is about to disappear.

It will not happen dramatically. No announcement will describe it that way. A capable AI system now drafts the analyses, reconciles the records, prepares the briefings and answers the routine questions. The person who used to do those things moves on, and the position is not backfilled.

The outputs still arrive. The monthly report is on time. The correspondence still goes out. The dashboard still updates.

From a distance, nothing looks different.

That is what should concern us.

Most of the public conversation about AI and work asks which jobs will be automated. It is an understandable question. It is also, I think, the wrong level of analysis. The evidence increasingly points somewhere else: not at jobs, but at tasks; not at what AI can do, but at what the institution quietly stops being able to do.

The question I want to explore is this:

If AI increasingly does the work, what must the institution still retain the capability to do for itself?

I do not think the answer is obvious. But I think the path to it runs through a series of distinctions we have been blurring.

A task is not a role

The first distinction is between what a person does and what a role carries.

Labour-market research has been making this point for a decade. Automation acts on tasks, not jobs. David Autor's well-known paradox of automation observes that technology substitutes for workers in routine, codifiable tasks while complementing them in everything that requires flexibility, judgement and common sense — and that journalists and experts alike systematically overstate the substitution and miss the complementarity.[1]

The same pattern holds in how we measure AI exposure. The most widely cited estimate found that around 80 per cent of the US workforce could have at least 10 per cent of their tasks affected by language models, and around 19 per cent could see half or more of their tasks affected.[2] Those numbers are usually read as a threat assessment. They are actually a task assessment. The authors are explicit that exposure at the task level does not translate one-to-one into job automation.

Current usage data tells a similar story. Anthropic's Economic Index finds that real-world use remains concentrated in a surprisingly small set of tasks — the ten most common tasks account for roughly a quarter of conversations on Claude.ai and nearly a third of enterprise API traffic.[3] On the consumer product, collaborative use slightly exceeds fully delegated use. The work being done is real, but it is narrower than the rhetoric.

The more interesting finding is what happens when tasks leave a role. Anthropic's analysis distinguishes occupations by which tasks AI absorbs. Remove AI-performed tasks from a travel agent's role and what remains is routine ticketing and payment collection — the role is deskilled. Remove them from a property manager's role and what remains is contract negotiation and stakeholder management — the role is upskilled.[3]

Same technology. Opposite effects. The difference is not in the tasks. It is in what the role carried around them.

A role is not a bundle of tasks. It is a vessel for institutional knowledge, context, relationships, expertise, judgement, authority, accountability and the capacity to adapt when the situation stops matching the procedure. Task lists are observable. The carrying capacity of the role mostly is not — until it is gone.

The outputs can continue while the capability thins

This leads to a second distinction, and it is the one I find institutions most often miss.

AI can preserve the visible outputs of a role while obscuring the capability that disappears with it.

Consider the monthly risk report. For years, a senior analyst produced it. The document was the visible output. But the document was never really the point. The point was the monthly act of institutional attention: someone who noticed the unusual number, remembered the last time something similar appeared, knew which manager would give a straight answer, and understood which risks the template had never captured.

The AI system now produces the report. It is on time, well formatted and usually right.

What is gone is not the document. It is the attention.

The institution may not notice for years, because the visible outputs are the last thing to degrade. The report still arrives. The capability to know whether the report deserves to be trusted has quietly departed.

An institution can keep every output of a role and lose the role.

Repeatable, correct — and still not the control

The third distinction is about control, and it is where I think the most consequential confusion now lives.

Suppose a chief financial officer holds a financial delegation: commitments above a threshold require her approval. In the Australian public sector, this is not a custom. Under the Public Governance, Performance and Accountability Act 2013, an accountable authority delegates power by written instrument, some powers cannot be delegated at all, and the instrument can be directed, varied and revoked.[4] The delegation is an enacted artefact of authority. It exists whether or not anyone behaves according to it.

Now suppose an AI system is asked to apply the delegation rules. It reads the instrument, interprets each transaction and predicts whether approval is required.

Three different properties are easily conflated here.

Repeatability. The model applies the rule the same way across ten thousand transactions.

Correctness. When audited, its applications match the instrument.

Enforcement. When a transaction exceeds the delegation, the transaction does not proceed.

The first two are behaviours. The third is a property of an authoritative mechanism. A frontier model can exhibit repeatability and correctness at impressive levels and never possess the third. Take the limiting case: a model that applies the delegation perfectly, on every transaction, without exception. Even then, its behaviour would not constitute the control. Conformity to a boundary — even flawless conformity — is not authority over it, and a record of correct interpretation is not a mechanism of enforcement. The boundary has to hold independently of how the model behaves: on the transactions it interprets correctly, and equally on the ones it does not.

The model does not refuse the transaction. It predicts that refusal is appropriate — and a prediction about a boundary is not the boundary.

The finance system that will not post the payment is the control. The written instrument is the authority. The model is, at best, an interpreter of both.

This is why I want to test a proposition that I think the evidence supports:

Deterministic controls can constrain what a system is permitted to do. They cannot by themselves ensure that the institution continues to know what it ought to do.

The controls matter enormously. Every delegation should be enforced by mechanisms that do not negotiate. But notice the quieter risk in the example: the control can keep holding while the institution forgets why the limit exists, when it should change, and what would happen if it were wrong. Enforcement without understanding is its own kind of fragility.

Why the pilot impresses and the operation disappoints

If models are so capable, why not simply let them carry more?

The honest answer begins with genuine respect for the capability trajectory. METR's long-task evaluations show the length of task a frontier agent can complete with 50 per cent reliability doubling roughly every seven months since 2019. By early 2026 the public frontier reached tasks that take a human expert around twelve hours, with the leading measured model exceeding sixteen hours — at the edge of what the evaluation suite can reliably measure.[5]

That is extraordinary progress. It is also exactly the number that misleads.

The same evaluations show the 80 per cent reliability horizon is roughly an hour and a half.[5] A system that can sometimes complete a sixteen-hour task is a system that reliably completes a ninety-minute one. Demonstration capability and operational reliability are different properties, separated by an order of magnitude.

The boundary is also invisible in advance. In the best-known field experiment on knowledge work, 758 consultants using AI inside its capability frontier completed more tasks, faster, at markedly higher quality. On a task selected to sit just outside that frontier — similar in apparent difficulty — consultants using AI were nineteen percentage points less likely to reach the correct answer.[6] Nobody could see the cliff from inside the pilot.

And then there is the problem of the trajectory. I want to be careful here, because a crude version of this claim — that more context simply makes AI more random — is wrong. The accurate version is more serious. Chroma's long-context research evaluated eighteen frontier models and found performance degrades non-uniformly as input grows, even on simple retrieval and replication tasks.[7] Long-horizon agent research finds models abandoning tasks or producing uncertain wrong answers as interaction trajectories lengthen — well before their context windows are full.[8] Each step's output becomes the next step's input. Small errors compound. State accumulates. Exceptions multiply.

A pilot is a short, curated trajectory: bounded context, clean state, few exceptions. An operation is a continuing stateful interaction among context, tools, actors, exceptions and prior decisions — a trajectory whose shape at any point depends on the accumulated state of everything that has already happened.

Operational reliability is a property of the trajectory, not of the demonstration.

None of this argues against deployment. It argues that impressive bounded performance does not, by itself, establish operational reliability — and that the gap between the two is where institutional capability either survives or erodes.

What remains when the reasoning is gone

The fourth distinction is about memory.

When a human role disappears, the institution typically keeps the decisions and loses everything that made them intelligible: the context that was considered, the evidence that was weighed, the alternatives that were rejected, the reasoning that connected them, the authority under which the judgement was exercised, and — critically — the outcomes that later proved the judgement right or wrong.

I have started calling this judgement provenance. Most institutions have never had it. They have outcome records: minutes, approvals, file notes. They do not have the ecology of the judgement itself.

Without provenance, institutional experience cannot be retrieved, audited or learned from. The institution retains the decision and loses much of the experience that produced it. And because AI can keep the visible outputs flowing, that leakage leaves no gap in the record: the file looks complete precisely while the capability behind it thins. The institution has a history of decisions and steadily less access to its own experience.

With captured provenance, something important becomes possible. AI can do what it is genuinely good at: retrieve and amplify institutional experience. How have we handled situations like this? What did we consider? What happened afterwards? The model becomes a means of reaching the institution's accumulated judgement rather than a substitute for it.

Provenance also clarifies the automation boundary. When you can see how judgements were actually formed, you can see which were mechanical — the rule applied, the threshold checked — and which carried genuine discretion. The first category becomes a candidate for deterministic automation. The second is where institutional judgement lives.

This is the premise behind Judgement Governance™ and its artefact model: the Judgement Object™ as the lifecycle container for a matter, and the Judgement Passport™ as the provenance and accountability record for a single judgement within it. The discipline exists to keep judgement informed, visible, contestable and accountable — whether or not AI participated.

Reasoning is not judgement

The fifth distinction is the one most likely to be misread, so let me be precise.

I am not claiming AI cannot exercise something like judgement. Frontier models reason impressively about ambiguous situations, and they will get better at it. The question is not what models can do. The question is what makes judgement institutional.

Judgement, in the sense that matters here, is the accountable determination of what ought to be done where the appropriate outcome is not mechanically determined. Its legitimacy does not come from the quality of the reasoning alone. It comes from where the reasoning sits: within a structure of purpose, rules, authority, precedent, risk, responsibility and consequence.

A model may produce an excellent recommendation. It does not hold the delegation. It does not own the consequence. It cannot be answerable to the board, the regulator, the court or the citizen.

A recommendation becomes an institutional decision when someone with legitimate authority owns its consequences. That is not a limitation of current models. It is a property of what institutions are.

The longer question: who becomes the expert?

The final distinction is about time.

In 1983, Lisanne Bainbridge described the ironies of automation: the more capable the automated system, the more the human operator is left with the hardest problems — while the system, by handling all the routine ones, deprives the operator of the practice needed to solve them.[9] The irony is that automation does not remove the need for expertise. It removes the conditions under which expertise forms, and then demands expertise at the moment of failure.

We now have direct evidence that this is not theoretical. In a field experiment with nearly a thousand students, unguarded access to a capable AI tutor improved practice performance by 48 per cent — and when the access was removed, those students performed 17 per cent worse on their own than students who never had it. A guarded tutor, designed to give hints rather than answers, improved practice performance even more and avoided the harm entirely.[10]

The tool was the same. The design determined whether capability accumulated in the human or only in the session.

The institutional version of this question is uncomfortable. If junior staff never work through the cases, exceptions and corrections that routine financial work throws up, where does the next CFO's judgement come from? If analysts never write the first draft badly — and never have it corrected — where does the next chief analyst's taste come from? Expertise forms through the work: not through any single task, but through the accumulation of cases handled, exceptions met and errors corrected. Organisations learn through correction: outcomes traced back to judgements, judgements revised, rules updated through legitimate authority. Progressive automation can sever every link in that chain while every output continues to look correct.

This is where I want to preserve an honest tension. The evidence does not say stop automating. The guarded tutor did not harm learning; the unguarded one did. The difference was design, not abstinence. What the evidence does not yet tell us is how much provenance is enough, how much practice is enough, or where exactly the capability floor sits for a given institution. Anyone who claims certainty about that is ahead of the evidence.

What the institution must still be able to do

Let me return to the opening question.

If AI increasingly does the work, the institution must retain the capability to:

establish what is true, current and authoritative — rather than inferring it;

hold authority — and bound it — through mechanisms that enforce rather than predict;

exercise judgement where outcomes are not mechanically determined, with legitimate authority and accountability;

remember how its judgements were formed — the context, evidence, alternatives, reasoning and outcomes;

correct itself — trace outcomes back to judgements, and turn resolutions into reviewed institutional learning rather than silent model drift;

and change its tools without losing itself — replace models, vendors and systems without reconstructing what it means, knows or is authorised to do.

That list is an architecture, and it is the direction we have been developing as the Sovereign Operational Capability architecture. AI operates at two boundaries: interpreting human expression and situations at the edge, and observing operations to propose improvements under governance. The institutional core — meaning, capability, knowledge, rules, authority, evidence and human judgement — remains institution-owned and model-independent. Needs are resolved against meaning the institution owns. Capabilities execute through bounded mechanisms and replaceable tools. Cognitive aides assist authoritative roles without inheriting their authority.

The principle is short enough to fit on one line:

AI at the boundaries. Institutional capability at the core.

This is not an argument for less AI. An architecture that keeps the institution authoritative can use more AI, more confidently, in more places — because what must endure is never confused with what can be replaced. Tools are replaceable. Capabilities are durable.

The work continues

The role I described at the beginning will keep disappearing. That is not, by itself, a failure. Some roles should change; some tasks should never have consumed human attention at all.

But the outputs will keep arriving either way. That was never the hard part.

The hard part is whether anyone can still tell whether those outputs deserve to be trusted. Whether the institution remembers how it decided. Whether it is still forming the expertise it will need in ten years. Whether it can change its mind, on purpose, with authority.

Whether, five years from now, the institution is more capable — or only more dependent.

If AI increasingly does the work, the institution must retain the capability to know what it ought to do.

References

[1] David H. Autor, "Why Are There Still So Many Jobs? The History and Future of Workplace Automation", Journal of Economic Perspectives, 29(3), Summer 2015, pp. 3–30. https://doi.org/10.1257/jep.29.3.3

[2] Tyna Eloundou, Sam Manning, Pamela Mishkin and Daniel Rock, "GPTs are GPTs: Labor Market Impact Potential of LLMs", Science, 384(6702), 2024. https://doi.org/10.1126/science.adj0998

[3] Anthropic, "Anthropic Economic Index Report: Economic Primitives", January 2026 (November 2025 usage sample). https://www.anthropic.com/research/anthropic-economic-index-january-2026-report

[4] Public Governance, Performance and Accountability Act 2013 (Cth), s. 110 (delegation by accountable authority by written instrument; non-delegable powers); Department of Finance, "PGPA legislation, associated instruments and policies". https://www.finance.gov.au/government/managing-commonwealth-resources/pgpa-legislation-associated-instruments-and-policies

[5] METR, "Measuring AI Ability to Complete Long Tasks", arXiv:2503.14499, 2025; and METR, "Task-Completion Time Horizons of Frontier AI Models" (Time Horizon 1.1), updated May 2026. https://metr.org/time-horizons/

[6] Fabrizio Dell'Acqua et al., "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality", Harvard Business School Working Paper 24-013, 2023. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321

[7] Kelly Hong, Anton Troynikov and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", Chroma Technical Report, July 2025. https://www.trychroma.com/research/context-rot

[8] Shijie Xia, Yikun Wang, Zhen Huang and Pengfei Liu, "Diagnosing and Mitigating Context Rot in Long-horizon Search", arXiv:2606.29718, 2026. https://arxiv.org/abs/2606.29718

[9] Lisanne Bainbridge, "Ironies of Automation", Automatica, 19(6), 1983, pp. 775–779. https://doi.org/10.1016/0005-1098(83)90046-8

[10] Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman, "Generative AI without Guardrails Can Harm Learning: Evidence from High School Mathematics", Proceedings of the National Academy of Sciences, 2025. https://doi.org/10.1073/pnas.2422633122


Dakshan Pothuhera
Founder, DataMPowered®

This research informs how ifCEM supports governed work, with reviewable workflows designed for accountable adoption in organisations.

Explore ifCEM →

Want to discuss how ifCEM could support your organisation? Let's talk.

Start a conversation
If AI Does the Work, What Must the Institution Still Do? | DataMPowered