Editorial diagram contrasting model-centric autonomy with institution-centred authority structures

Research · Essay

Is AI Easier to Weaponise Than to Institutionalise?

The uncomfortable gap between what AI can do and what institutions can responsibly allow it to do

Dakshan Pothuhera

Founder & Chief Strategist

16 min read

AI can look extraordinary when raw capability is enough. Inside an institution, capability must also operate within identity, authority, policy, accountability and judgement. What if that explains why AI can be easier to weaponise than to institutionalise?

Something strange is happening with AI.

The models are getting better.

They can reason across harder problems. They can write and debug software. They can search, analyse, summarise and use tools. They can work across longer sequences of activity. Give them more compute, better context and better ways to verify their work, and they often become more capable again.

In cybersecurity, this is becoming serious enough that governments, researchers and AI companies are warning about the growing offensive capability of frontier models.

But walk into a large organisation and you often see a very different picture.

There are impressive demos.

There are pilots.

There are copilots.

There are agent experiments.

There are strategy decks promising transformation.

Then comes production.

Suddenly the impressive AI system runs into identity systems, fragmented data, permissions, policies, legacy applications, privacy obligations, records management, approval processes, professional responsibilities and people who ultimately have to answer for what happened.

The magic seems to fade.

That raises an uncomfortable question:

Is AI easier to weaponise than to institutionalise?

Not because AI somehow prefers bad actors.

And not because enterprises are incapable of adopting technology.

The reason may be simpler.

A bad actor can extract value from capability. An institution has to establish whether that capability can be trusted with responsibility.

Those are two very different problems.

And that difference may explain why AI can look extraordinary in a demo and much less extraordinary once it enters the real world.


Perhaps we have misunderstood the breakthrough

We usually describe foundation models as a breakthrough in artificial intelligence.

That is true.

But I think it may also be hiding something equally important.

For most of computing history, if we wanted a machine to solve a problem, we first had to formalise the problem.

Someone had to understand the requirement.

Data had to be structured.

Rules had to be encoded.

Software had to be written.

Interfaces had to be designed.

Only then could the computer repeatedly perform the work.

Foundation models changed that relationship.

We can now describe a loosely formed problem in ordinary language and cause a huge amount of computation to be applied to it.

The problem does not always have to be fully formalised first.

The person does not necessarily need to know how to code.

The desired output may not even be precisely specified.

In effect:

Language became an interface to computation.

That is a profound change.

And it may be one of the most important things we have overlooked in all the excitement around AI.

What we experience as "the intelligence of the model" is increasingly not just the model itself.

There is the model.

But there is also compute.

Context.

Search.

Retrieval.

Tools.

Code execution.

Repeated attempts.

Candidate answers.

Verification.

Critique.

Sometimes other models.

And underneath all of it is a large computing infrastructure doing an enormous amount of work.

Research into inference-time or test-time scaling has shown that additional computation at the point of solving a problem can materially improve performance on some difficult tasks [11].

It is not as simple as "more compute equals more intelligence".

How that compute is used matters.

Search matters.

Verification matters.

The way the system reasons matters.

But the broader point remains.

What we are interacting with is not just a model.

It is a computational system.

And perhaps that is the real breakthrough.

We may have created a general-purpose linguistic interface through which huge amounts of computational capability can be directed at problems that were previously too ambiguous, too expensive or too difficult to turn into software.

That is extraordinary.

But it also creates a problem.

Because computational capability is not the same thing as institutional capability.


An attacker does not need an institution

Cybersecurity makes this difference easier to see.

The Australian Signals Directorate has been tracking meaningful increases in the cyber capabilities of frontier models [1].

The Bank for International Settlements has examined whether frontier AI could create an asymmetric advantage for attackers [2].

The Bank of England has also pointed to the long-standing imbalance in cyber defence: an attacker may only need to find one overlooked weakness, while defenders have to protect thousands of possible points of failure [3].

Recent threat-intelligence reporting also suggests that AI is moving beyond simple assistance into longer sequences of malicious activity [4].

None of this means AI has suddenly become an autonomous cyber weapon.

Humans still matter.

Targets matter.

Access matters.

Objectives matter.

Judgement still matters.

But there is another point underneath this.

An attacker does not necessarily need AI to be reliable.

If a system works some of the time, that may still be useful.

For an institution, unreliable performance can be a major problem.

For an attacker, it may simply be a conversion rate.

Generate enough attempts.

Probe enough targets.

Produce enough variations.

Discard the failures.

Exploit what works.

The system does not need to survive a board review.

It does not need procedural fairness.

It does not need a defensible audit trail.

It does not need to prove that the person requesting the action had delegated authority.

It does not need to comply with records-management obligations.

It does not need to explain which policy applied.

It does not need to offer someone a way to contest the outcome.

It only needs to work often enough to be useful.

That is a very different threshold.

And it may explain part of the asymmetry we are beginning to see.

Bad actors can sometimes monetise probabilistic capability directly. Institutions cannot.


The institution asks different questions

Now take that same AI capability and place it inside a bank, hospital, insurer, university, government department or large company.

Ask a seemingly simple question:

Can AI handle this?

Very quickly, the question changes.

Who is asking?

Who are they acting for?

Are they authenticated?

What authority do they have?

What information can they access?

Where did that information come from?

Is it current?

Can it be shared?

Which policy applies?

Which jurisdiction applies?

Does another organisation need to participate?

Is the system actually allowed to perform the action?

Does a human need to approve it?

What happens if two sources conflict?

What happens if the situation changes?

What happens if the model is uncertain?

What happens if the model is confidently wrong?

What record needs to be retained?

Can the outcome be challenged?

Who owns the consequence?

And eventually:

Who is accountable?

None of these questions disappears because the model becomes smarter.

In some cases, they become more important.

This is where the distinction matters.

Computational capability asks:

Can the system do the task?

Institutional capability asks:

Can the task be done by the right actor, using the right information, under the right authority and rules, with an appropriate chain of responsibility and a defensible outcome?

Those are not the same thing.

The first is advancing incredibly quickly.

The second is where much of the hard work of enterprise AI begins.


The pilot proves computational capability

This also helps explain why AI demonstrations can be so impressive.

A demo is usually built around the strengths of the model.

The problem is bounded.

The context is curated.

The required data is available.

The desired outcome is visible.

Failure is relatively cheap.

Someone is often nearby to intervene when something goes wrong.

Under those conditions, frontier AI can be remarkable.

The organisation watches the demo and understandably asks:

If it can do this, why can't we deploy it?

Then the pilot becomes a production system.

And production reveals everything the demonstration was able to ignore.

Identity.

Access.

Data quality.

Integration.

Policy.

State.

Authority.

Privacy.

Security.

Exceptions.

Legacy systems.

Records.

Human escalation.

Audit.

Cost.

Reliability.

Accountability.

Suddenly the problem looks very different.

What appeared to be an AI problem starts to look like an institutional architecture problem.

That is why I think this distinction matters:

The pilot proves computational capability. Production demands institutional capability.

Or more simply:

The pilot demonstrates the model. Production reveals the institution.


The evidence of this gap is starting to build

This is not just anecdotal anymore.

Gartner has reported that at least half of generative AI projects had been abandoned after proof of concept by the end of 2025, citing poor data quality, weak risk controls, rising costs and unclear business value [6].

McKinsey's work on agentic AI points to wide experimentation but far fewer organisations achieving scaled, tangible value [7].

Forrester has observed a similar gap between enterprise enthusiasm and meaningful production deployment [8].

In Australia, Deloitte's 2026 research found that only 28 per cent of surveyed organisations had moved at least 40 per cent of their AI pilots into production [9].

That does not mean enterprise AI is failing.

There are plenty of areas where it is working well.

Coding is a good example.

AI can generate code.

The code can be compiled.

Tests can be run.

Failures can be detected quickly.

Outputs can be compared.

Feedback loops are fast.

The problem is often easy to break down.

That makes software development unusually friendly to computational capability.

And unsurprisingly, coding is one of the areas where some of the clearest productivity gains are being reported.

That tells us something important.

AI seems strongest where the problem can be bounded, the feedback loop is strong and success can be verified.

Institutions contain many such problems.

But institutions are not made entirely of them.


Real work is messy

Research from METR is useful here [10].

Frontier AI agents have become increasingly capable of completing longer and more difficult software engineering tasks.

That is significant.

But METR also makes an important point [10].

Real jobs are not just collections of clean, well-scored tasks.

Objectives can be ambiguous.

Success can be subjective.

Information can be incomplete.

Priorities can conflict.

People disagree.

Situations change.

Authority matters.

Judgement matters.

And AI systems tend to struggle more as tasks become messier.

This is easy to underestimate.

A benchmark can ask:

Did the system get the right answer?

An institution may have to ask:

Did it get the answer in the right way?

Those are very different questions.

An insurance decision, procurement exercise, regulatory determination, welfare assessment, financial transaction or clinical process may produce an apparently sensible result and still be institutionally unacceptable.

The wrong information may have been used.

The person may not have had authority.

A mandatory step may have been skipped.

Another party may have needed to be consulted.

A policy may have changed.

The decision may have required professional judgement.

The outcome may even be correct for the wrong reason.

The model may see a successful answer.

The institution may see a governance failure.


And then came the agents

The industry's answer to many of these limitations is increasingly agentic.

If a model can reason, perhaps it can plan.

If it can plan, perhaps it can use tools.

If it can use tools, perhaps it can take action.

If it can take action, perhaps several agents can work together.

And if enough of these capabilities are assembled, perhaps the system can perform significant parts of enterprise work autonomously.

There is real technological progress behind this.

But there is also a question that deserves more attention.

What if autonomy is not the capability the enterprise is missing?

NIST's work on AI-agent identity offers a clue [5].

As agents become more capable of taking actions, the discussion quickly moves to identification, authentication, authorisation, auditing and non-repudiation.

That is not accidental.

The moment software stops merely suggesting an action and begins to perform one, a new question appears.

On whose authority?

And once the action has consequences:

Who answers for it?

So enterprises start surrounding agents with identities, permissions, policies, audit, monitoring, approval boundaries and human oversight.

But notice what has happened.

We started with an intelligent model.

Then we gave it tools.

Then autonomy.

Then identity.

Then permissions.

Then memory.

Then workflow.

Then policies.

Then governance.

Then audit.

Then human escalation.

At some point, we should ask:

Are we making the model smarter, or slowly rebuilding the institution around it?


More intelligence does not create authority

This is where I think the current AI trajectory deserves challenge.

When models hit limits, our instinct is to add more.

More parameters.

More compute.

More context.

More reasoning.

More tools.

More agents.

More autonomy.

All of these can improve computational capability.

But none of them automatically creates institutional authority.

A more intelligent model does not become the authoritative source of someone's identity.

A longer context window does not create consent.

An agent does not become a case officer simply because it can perform similar steps.

A model capable of interpreting policy does not automatically gain the authority to make every judgement described by that policy.

More inference-time compute may produce a better answer.

It does not make that answer legitimate.

That distinction matters more as AI becomes more capable.


Perhaps we are trying to make the model become the enterprise

Most established organisations already contain enormous amounts of capability.

Identity platforms establish who people are.

Access systems determine what they can do.

Databases and registries hold authoritative information.

Rules engines apply deterministic logic.

Workflow systems coordinate activity.

Case-management systems maintain state.

Payment systems move money.

Document systems preserve records.

Notification systems communicate.

APIs connect organisations.

Professionals exercise delegated authority.

Teams make judgements.

Institutions maintain relationships with other institutions.

None of these capabilities disappeared because a model learned to reason.

Yet much of the current enterprise AI conversation seems to assume that more and more of this environment should somehow be absorbed into an intelligent layer.

Perhaps that is the wrong direction.

The model does not need to become the database.

It does not need to become the identity system.

It does not need to become the rules engine.

It does not need to become the payment platform.

It does not need to become the institutional record.

And it does not necessarily need to become the person authorised to judge.

Perhaps the opportunity is simpler.

Let each capability do what it is good at.


Use the model for what it is

Foundation models are extraordinary at working with language.

They can interpret ambiguity.

They can synthesise information.

They can recognise patterns.

They can translate between different forms of information.

They can generate possibilities.

They can help reason through complex situations.

They can make difficult systems much easier for people to interact with.

Those are remarkable capabilities.

We do not diminish them by defining their role more carefully.

If language has become an interface to computation, then perhaps we should use that interface to help people express what they are trying to achieve.

From there, systems can work out which capabilities are relevant.

Some of those capabilities may use models.

Many will not.

An identity system can establish identity.

A registry can establish an authoritative fact.

A rules engine can evaluate a deterministic rule.

An API can retrieve information.

A payment platform can transact.

A workflow can maintain state.

A professional can exercise authority.

A human can make a judgement.

And a model can help understand, coordinate, interpret, explain and reason around all of those capabilities where its strengths are useful.

The architecture begins to look less like this:

Language → Model → Everything

and more like this:

Language → Intent → Context → Capabilities → Work → Judgement → Outcome

The model remains important.

It just stops pretending to be the institution.


This changes the enterprise AI question

For the last few years, much of the enterprise discussion has effectively asked:

What work can we give to AI?

Perhaps the better question is:

What is someone trying to achieve, and what combination of computational, institutional and human capabilities should respond?

That is a very different way to think about enterprise architecture.

It begins with the situation rather than the application.

It begins with intent rather than the interface.

It treats existing institutional capability as something to discover and assemble, rather than something that should automatically be recreated inside a model.

And it preserves a distinction that I think will become more important as AI improves:

Reasoning is not authority.

An AI system may be able to identify what appears to be the best course of action.

That does not necessarily mean it has the authority to take it.

Nor does greater reasoning capability mean that human judgement has somehow disappeared.


So, is AI easier to weaponise than to institutionalise?

Perhaps.

But not for the reason the phrase first suggests.

AI does not inherently favour attackers.

And enterprise AI is not doomed to disappoint.

The asymmetry sits somewhere else.

A bad actor can sometimes extract value directly from computational capability.

An institution has to transform computational capability into something that can operate inside structures of identity, authority, policy, provenance, responsibility and judgement.

That is much harder.

And making the model more intelligent does not automatically solve it.

That may explain why two apparently contradictory things can happen at once.

AI can become more capable.

Its potential for misuse can increase.

And enterprises can still struggle to turn that same capability into dependable operational value.

There is no contradiction if the underlying problems are different.

The attacker asks whether AI can act.

The institution must ask whether it should act, whether it may act, and who answers for it when it does.

That is a much higher bar.


The next breakthrough may not be a bigger model

There will be better models.

There will be more compute.

There will be better reasoning.

There will be better agents.

There will be extraordinary things built with them.

But enterprise transformation may ultimately depend on something less spectacular and more difficult.

Architecture.

How do we connect language to intent?

How do we understand enough of a situation without unnecessarily centralising everything known about a person or organisation?

How are relevant capabilities discovered?

How do those capabilities work across organisational and jurisdictional boundaries?

How is authority established?

How is provenance retained?

How does an actor retain visibility and control?

Where does deterministic computation end and probabilistic reasoning begin?

Where does human judgement remain necessary?

And how do we preserve responsibility as more surrounding work becomes computational?

These are not only model questions.

They are institutional questions.

And perhaps this is why enterprise AI can feel disappointing when compared with its demonstrations.

We may be measuring one kind of capability and expecting another.

The pilot proves computational capability. Production demands institutional capability.

The model does not need to become the enterprise.

It needs to know how to work with one.


References

[1] Australian Signals Directorate, “Frontier AI models and their impact on cyber security.” https://www.cyber.gov.au/about-us/view-all-content/news/frontier-models-and-their-impact-on-cyber-security-update

[2] Bank for International Settlements, “A Mythos moment? Frontier AI and cyber risk,” BIS Bulletin No. 129, July 2026. https://www.bis.org/publications/bulletin-129-mythos-moment-frontier-ai-and-cyber-risk

[3] Bank of England, Financial Stability Report, July 2026. https://www.bankofengland.co.uk/financial-stability-report/2026/july-2026

[4] Anthropic, “Detecting and countering misuse of AI: September 2026.” https://www.anthropic.com/threat-intelligence-report-september-2026

[5] NIST National Cybersecurity Center of Excellence, “Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization,” concept paper, February 2026. https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd

[6] Gartner, “Why 50% of GenAI Projects Fail — And How to Beat the Odds.” https://www.gartner.com/en/articles/genai-project-failure

[7] McKinsey & Company, “The State of AI: Global Survey 2025.” https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai

[8] Forrester, “The State Of Agentic AI In 2026: Companies Are Chasing, Few Are Catching.” https://www.forrester.com/blogs/the-state-of-agentic-ai-in-2026-companies-are-chasing-few-are-catching/

[9] Deloitte Australia, “Australian organisations lag global peers in realising AI’s transformational potential,” press release on State of AI in the Enterprise, 2026. https://www.deloitte.com/au/en/about/press-room/australian-organisations-lag-global-peers-realising-ai-transformational-potential-120226.html

[10] METR, “Measuring AI Ability to Complete Long Software Tasks,” March 2025; see also METR time-horizon methodology and limitations. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

[11] Microsoft Research, “Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead,” MSR-TR-2025-16, March 2025. https://www.microsoft.com/en-us/research/publication/inference-time-scaling-for-complex-tasks-where-we-stand-and-what-lies-ahead/


Dakshan Pothuhera
Founder, DataMPowered®

DataMPowered explores Operational Intelligence, Judgement Governance™ and the architecture required to connect AI, institutional capability and human judgement.

© 2026 DataMPowered Pty Ltd. All rights reserved.

This research informs how ifCEM supports governed work in the public pilot, with reviewable workflows designed for accountable adoption in organisations.

Explore the ifCEM public pilot →

Want to discuss how ifCEM could support your organisation? Let's talk.

Start a conversation

Related research

Is AI Easier to Weaponise Than to Institutionalise? | DataMPowered