You can't run your business on AI that's governed like a side hustle.
Most of what's sold as AI governance is theatre: an ethics board, a set of principles, a policy document, a workshop for the leadership team. The board shrugs and approves, lots of people talk about bias, HR create a community page, and it all melts on contact with the actual technology. A business that barely knows what it has deployed gives governance no grip. So a framework full of "shoulds" meets a team that deals in specifics innovating with a technology that evolves daily. The result is foreseeable.
So what is AI governance?
AI governance is the set of controls, evidence and decision rights that make an organisation's use of AI deliberate and accountable. It is operational, and it's based on confidence and understanding. You cannot govern what you cannot evaluate, you cannot evaluate what you do not measure, and you cannot provide useful measures for a system that you don't really grasp.
87% of executives claim to have clear AI governance frameworks, but fewer than 25% have fully implemented the technical tools to manage risks.
Source: IBM Newsroom
A framework is a structure, and a structure is easy to approve. But unconnected to technical implementation it's window dressing. A governance function that only has the ability to advise means the delivery timeline decides what ships. This is why AI governance frameworks fail. Real governance is the work of making your organisation's AI practice intentional.
"Should" is not a control. It never will be.
Impact
Everyone is deploying. Almost no one has it running.
This isn't a failure of intent. Most organisations have the framework, the principles, the committee. But it stalls when governance has to leave the document and start running inside the technology: continuous evaluation, instrumentation, controls that live in the stack.
That is a different order of difficulty and well outside the comfort zone of most AI governance programs. It is the edge of what they can currently execute, not the edge of their intent
AI doesn't wait for them to catch up. The gap widens with every model pushed to production without triage, every vendor claim taken on trust, every system running that no-one is watching, every new model release.
are already using or plan to use agentic AI this year
Source: Cloud Security Alliance & Google Cloud
<28%
of respondents are confident they can secure AI used in core business operations
Source: Cloud Security Alliance & Google Cloud
<1%
of organisations have fully operationalised responsible AI
Source: World Economic Forum
The work product
What operational governance is made of.
Governance requires instrumentation that fits the reality of AI innovation,
the profile of current AI, and its evolving risk surface.
None of this is a slide deck or a policy pack. Each is a working component that runs
inside your systems, tuned to your models and your risk. Built to your environment,
not lifted from a template.
01
Tiered AI Impact Assessment
Scrutiny scaled to consequence: heavy where it matters, light where it doesn't.
A triage that sorts every use case by what happens if it goes wrong, so a customer-facing credit model and an internal meeting-notes summariser don't carry the same paperwork. Each tier sets its own decision rights, escalation path, and evidence bar. A separate procurement track puts vendor "responsible AI" claims through real due diligence before the contract is signed, not after the incident.
02
AI Bill of Materials
A live inventory of every model, dataset, prompt, MCP, and agent in production, each with an owner.
Most large organisations cannot name every model they have running, let alone the prompts, agents, and datasets feeding it. This is the discovery exercise that finds them, then tracks lineage, provenance, and dependencies as they change. It is the difference between an estate you manage and one you hope is behaving.
03
AI Risk Taxonomy
A risk model drawn from your actual systems, mapped to the standards you'll be audited against.
Generic risk registers list harms that could happen to anyone. This one is built from your own models, agents, and data flows, then mapped against ISO 42001, NIST AI RMF, the EU AI Act, and OWASP CycloneDX so every risk traces to a specific system and a specific control. It is the artefact that lets your risk committee and your engineers point at the same row and mean the same thing.
04
Reusable Guardrails Library
A versioned library of controls: input filters, output checks, and policy enforcers, tested and ready to drop in.
Without a shared library, every team rebuilds the same input filters and output checks slightly differently, none of them audited, none reusable. This collects them as versioned, tested controls any team can pull into a deployment, and gives the Bill of Materials something real to point at. As it grows it becomes the shared standard a community of practice forms around, rather than a policy nobody reads.
05
Evaluation & Observability Pipeline
Continuous evaluation, tracing, and drift detection wired into your CI/CD, running on every change.
Before a model ships, evaluation suites test it for accuracy, bias, and capability limits. Once it is live, tracing and drift detection watch its real behaviour and route a failure into your incident process the same way a Sev-1 outage would. The audit trail a regulator asks for is generated automatically, as a by-product of running the thing.
Frontier06
Agentic Constraint Architecture
Decision-traceability, hard constraint layers, and red-teaming for systems that act without a human in the loop.
When an agent can place an order, send a message, or change a record on its own, a wrong step stops being a wrong answer and becomes a wrong action that has already happened. This bounds what an agent is allowed to do, traces why it did what it did, and red-teams it for failure modes a conventional risk framework never anticipated. The constraints have to hold at machine speed, because nobody is reading along.
Each of these runs continuously, and each produces a decision that lands on someone's desk: a drift alert, a model nobody had catalogued, an action just blocked in a live system. They arrive at commit speed; governance runs at committee speed. And there is nowhere for them to land.
How this plays out
Insight you can't act on changes nothing.
The instruments produce insight: what each model is doing, where it is drifting, which risks are live. But insight changes nothing if no one can interpret it, decide on it, and act in time. So governance takes more than the instruments. It takes an organisation that can use them: reason about the risk, decide what is acceptable, and act while it matters.
That capability is rarely already in place, and it is not an ethics board or a set of principles. It is an intentional uplift: a shared way to reason about AI risk, the authority to decide where the work happens, and language that carries risk from the engineer to the board. Each engagement below started with the recognition of missing comprehension, not missing tools.
01
Scenario:
A regulator had asked how an AI decision was reached. The answer didn't exist.
The signals to answer were already there, scattered across logs, dashboards, and tickets. But every conversation was about bias, fairness, and transparency: real concerns, and the ones the organisation knew how to hold. No one was turning the fragments they already had into a plain account of what the system was doing and whether it was working. What was missing was not another tool. It was the comprehension to know which signals mattered, and what they were saying.
Once they could read their own signals, there was something real to govern.
02
Scenario:
The board was asked to set risk appetite for AI. No one in the room could.
The people being asked to set appetite had no way to reason about what they were setting it for. They had sat through training that left them more confident and no more capable: specialists explaining their specialisms, none of it joined up, none of it pitched where a board actually decides. Appetite was impossible because the risk could not be reasoned about above the engineering floor. The work was building that shared comprehension: enough to ask precise questions, weigh the trade-offs, and decide without deferring to the people they were there to govern.
Once the board could reason about the risk, it could set its own appetite.
Frontier
03
Scenario:
The systems were already past where the field had answers.
This team was strong, and its practice was already deliberate. Its questions sat at the real frontier: agentic systems that act on their own, multiple agents coordinating, models that change their own behaviour, with failure modes no one has charted yet. The governance specialists they had brought in could not follow. The work was the architecture that constrains what such systems are allowed to do, traces why they did it, and red-teams them for failures no conventional framework anticipates, because here there is no established practice to lean on.
The result was control over systems no standard yet covers.
About
From neural networks research to international AI standards.
Work like this only holds when several disciplines meet in one person: safety research at the
frontier, a hand in writing the standards, the hard-won craft of operationalising governance inside
large regulated organisations, and the change practice that decides whether any of it gets adopted.
Most people have one. The engagements that fail, fail at the seams between them.
Fifteen years in AI and machine learning, a decade specifically in AI governance and safety, and more than thirty enterprise engagements across banking, insurance, energy, healthcare, and telecom.
"I help executives turn a vision into reliable assets, and engineers turn experiments into trusted systems."
That's what makes this work. A risk taxonomy that maps to your actual technology. Governance
requirements that your engineers can implement without interpretation. Board assurance that is
technically honest, and a board who are confident in having those conversations.
At SingularityNET I built safety frameworks for frontier AGI systems. At nib I'm
operationalising ISO 42001 and NIST AI RMF through governance boards, LLM evaluation pipelines,
and agentic observability infrastructure, working with engineers on exactly how
guardrails land in their stack and with the risk committee on what that means
in terms they can act on.
You can't assert one correct answer for a system that's never the same twice. What AI agent evals are, the three ways to judge quality, and why you grade the path, not just the output.
Why Your Monitoring Stack Can't See an Agent Failing
An AI agent can return a clean 200, pass every health check and still be wrong. Why agent reliability needs observability and evals, not just uptime dashboards.
AI Audit Trail vs Logging: Why Logs Aren't Evidence
A log proves an AI system ran. An audit trail proves it was governed. Why technical logs do not meet AI compliance requirements, and what to capture instead.
Audit Trails for AI Agents: Reconstructing What an Agent Did
An agent's audit trail has to reconstruct each step it took, not just its final output. What to capture for systems that act without a human reading along.
Audit Committees Are Asking the Wrong AI Questions
Audit committees know how to ask whether AI is compliant and whether vendors are managed. The questions that create real accountability - the error rate on decisions that affect customers, whether any decision can be reconstructed, who can halt a deployment — rarely get raised.
Should Is Not a Control: How AI Ethics Built Its Own Graveyard
The discipline that was supposed to prevent AI harm produced frameworks, principles, and declarations - then largely watched as the industry did whatever it was going to do anyway. The same pattern is repeating in AI governance right now.
Using large language models to evaluate large language models introduces a class of systematic bias that most evaluation pipelines are not designed to detect.
The Inference Audit Trail: Making Every AI Decision Accountable
What is an AI audit trail, and why are logs not enough? How to build audit trails for AI models that reconstruct decisions for compliance and governance review.
Hard constraints, soft constraints and approval gates: why enterprise AI agent safety belongs at the tool level, not the model level, and how to build it.
Ethical debt in artificial intelligence refers to the accumulated cost of unresolved ethical issues that arise during a system's design, development, and deployment.
Human Resources: Empowering the organisation to make responsible-use decisions
The rapid advancement of Artificial Intelligence (AI) and other emerging technologies has created new opportunities and challenges for businesses, requiring them to adapt and evolve.
Emerging technologies, particularly artificial intelligence (AI), machine learning (ML), and generative AI, present both opportunities and challenges for businesses across industries.
Generative AI offers significant opportunities for business transformation through automating content generation, enhancing creativity, and unlocking new revenue streams.
Thank you for providing to the people of Australia an opportunity to respond to key questions regarding Australia's AI Strategy in the form of the AI Action Plan Consultation Paper.
This reality of AI tech adoption in our private and public institutions struggles to fit into frameworks created to promote responsible and ethical use which focus heavily on the development process as the locus to effect positive outcomes.
MLOps, the big-bet for scaling Machine Learning, promises seamless development through to in-life use of models using automated, DevOps CI/CD workflows, but what does this mean for an AI Ethics discipline that has focused its fire-power on single-shot development projects?
If you're a CRO, CISO, CDO, or board member trying to get ahead of AI risk, rather than
catch up to it, I'd like to hear from you.
That might be building your first governance framework, hardening an enterprise-scale agentic
deployment, stress-testing what your engineers have already built, or preparing your leadership
team for what's coming. Bring the specific problem, and I'll tell you directly where
I can help.