Applied AI Governance: From Requirements to Systems That Work

Principles are the easy part. The difficulty is what happens between a stated commitment and a system that actually honours it
There is no shortage of AI principles. Governments have published them, companies have adopted them, international bodies have convened around them, and the Church has contributed its own through the Rome Call for AI Ethics and the reflection gathered in Antiqua et Nova. The vocabulary is now broadly shared: transparency, accountability, fairness, human oversight, privacy, reliability.
The convergence is genuine and worth something. But a gap has opened that principles alone cannot close. Organisations can affirm every one of these commitments and still deploy systems that harm people, and often the affirmation and the harm coexist without anyone acting in bad faith.
This is the practical question now facing anyone serious about ethical technology. Not what should the principles be, but what has to happen for a principle to become a property of a working system.
Where the Translation Fails
Requirements are written in the language of values. Systems are built in the language of specification. Something is lost in the passage between them, and the loss is not accidental.
Take transparency. As a principle it is unambiguous. As a specification it fractures immediately. Does it mean publishing a model card describing training data and known limitations? Logging every decision for later audit? Telling an affected person that a system was involved in their case? Explaining why their particular application was declined, in language they can act on? Opening the model to independent inspection?
These are five different engineering commitments with five different costs, and an organisation can implement the cheapest one, describe itself as transparent, and be telling a kind of truth. The person refused a loan at midnight still has no idea why.
The same fracture appears everywhere. Fairness across which groups, measured how, at what cost to which other group? Human oversight by a person with what training, what authority to override, and how many cases per hour? Accountability to whom, through what mechanism, with what remedy?
A requirement that has not been resolved into a specification is not yet a requirement. It is an intention. And intentions do not constrain systems.
Governance as Gate, Governance as Design
The second failure is one of timing.
In most organisations, ethical review arrives near the end. The system has been built, the budget spent, the launch date set, and the review is a checkpoint to clear. In that position, a reviewer has one option: approve, or become the person who killed a project everyone has worked on for a year. The structural pressure is overwhelming, and it does not require anyone to behave badly for the outcome to be predictable.
Governance placed at the end can only veto, which means in practice it can only approve.
Governance placed at the beginning does something different. It shapes what gets built. It asks, before the architecture is chosen, whether this system should exist, who it will affect, what data it requires, and what happens to the people it gets wrong. Those questions are cheap to answer in a design document and enormously expensive to answer after deployment.
This is not a plea for slower innovation. It is the observation that ethical constraints are like structural ones: identified early, they shape a design elegantly; identified late, they require demolition.
The Measurement Trap
There is an old management maxim that you cannot manage what you do not measure, and AI governance has embraced it enthusiastically. Fairness metrics, bias audits, performance dashboards.
The maxim is half right, and the missing half causes real damage.
Some fairness definitions are mathematically incompatible with one another. A system cannot simultaneously equalise false positive rates across groups and maintain equal predictive value across groups, except in cases that do not occur in practice. Choosing which fairness definition to optimise is therefore a moral decision disguised as a technical one, and it is frequently made by whoever wrote the evaluation script.
More broadly, what is measurable tends to displace what matters. A system can pass every fairness audit and still be deployed in a context where its very existence is an injustice. A model can be accurate and well-calibrated while the decision it automates should never have been automated at all.
Metrics are necessary. They are not sufficient, and treating them as sufficient produces governance that is rigorous about the wrong things. The question of whether a system should exist cannot be answered by measuring how well it performs.
The Accountability Gap
When an automated system causes harm, responsibility tends to dissolve.
The developer implemented a specification. The product manager followed a business requirement. The executive relied on the assurance of a compliance team. The compliance team reviewed a document produced by the developer. The vendor points to its terms of use. Each account is individually plausible, and collectively they add up to nobody.
This is the deepest problem in applied governance, and it is not solved by another framework. It is solved by naming people.
A system that materially affects a person’s livelihood, liberty, health, or access to services should have a named individual who is answerable for it. Not a committee, not a function, a person, with the authority to halt the system and the obligation to explain it. Distributed responsibility is a synonym for no responsibility.
The Catholic tradition is unusually clear on this point, and its clarity is useful even to those who do not share its premises. Conscience is exercised by persons. It cannot be held by an institution, delegated to a process, or satisfied by a policy. An organisation that cannot say who is answerable for a given system has not delegated responsibility. It has abandoned it.
The Half That Is Missing
Nearly all AI governance frameworks concern themselves with what happens before deployment. Very few say anything serious about what happens after.
But systems fail. They fail in ways their designers did not anticipate, in contexts they were not tested in, against populations that were not represented in evaluation. This is not a sign of negligence. It is the ordinary condition of complex systems meeting the real world.
The measure of a governance regime is therefore not only whether it prevents foreseeable harms. It is whether the organisation can detect harm it did not foresee, and respond once it does.
That requires concrete capacity: a channel through which affected people can report that something went wrong, and evidence that reports are read. Monitoring that would actually surface a problem rather than confirming that performance metrics remain steady. Authority to suspend a system quickly, held by someone who will not be punished for using it. A commitment to remedy, meaning that people harmed are made whole rather than merely apologised to.
Subsidiarity matters here more than anywhere. The people best positioned to notice that a system is failing are the people it is failing. If they have no route by which to say so, and no reason to believe saying so will change anything, the organisation has blinded itself and called it stability.
What a Working System Looks Like
Pulling this together, applied governance means something reasonably concrete.
Requirements are resolved into specifications before building begins, so that transparency names a particular artifact and fairness names a particular definition, chosen deliberately and documented as a moral choice rather than a technical default.
Review sits at design time, where it shapes architecture, rather than at launch, where it can only bless.
Metrics are used but not trusted alone, and the prior question of whether a system should exist is asked separately from how well it performs.
Every consequential system has a named person answerable for it, with real authority to stop it.
Affected people have a genuine route to contest a decision and report a failure, and that route is monitored by someone who can act.
Failure is planned for rather than merely feared, with detection, suspension, and remedy defined before deployment rather than improvised during a crisis.
None of this is technically difficult. It is organisationally difficult, which is a different thing, and harder.
The Test
The distance between principles and practice is where ethics either becomes real or reveals itself as decoration. An organisation’s published commitments cost nothing. What it does when honouring them would be expensive, inconvenient, or embarrassing is the whole of its actual ethics.
This is not cynicism about principles. Principles matter, and the international convergence on them is a genuine achievement worth defending. But a principle that never becomes a specification, never constrains a design, never names a responsible person and never survives contact with a deadline has not governed anything.
The work now is unglamorous. It is specification documents, escalation paths, named owners, contestation channels and incident procedures. It looks nothing like the philosophical debates that shaped the principles, and it is where the human dignity those principles were written to protect is either honoured or quietly traded away.

