I gave a talk a few weeks ago to the internal audit teams of PLDT and the Gokongwei Group. Not a crowd I expected to find myself in front of. But the more AI systems we deploy into real institutions, the more I find myself caring about what happens to them once they leave our hands.
When the Model Meets the World
Let me start with something from complex systems. There is a phenomenon called emergence. Thousands of starlings moving together in the sky, forming shapes that look almost coordinated. No single bird is directing the flock. Each bird is just responding to the birds immediately around it. From those local interactions, something coherent appears at the level of the whole; something no single bird planned, and no single bird could have produced alone.

You see the same thing in traffic, in social networks, in organizations, and increasingly, in AI systems interacting with people and institutions.
We saw this in our urban mobility work for Singapore. We built a simulation of the Rapid Transit System that models human mobility down to the level of individual commuters. Government planners, investors, and contractors used it to test what-if scenarios: Would this decongestion strategy work? Would a new train line reduce pressure where it mattered most? If we changed one part of the network, what happens somewhere else?
That is the power of modeling complex systems. You get to test decisions before they become expensive, physical, and hard to reverse.
But the moment a model starts informing real decisions, a new set of questions appears.
What if the assumptions are outdated? What if commuter behavior has shifted since the model was trained? What if the model performs well on average, but fails for a specific station, a specific community, a specific peak-hour condition? What if the investment is technically justified but doesn’t deliver the value promised?
I carried those same questions into AI. And on a meta level, this is emergence in practice: local behaviors, assumptions, and incentives producing outcomes no one person fully controls or anticipated.
The builder’s intention and the reality of deployment are two different things. Deployment drift, changing behavior, edge cases, accountability, and consequences that only appear once the system is out in the world interacting with real people. That is not because builders are careless. It is because no team can truly see the whole system from inside the building process.
That is when audit started making real sense to me.
The shift in posture
Most people I know, including my own AI scientists and engineers, hear the word “audit” and feel something between inconvenience and dread. Scrutiny. Slowdown. Someone finding out what you’d rather keep hidden.
I have been in rooms where there’s a shift in posture when the internal audit team arrives. I understand the reaction because I was not immune to it myself.
But I have come to think that this reaction is a risk signal. When an organization treats scrutiny as something to avoid, it often means speed has become more important than accuracy. The cost of being examined starts to feel higher than the cost of being wrong. In complex systems, that is a very expensive calculation.
At ECAIR, before any high-risk AI system moves too far from idea to deployment, we stress test as much as we can. No regulation requires it. We do it anyway, because builders who take deployment seriously should. We bring in people with no stake in our success: AI ethics practitioners, governance specialists, technically sharp AI scientists from outside the country who have no reason to tell us what we want to hear.
We ask them to look at our process, assumptions, algorithms, even our data pipeline, and to tell us where they see problems. We do it because the earlier someone challenges the assumptions, the less expensive it is to correct the system.
To me, the most valuable reviewer is not the one who confirms we got it right. It is the one who helps us see what we missed.
That is the relationship I want more of between builders and auditors: not opposition, but partnership. The builder’s job is to make the system work. The auditor’s job is to help make sure it continues to work, works as intended, and works without creating unacceptable harm.
What comes next
The technologies entering organizations now are only becoming more complex. AI is being embedded into hiring, lending, medical decisions, and infrastructure management. Quantum computing is approaching, and most of the attention is on what it will enable; less attention is being paid to what it may undo. A significant portion of the encryption we rely on today is based on mathematical problems that quantum systems may eventually solve quickly. This is a known trajectory, by the way.
The audit function that will matter in the next decade needs enough technical fluency to ask the right questions about these systems. Not to become AI engineers or quantum physicists. But to know enough to know where the limits are; to ask better questions, and to bring in the right people when those limits are reached.
If you are in audit and reading this, here is what I’d suggest.
Map which AI systems in your organization affect real decisions about real people, such as hiring, credit, access, and performance evaluation. These are the systems that warrant the closest scrutiny, and where governance gaps are most likely to carry real human cost.
Ask three questions for each one: Who owns it? When was the model last reviewed against actual outcomes? Is there a documented process for what happens when it fails?
If any of those answers are unclear, that is where to start.
You don’t need to audit the algorithm to audit the governance. Just know whether anyone is watching, what they are watching for, and what they are empowered to do when they see something wrong.
The more seriously we take the work of building systems that operate in the real world, the more seriously we take the question of how those systems are checked, challenged, and corrected.




