When AI models take actions, read production databases or write to third-party APIs, standard software security checks miss the risk. We audit the data lineage, the model evaluation set, the agent permissions and the human-in-the-loop controls.
We map the whole path rather than testing the model on its own, because in practice the damage happens at the joins and at the point where something acts.
Every model, copilot, agent and API call in production, found through code, cloud accounts and expense records rather than by asking people to self declare.
For every agent that acts: what it can read, what it can write, what it can trigger, and what the worst outcome is if someone talks it into doing the wrong thing.
Where the model came from, what it was trained or tuned on, and whether you hold the rights to use it the way you are using it.
Every route a customer record can take once it enters a prompt: into the index, into logs, into a third party processor, into a jurisdiction you did not intend to be in.
Prompt injection through content the system reads, jailbreaks against the guardrails, tool permission abuse, and poisoned or unpinned dependencies.
How your teams actually build and change these systems: approval before deployment, change control on prompts and thresholds, and incident logging.
That refusal is itself a finding and we write it as one. We then assess what can be established from outside, and state plainly what the gap means for the risk you are carrying and who is holding it.
Tier 3 includes adversarial testing, but the engagement is wider. A penetration test asks whether someone can get in. This asks whether the system should be trusted with the decision you have handed it.
Usually harder. Access is better and documentation is worse, and the evaluation set is often missing entirely because the team was shipping under pressure.
Not on the same engagement. Any connected firm that could do remediation is disclosed on page one, and the assessor takes no part in that work.
Discovery is the first week of an AI governance assessment, and it is usually the week that changes the conversation.