On-demand WEBINAR

The Agentic Assurance Shift: How MindBridge and Holistic AI are Advancing AI Governance and Testing | Vision 2026

Raj Patel

VP of AI Transformation Holistic AI

Wenzel Reyes

Head of AI Governance and Industry Engagement MindBridge

Raj Patel

VP of AI Transformation

Wenzel Reyes

Head of AI Governance and Industry Engagement

As AI moves from generating answers to taking actions, organizations need to govern more than the final output. Wenzel Reyes of MindBridge and Raj Patel of Holistic AI examine how agent authority, human oversight, step-level testing, adversarial testing, and independent assurance can help organizations evaluate increasingly agentic AI systems. 

Key Learnings

  • How agentic AI changes governance by introducing greater autonomy, delegation, and multi-step decision-making.
  • The five elements of meaningful human oversight: information, authority, competence, time, and accountability.
  • How step-level, adversarial, and independent testing can reveal risks that output testing alone may miss.

Session Summary

As AI moves from generating answers to taking actions, finance and audit teams face a different governance challenge. The central question is no longer only whether an AI model produces the right output. Organizations also need to understand what an AI agent is allowed to do, what information it can access, when it should stop or escalate, and how its actions can be independently tested and reconstructed.

In this Vision 2026 session, Wenzel Reyes, Head of AI Governance and Industry Engagement at MindBridge, and Raj Patel, VP of AI Transformation at Holistic AI, examine how governance and assurance need to evolve as organizations adopt more agentic AI workflows.

Reyes begins by distinguishing agentic AI from the generative AI workflows many finance and accounting teams already use. Rather than simply responding to a prompt, an agent can pursue an objective across multiple steps, retrieve information, use tools, evaluate intermediate results, interact with other systems, and initiate actions. As that authority increases, he argues, organizations need greater confidence in how those decisions are delegated and governed.

Using a revenue audit example involving thousands of contracts, Reyes illustrates how multiple agents could extract contract terms, evaluate risk indicators, retrieve accounting guidance, score contracts, and determine which items should be escalated to an auditor. That creates three central governance questions: what the agent is allowed to decide, when the agent must stop and escalate, and whether the auditor can reconstruct the path the agent took.

The discussion then turns to human-in-the-loop oversight. Reyes argues that simply adding a human reviewer does not automatically create meaningful governance. Effective human oversight depends on five elements: information, authority, competence, time, and accountability. Reviewers need to understand what happened, have the authority to challenge or override the work, possess the necessary technical and professional knowledge, have enough time to perform a meaningful review, and know who ultimately owns the decision.

Patel builds on this by explaining that human review must be deliberately designed into an AI workflow. A review step that gives someone insufficient information or time can become a weak control rather than an effective safeguard. The goal is not to place people at every step, but to determine where human judgment is most valuable and where automated validation can provide reliable checks.

The session then examines how AI agents should be tested. Traditional model testing can often compare an input and output against an expected result. Agentic systems are more complex because an agent can plan, retrieve data, generate content, check its own work, retry failed steps, and interact with other tools before producing a final result. Patel explains that these individual components need to be evaluated differently rather than treated as a single pass-or-fail test.

For example, retrieval tasks can be measured against labeled data, while open-ended generation requires predefined scoring criteria and benchmark data created by qualified practitioners. The path the agent follows also matters. Two agent runs may produce the same correct answer even though one followed the expected process and another repeatedly failed, retried, or accessed information outside the intended scope.

Patel outlines an AI assurance methodology that includes defining the system’s scope and threat model, reproducing and measuring its behavior, performing adversarial testing, and establishing monitoring and release gates. Adversarial testing deliberately pushes an AI system beyond its intended boundaries to identify weaknesses that conventional testing may miss.

The depth of assurance should also reflect the authority and potential impact of the AI system. An agent that only reads information and drafts content may require documentation and logging. An agent that calls tools or writes records requires stronger tracing, access restrictions, and review. A supervisory agent that directs other agents or can materially affect systems requires deeper testing, monitoring, and human involvement.

The session also details how MindBridge and Holistic AI are evolving independent testing of MindBridge AI systems. Holistic AI has historically conducted annual audits of specified MindBridge control points. The organizations are moving toward multiple testing points during the year so that independent testing can more closely align with changes to those control points.

MindBridge can initiate a testing run, but Holistic AI retains control of the testing methodology, evaluation criteria, thresholds, and final opinion. Results are not available for external reliance until Holistic AI has reviewed and signed off on them. Patel emphasizes that Holistic AI does not design, develop, or operate the MindBridge systems it audits, and its compensation is not tied to the outcome of the testing.

A live demonstration shows how Holistic AI visualizes the actions taken by a multi-agent system, including the sequence of tools, sub-agents, checks, and interactions involved in completing a task. Patel then demonstrates adversarial testing designed to identify weaknesses such as hallucinations and jailbreaking attempts.

The session closes by examining what agentic AI means for the future of audit and assurance. Reyes and Patel distinguish between repetitive, lower-judgment work that may be appropriate for agents and areas where professional judgment, skepticism, supervision, and accountability should remain with auditors. As agents take on more work, assurance must evaluate both what an agent produces and how the agent itself operates.

The central takeaway is that organizations do not need to choose between automation and human judgment. Effective agentic assurance depends on deliberately defining authority, testing AI systems at the appropriate level of risk, preserving independent verification, and placing professional judgment where it matters most.

Chapters

01:25 The agentic assurance shift
03:29 Trust and assurance as AI gains authority
06:45 How an agentic audit workflow works
08:07 Three governance questions for AI agents
11:42 Designing effective human-in-the-loop oversight
15:47 AI governance as an operational discipline
18:23 How to test an AI agent
21:37 Different failure modes in agentic systems
24:59 Scope, measurement, adversarial testing, and monitoring
26:55 Matching assurance to AI autonomy and impact
29:34 Evidence and deliverables from AI assurance
35:47 Moving toward more frequent independent assurance
42:49 Demo: agent graphs and adversarial testing
50:26 Q&A: independence, agent behavior, and audit evidence
56:03 Key takeaways for governing AI agents

On-demand webinar

The Agentic Assurance Shift: How MindBridge and Holistic AI are Advancing AI Governance and Testing | Vision 2026

Please enter your email to proceed

Library item

The Agentic Assurance Shift: How MindBridge and Holistic AI are Advancing AI Governance and Testing | Vision 2026

Please enter your email to proceed

Library item

The Agentic Assurance Shift: How MindBridge and Holistic AI are Advancing AI Governance and Testing | Vision 2026

Please enter your email to proceed