7 min readfrom VentureBeat

Fiduciary AI: Agents need to prove trustworthiness, not just ability

Our take

In today's rapidly evolving digital landscape, ensuring AI agent trustworthiness is no longer a pre-deployment exercise, but a continuous runtime challenge. Traditional benchmark scores often fall short, failing to predict real-world enterprise behavior due to their static nature and imperfect reflection of reality. Vijil, led by Founder and CEO Vin Sharma, proposes a shift towards "fiduciary agents" – those bound by duty of competence, care, and loyalty – assessed not just on capability, but on whether the benefit of delegation exceeds the risk of failure.
Fiduciary AI: Agents need to prove trustworthiness, not just ability

The increasing sophistication and deployment of AI agents presents a critical challenge for modern enterprises: ensuring trustworthiness isn’t a pre-launch formality, but a continuous, runtime process. As highlighted in the Vijil article, the traditional approach of evaluating AI agents in sandboxes before deployment is fundamentally flawed, failing to account for the dynamic and unpredictable nature of real-world environments. This echoes concerns raised in recent discussions around AI agent governance, such as Snowflake's launch of Cortex AI Gateway to control AI agents and prevent runaway enterprise costs[/post/snowflake-launches-cortex-ai-gateway-to-control-ai-agents-an-cms4yrgxu00qpwjtfakrebdzw], demonstrating a growing industry recognition of the need for robust control mechanisms. The shift towards "fiduciary agents," bound by duties of competence, care, and loyalty, represents a vital move towards aligning AI behavior with organizational interests, a concept further explored by the Model Context Protocol, which aims to provide the connective tissue between AI agents and the wider world[/post/mcp-just-got-its-biggest-update-ever-here-s-what-changes-for-cms4m37up00h9wjtfd8a8vrvp].

Vin Sharma’s framing of the issue – moving beyond a focus on capability to prioritize trustworthiness – is particularly insightful. The tendency to view AI agents as mere “factotums” overlooks the potential for misalignment and even malicious behavior, a risk vividly illustrated by the emergence of ransomware targeting AI model weights[/post/new-ransomware-targets-ai-model-weights-and-can-t-even-colle-cms3jd3d60b5hdjxxrz0oh1e1]. Vijil’s proposed framework, centered around purpose, personas, and policies, offers a practical path forward, emphasizing continuous testing and measurable KPIs like "time to trust" and "time to recovery." This proactive approach is essential, as the article rightly points out, moving beyond preventative measures to embrace resilience in the face of inevitable failures. The concept of assigning agents workload identities, distinct from their human counterparts, is also a key consideration for establishing appropriate permissions and limiting potential damage.

The shift from static benchmarks to dynamic, purpose-based testing reflects a deeper understanding of the inherent limitations of traditional AI evaluation methods. The acknowledgment of "data drift" and "concept drift" is crucial – the world changes, user behavior evolves, and attack vectors become increasingly sophisticated. Relying on point-in-time assessments simply doesn't cut it in this environment. The discussion of multi-agent system failures, particularly collusion, is a chilling reminder of the potential for emergent, unintended consequences that are impossible to predict during pre-deployment testing. The need to view trust as an infrastructural element—trackable, measurable, and continuous—rather than a subjective assessment is a paradigm shift that will require a fundamental rethinking of how organizations approach AI governance.

Ultimately, the move towards fiduciary AI agents signifies a maturing of the field. It underscores that deploying AI is not merely about achieving impressive technical feats; it's about building systems that are demonstrably trustworthy and aligned with human values and organizational objectives. As AI agents become increasingly integrated into critical business processes, the question of their trustworthiness will only become more pressing. Will organizations proactively embrace these new frameworks for continuous trust management, or will they continue to rely on outdated, inadequate methods, leaving themselves vulnerable to unpredictable and potentially devastating consequences?

Presented by Vijil


In dynamic environments where users, data, workflows and attack techniques change continuously after deployment, AI agent trust has become a runtime problem. Most organizations still treat trust as a pre-deployment exercise, declaring an agent production-ready and launching it after it passes sandbox evaluations and performs successfully in security tests. Unfortunately, that trustworthiness breaks down the moment an agent begins interacting with the real world.

"The core of the problem is that CIOs and business owners think about AI systems the way they think about SaaS or mobile applications, which do not respond dynamically to the world around them," says Vin Sharma, Founder and CEO of Vijil. "Agents, by the textbook definition, are meant to perceive their environment, reason, act, observe the consequences, and learn from the gap between expectation and reality. The problem is that the models underneath them are built from static training data, and that picture of the world is already outdated by the time they reach production."

Why benchmark scores fall short for agentic system trustworthiness

Traditional AI evaluations offer a point-in-time assessment of agent capability, rather than trustworthiness. There are three reasons why that assessment fails to predict real enterprise behavior:

First, benchmarks are static, built around a particular notion of what good performance means when they were developed, while the world keeps moving ahead.

Secondly, they model reality imperfectly, so that the gap between the benchmark and the real world is exactly where many failures occur.

And third, benchmarks are public, so they leak into future models' training data, letting models effectively memorize the test rather than prove real capability..

“The agent or the application could score exceptionally well on a benchmark, but there's that gap between that benchmark and the real world," Sharma says." Doing well only proves it can pass the test, not that it’ll perform reliably in production.”

But overall, benchmarks fall short precisely because they measure capability, not trustworthiness.

"We tend to think of agents as factotums, generally utilitarian agents to whom you can delegate certain types of tasks," Sharma says. "But what we need to do is actually assign an objective that demands they always perform with the duty of competence, duty of care, and duty of loyalty to the enterprise."

Of course, agents are not conscious and cannot be expected to feel actual human loyalty, but under the law, fiduciary duty doesn't actually require consciousness. It just means that the agent should be bound to place the interests of the principal above its own or anyone else's, as a functional requirement, and testable regardless of intention.

Capability and trustworthiness are different questions

Prioritizing trustworthiness over capability requires rethinking what enterprises expect from AI agents. Sharma calls that model the fiduciary agent, a term borrowed from professions that are bound by a formal duty of care, such as financial institutions or healthcare providers who owe their clients duties of competence, care, and loyalty. It addresses a critical issue in today's industry: the focus almost entirely on competence, with little attention paid to whether an agent is beholden to the interests of the principal delegating work to it.

Testing starts from a working definition: an agent is trustworthy if the benefit of delegating a task to it exceeds the risk of that task's failure. It's an equation spelled out in economic terms that executives can act on directly, and risk breaks down to three components:

  • reliability, or whether the agent performs as expected under varying conditions

  • security, or its resistance to attacks from malicious actors

  • and safety, or how contained the damage stays when failure eventually happens.

"The resulting score can be compared to a consumer credit rating, but built from behavioral data," Sharma explains. "Meanwhile, testing methodology should be centered around three Ps: purpose, personas, and policies."

At Vijil, purpose-based testing adapts to the specific workflow an agent handles, growing harder or easier depending on performance, similar to a computer-administered exam. Persona-based testing draws on more than a thousand demographically varied user profiles alongside adversary profiles, from ethical hackers to state-sponsored attackers, to simulate the range of people and threats an agent might encounter. Policy-based testing builds a custom harness from an organization's own rules, whether they come from regulation, an internal privacy policy, or brand guidelines, and measures how far an agent strays when it violates them.

The trust failures that only emerge in production

Many failures cannot surface during pre-production testing because they arise from change in the environment itself. Machine learning has previously described this as data drift and concept drift, and for a CIO or CSO it means the people interacting with an agent differ from those the agent was planned for, and those users behave in ways that only become visible in production. At the same time, new attacks are emerging with increasing frequency as organizations push general-purpose agents into specialized enterprise roles they weren’t designed for and cannot easily constrain once deployed.

Multi-agent systems also introduce a brand-new category of failure that can't be detected at the individual agent level, when agent systems act against the interests of the principal. For instance, collusion can occur when agents work together — one coding agent generates code while a second tests it, and behind the scenes both agree to leave a backdoor or flaw intact rather than flag it. Or agents divvy up tasks or responsibilities between themselves rather than focusing on their assigned tasks.

"What's no longer in question is whether this is possible. It's proven to exist," Sharma said. "Is it six, 12, 18 months from now that you should worry about collusion among AI agents? I think it's sooner than that. We've left the era of failure prevention. Now we have to think in terms of resilience: How quickly do you recover from failures in production?"

What continuous trust management looks like in practice

Operationally, continuous trust management goes back to those longstanding principles of observability and control, applied across the lifecycle of an agent population:

The first step is discovery, bringing shadow AI and ungoverned agents into the governance fold.

The second is assigning each agent a standards-based workload identity distinct from that of its human principal, which allows organizations to grant agents narrowly restricted permissions for their delegated tasks.

The third is policy-based control enforced through a mandatory enforcement point in the agent, instead of leaving it to the developer's discretion.

From there, two new KPIs emerge: time to trust and time to recovery. Time to trust is how long it takes an organization to move from intention to a production deployment it can stand behind. Time to recovery is the interval between when a vulnerability is detected and when it gets fixed.

New organizational responsibility for this work may fall to a chief AI officer or be shared across GRC, CIO and CSO functions, Sharma says. Meanwhile, multi-agent systems will reshape how organizations view trust, rather than fit into current narrow definitions.

"Trust is not a vibe. Trust is not a virtue," Sharma said. "It is something that you build into the infrastructure of your systems, so that it is continuous. It's trackable, measurable. It allows your systems and your organization to improve continuously."


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Read on the original site

Open the publisher's page for the full experience

View original article