The Science of Agentic Trust: Moving Beyond AI Vanity Metrics with Pass@k

by Stephen Fishman
Published Jul 29, 2026

Key Takeaways

  • To build trust in autonomous AI agents, enterprises must move beyond vanity metrics and adopt mathematical reliability measures: Pass@k to confirm an agent can solve a problem (capability) and Pass^k to confirm it will solve it consistently (reliability/stability).
  • To maintain safe operations, Boomi’s framework uses the Decision Delegation Index (DDI) as a transmission setting to control the level of authority (or “leash”) granted to an agent, while the Policy Violation Rate (PVR) acts as traction control to detect deviations from compliance and standard procedures.
  • The goal is to reach the “Sweet Spot”—a state of high reliability and high delegation where agents operate with accountable autonomy, closing the gap between AI reasoning potential and actual business value.

In our previous post, “Closing the Impact Gap“, we explored how the Impact Gap shows up in agentic initiatives—as the distance between AI’s reasoning potential and its desired business value. To close this gap, enterprises must “harden” their AI agents, moving them from slow, effortful System 2 reasoning to fast, instinctive System 1 habits.

Quick reminder: In Daniel Kahneman’s classic book, “Thinking, Fast and Slow“, he called fast, automatic, and intuitive thinking System 1, and slow, deliberate, and logical thinking System 2. Many businesses today are trying to build and scale AI systems as though they were traditional deterministic IT systems designed to solve rote, predictable problems. But before you can trust agents to work on tasks quickly and automatically, you first need a systematic way of measuring how trustworthy your agents are. Only when an agent has proven itself trustworthy should it make the jump to working automatically with little or no human supervision. Once agents are proven to be trustworthy, you trust them with fast System 1 tasks.

How do you know when an agent is ready to make that jump? How do you measure “trust” in a system that is, by its very nature, probabilistic? That’s the question we’re going to explore in this blog post.

The answer isn’t found in vanity metrics like “total agents deployed” or “tokens consumed.” It’s found in the math of reliability as measured by two new metrics: Pass@k and Pass^k.

Pass@k from OpenAI and Pass^k by Anthropic have together blazed a trail for measuring agent reliability and establishing a reliability gate that can inform and enable trusted autonomy for agentic decision-making. These metrics serve as the analytical engine residing inside Boomi’s Agentic Evaluation Framework (available in our upcoming Agentic Score app — stay tuned!).

The “k” in each of these metrics is a critical variable—a sliding scale that provides an immediate indication of an agent’s trustworthiness.

Metric Technical Definition Business Meaning What does “k” mean?
Pass@k(pronounced “pass at k”) Measures the likelihood that an agent gets at least one correct solution in k attempts. Demonstrates that an agent can deliver the intended outcome As “k” rises, it means the agent needs more tries to get the preferred outcome.

1 is good, and 100 is not so good.

Pass^k (pronounced “pass power k”) Measures the probability that all k trials succeed. Demonstrates that an agent will deliver the intended outcome consistently As “k” rises, it means the agent will consistently deliver the preferred outcome.

100 is good, and 1 is not so good.

By leveraging the Pass@k and Pass^k methods at scale, enterprises can drive both autonomy and value without compromising on the risk of agents gone rogue.

The Road to Autonomy: A Quick Map

Before we dive into the math behind these metrics, let’s remind ourselves of both the destination and navigation tools that we shared in that earlier post: we use Pass@k and Pass^k to prove an agent is smart and stable, the Decision Delegation Index (DDI) to set the length of its leash, and the Policy Violation Rate (PVR) as the traction control.

null

(In the next blog post in this series, we’ll show how we’ve automated this entire reliability gate into a push-button Agent Evaluation Scorecard.)

Capability is Not Reliability

Most AI evaluations today measure capability: “Can the model solve this problem if given enough chances?” Pass@k measures theoretical potential. It asks if the correct answer is possible within the agent’s reasoning bounds. For example, if you ask an agent to solve a complex supply chain disruption and give it 10 attempts (k=10), and it gets it right at least once, it has “passed.”

For a research lab or a creative brainstorming tool, Pass@10 might be an impressive metric. It shows the model is “smart,” but in an enterprise context where good decisions drive financial results (and bad ones cost time and money), capability is not enough.

You need reliability because discovering a solution is not the same as delivering a service. A 10% pass confirms the agent has the logic, but lacks the discipline required for autonomy. This keeps you trapped in HITL (Human-in-the-loop) where the “cost to invoke” the agent includes the expensive time of a human auditor.

To reach Milestone 3 (Optimized Scale) in our Agentic Blueprint and move to HOTL (Human-on-the-loop) oversight, you need to transition from reasoning to hardened habit.

The Reliability Benchmark: Pass@1 & Pass^100

To move your agents from HITL (where the human validates every step) to Human-on-the-Loop (HOTL) (where the human audits by exception), you need to get closer to Pass@1 excellence while also ensuring that the agents perform consistently by scaling the number of trials to demonstrate reliability with Pass^100.

Metric Metric Definition What You’re Aiming For & What It Means
Pass@k The probability that at least one of k attempts is correct.

The “Can it?” metric.

Pass@1 with a probability of 95% or higher shows that your agent is very likely to get the right answer on the very first attempt.
Pass^k The probability the agent gets the right answer on every attempt.

The “Will it?” metric.

Pass^100 with a probability of 95% or higher shows the probability the agent is very likely to get the right answer on every attempt (with a sample size of 100 trials).

Pass^100 represents the ultimate stress test for Milestone 4: Autonomous Enterprise.

In Boomi’s Agentic Evaluation Framework, Pass@1 is the core reliability benchmark. However, an agent with a high Pass@1 can still be brittle. This is where we introduce Pass^k to measure decision stability.

While Pass@k looks at multiple attempts for one prompt, Pass^k measures the frequency of identical or equivalent outputs across k differently phrased requests. High Pass^k indicates “Decision Stability” — the hallmark of a hardened System 1 asset.

To move your agents toward true autonomy, we look for Pass^100 as the indicator of extreme decision stability. It represents the high watermark of a hardened habit; while your operational k may be lower for internal or low-risk tools, hitting excellence at a 100-trial scale is the gold standard for high-volume, mission-critical autonomous delegation. It indicates that the intended behavior has successfully transitioned from a variable response to a hardened habit.

When an agent demonstrates both a high Pass@1 score (typically >95% depending on the risk of the domain), and a high Pass^100 score (also >95% depending on the risk of the domain), you have earned the right to “lengthen the leash.” When these two KPIs are high enough, it is the proof that your “agentic engine” can handle a higher gear without causing the tires to slip.

The Autonomy Transmission: Quantifying the Leash with the DDI and PVR Formulas

OK, this section is going to add a little more math. And to set the context for this math, we’re going re-introduce a chart from our first blog post. The Digital Impact Maturity Model tracks the maturity of agentic AI from pilot projects to autonomous enterprise-grade operations delivering measurable business value. The model identifies four milestones along this path.

null

Boomi’s Agentic Blueprint recommends using two anchor metrics for managing the process of granting increasing levels of autonomy to agents. In the blueprint, these two anchor metrics define the “Safe Operating Envelope” of an agent: Decision Delegation Index (DDI) and Policy Violation Rate (PVR).

  • The Decision Delegation Index (DDI): Also known as “The Leash,” this acts as a transmission setting. Low gears (0.1–0.3) provide safety for Milestone 2 (Guided Action), while high gears (0.8–1.0) provide overdrive for Milestone 4 (Autonomous Enterprise).
  • The Policy Violation Rate (PVR): This is your traction control. It monitors if the agent attempts to deviate from established SOPs or compliance boundaries.

Strategic maturity requires the deliberate balancing of authority and evidence. Trust is not a static certificate granted at deployment; it is a decaying asset. If an agent hits a guardrail and the PVR spikes, it triggers a Leash Snap—a governance-driven intervention that downshifts the agent from an autonomous state back to HITL or Guided Action until the logic is re-hardened.

The Golden Rule of Boomi’s Agentic Blueprint: Scalable financial impact is only realized when your DDI increases while your PVR remains at background-noise levels, all underpinned by Pass@1/Pass^100 excellence.

In other words, scalable financial impact comes when you can trust agents to operate autonomously with minimal error rates, with performance tracked by metrics available to business and IT leaders alike.

1. The Decision Delegation Index (DDI) — “The Leash”

Think of the DDI not as a “warning light,” but as a transmission setting. Just as a vehicle uses different gears for different terrains, your enterprise uses different DDI levels for different levels of risk.

  • Low Gear (Low DDI / 0.1–0.3): This is “high torque” mode. It provides maximum human control and safety. You use this gear when the terrain is uncertain (i.e., a domain with high risk) or the agent is unproven (i.e., a suboptimal score in either pass@k or pass^k). This is the realm of Human-in-the-Loop.
  • High Gear (High DDI / 0.8–1.0): This is “overdrive.” It provides maximum efficiency and speed for well-paved, high-volume processes. This is the realm of Accountable Autonomy

The DDI quantifies how much authority is granted to an agent, weighted by the risk of the domain (Wr). It is calculated as:

DDI = Σ(Da * Wr) / Σ Wr

For the non-mathematicians in the room, the character capital sigma (Σ) means sum. In plain English, then: your DDI is a weighted average of how much decision-making authority you’ve handed to agents — where riskier domains count more. Think of it like a GPA where harder classes carry more credit hours. Automating a hundred trivial tasks barely moves your DDI; earning enough trust to let an agent run in a high-stakes domain moves it a lot. That’s deliberate — it stops an enterprise from looking “highly autonomous” just by automating the easy stuff.

Each agent’s Dₐ isn’t self-reported. It’s the gear the scorecard (detailed in the final blog post in this series), cleared that agent for. An agent only contributes a high number after passing its capability and security gates.

2. The Policy Violation Rate (PVR) — “The Guardrail”

If DDI is your gear selection, the Policy Violation Rate (PVR) is your traction control. It monitors if the tires are slipping — in this case, if the agent is attempting to deviate from established SOPs (standard operating procedures), compliance requirements, or data boundaries.

PVR = (Violations / Total Decisions) * 100

These two metrics can be combined with the Pass@k and Pass^k methods to allow an enterprise to systematically advance their agentic capabilities without having to compromise on predictability or safety.

For the non-mathematicians: Out of every 100 decisions an agent made, how many broke a rule. If your agents made 5,000 decisions and 3 violated an SOP, your PVR is 0.06 — background noise. If it’s 2 or 3, the tires are slipping and the Leash Snap kicks in.

If the PVR spikes (you lose traction), indicating you need to “downshift” your DDI back into a lower, human-governed gear until the agent is re-hardened.

The Agent Trust Matrix: Navigating the 2×2

To visualize the transition to autonomy, Boomi’s Agentic Blueprint utilizes a 2×2 maturity matrix. This grid maps Reliability (Pass@k/Pass^k) against Delegation (DDI), revealing four distinct zones:

null
  • Pilot Purgatory (Low Reliability / Low Delegation): This is where most AI initiatives stall. The agent is unproven, and humans are doing all the work. There is high activity but zero impact on the P&L.
  • Running with Scissors (Low Reliability / High Delegation): A dangerous zone. Here, agents are given high authority (high DDI) without the empirical evidence of reliability. This leads to policy violations and “leash snaps” that can derail an entire AI program.
  • The Opportunity Gap (High Reliability / Low Delegation): A state of missed value. The agent has reached Pass@1 excellence, but the enterprise has not yet “lengthened the leash.” The human remains a bottleneck for a process that is ready for autonomy.
  • The Sweet Spot (High Reliability / High Delegation): The target state for Milestone 4. Authority is granted based on evidence, and the agent operates with accountable autonomy, closing the Impact Gap.

From Math to Reality

Understanding the math is the first step. The second step is instrumentation.

In our final post, we’ll move from the scientist’s lab to the industrialist’s factory floor. By leveraging the power of the Boomi Enterprise Platform with the methodology of the Agentic Evaluation Framework, we’ll automate and deliver the Boomi Agent Evaluation Scorecard to accelerate your journey to the “Autonomous + Accountable” state.

Ready to move from proving your agents are reliable to actually putting them to work? Explore Boomi Platform Agents and see how to operationalize trusted, autonomous AI across your enterprise. Explore Boomi Platform Agents.