Sierra's researchers ran function-calling agents through customer-service tasks with real rules. The best of them, gpt-4o, succeeded on fewer than half. Asked to do the same task eight times, it got all eight right on fewer than a quarter. An agent that works in the demo is a turkey on day thirty.
[1]Thesis · v1.0
the model is not the control.
An agent that succeeded a thousand times has told you nothing about the thousand-and-first. So the things that must never go wrong are not decided by the model. They are decided before it runs.
Abstract
Abstract
Companies are about to let software spend their money and speak in their name. The usual answer to "is it safe" is a better model. This paper argues that the question is aimed at the wrong component.
A model's behavior is a distribution, and you can't put a distribution in charge of a bank account. What an agent may do, what it may spend, which actions wait for a person and what gets written down have to be settled outside the model, in code that runs before the side effect.
Below are six theses, the published evidence behind each, how OneShot enforces it, and the places where it stops working. We are unsure about most of the future of agents. We are sure about the part that is enforced.
- action
- A call with a side effect outside the system: an email sent, a call placed, a purchase made. Reading is not an action.
- policy
- A rule evaluated in code before an action runs. It returns allow, deny or hold. The model never sees it and cannot argue with it.
- spend
- Money committed by an action, quoted in USDC before the call and capped per call and per day.
- approval
- A human decision bound to one specific held call, used once.
- receipt
- A signed record of one action that a third party can verify without asking us.
01 · The evidence
a thousand good days prove nothing
A turkey is fed every day for a thousand days. Every feeding raises its confidence that the farmer means well. Day 1,001 is the Wednesday before Thanksgiving.
Taleb borrowed the bird from Russell, who used a chicken. The point survives the change of species: a record of success says nothing about the case that matters.
Text an agent reads can carry instructions, and the agent can follow them. Willison's trifecta is access to private data, exposure to untrusted content and the ability to communicate externally. An outbound agent has all three by job description: it reads the prospect's site, and then it sends the email.
[2][3][4]Gartner's forecast names three causes: escalating costs, unclear business value and inadequate risk controls. Forecasters are paid to sound sure, so read the number as a mood. The third cause is the only one an engineer can do anything about.
[5]02 · Six theses
what gets decided before the model runs
T1 permissions are not a prompt
What an agent may do is decided by a policy outside the model, on every call.
Tell a model "never email more than ten people" and you have made a request. A policy that returns 403 is a fact. Most of agent security is the work of turning the first into the second.
OWASP files this under Excessive Agency and says it without decoration: implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed. NIST opened a project on agent identity and authorization in February 2026 for the same reason.[6][7]
How OneShot does it
On the four routes that reach the outside world on someone's behalf (email send, SMS send, voice call, commerce buy), the action policy runs after the quote is validated and before any money moves. It returns allow, deny or hold.
A denied or held call is never charged. If the policy of an agent known to have one can't be read, the call is held. It fails closed.
VerifyA denied call returns 403 action_denied_by_policy and leaves no charge on the ledger.
T2 skin in the game, in USDC
Every paid action is priced before it runs and capped before it signs.
Hayek's point about prices was that one number carries more information than any planner can. It turns out the math doesn't check your species. An agent that sees the price picks the $0.02 email over the $0.35 call without being told.
Seeing the price is one thing; being stopped by it is another. A budget reconciled at month end tells you what already happened. The cap has to bite before the payment is signed.[8]
How OneShot does it
Every paid call is quoted first: the x402 response carries the price, and the agent signs only that amount.
Each agent can carry a per-transaction cap and a daily cap in USDC. Both are checked against the receipts ledger, plus calls still in flight, before payment verification. An over-budget call never signs.
VerifyGET /v1/agents/me/budgets returns the caps and today's utilization, ledger and in-flight separately.
T3 approvals are for the irreversible
A human approves the actions that can't be undone, and only those.
Taleb calls harm done by the healer iatrogenics. A person asked to approve a $0.0004 email isn't doing oversight. They are being trained to click yes, and by the hundredth click the approval means nothing.
The EU AI Act asks for something narrower and harder: a person able to override the system, or halt it with a stop button. The Act wants someone with the power to say no. Asking for yes a thousand times a day wears that power down.[9][6]
How OneShot does it
A policy can hold a call for approval. The approver is paged once per distinct call. The approval is bound to the agent and a hash of that exact request, expires, and is consumed atomically, so it buys one execution and no more.
Until it is approved, the held call costs nothing and does nothing.
VerifyWithout an approval the call returns 403 approval_required with the pending approval. Replaying a used approval is refused.
T4 a receipt you don't have to trust us for
Every action leaves a signed record that anyone can check offline.
Pacioli wrote down double entry in 1494. It has outlived every merchant who used it. A ledger is the most Lindy thing in commerce.
Article 12 of the AI Act asks high-risk systems to record events automatically. But a log we keep proves only that we kept a log. A signature proves it to someone who has every reason to doubt us.[10][7]
How OneShot does it
Every terminal receipt is signed with Ed25519. In production an unsigned receipt counts as an incident.
The signer and the verifier share one canonical payload, so they can't drift apart. Verify with the SDK's verifyReceipt() or the oneshot-verify-receipt CLI against our public key. No call to us needed.
Verifynpx oneshot-verify-receipt <receipt.json>
T5 payment is a protocol, not a wallet we hold
Agents pay over open standards, so the rail outlives the vendor.
HTTP has carried a status code for machine payments since 1999. 402, Payment Required, sat reserved for future use for over twenty-five years because nobody had a buyer who wasn't a person clicking a button.
Now there are two serious answers. x402 is an open standard for payment over HTTP with a foundation behind it. Stripe's Agentic Commerce Protocol, built with OpenAI, keeps the merchant of record in control of what is sold and how. Pick the rail most likely to outlive whoever sold it to you.[8][11]
How OneShot does it
OneShot takes both. USDC on Base over x402, verified and settled on-chain. Cards over Stripe ACP checkout sessions, paid with a Stripe shared payment token.
Neither rail requires you to park money in an account we control.
VerifyEvery paid route answers an unpaid request with 402 and a price you can read before you pay.
T6 one workflow, measured
Start with one task and a definition of success agreed before anything is built.
Anthropic's advice to people building agents is to find the simplest solution possible and increase complexity only when needed. Hardly anyone follows it, because complexity demos better.
MIT's NANDA group reported that 95% of organizations saw no business return from generative AI. Their method, 52 interviews, 153 surveys and a review of 300 public initiatives, has been fairly criticized. We cite it as a temperature, not a measurement. Our own number for your company doesn't exist yet, and anyone who quotes you one before running the work is selling you the feeling of certainty.[12][13]
How OneShot does it
Enterprise pilots start from one workflow and the measures that will judge it, written down before we build.
VerifySee the pilot workflows and their measures on /enterprise.
03 · Control map
for the reviewer with a spreadsheet
Our mapping, offered as a starting point for your own. It is not a certification or a compliance claim.
| Thesis | NIST AI RMF | EU AI Act | OWASP LLM Top 10 (2025) |
|---|---|---|---|
| T1 permissions are not a prompt | Govern, Manage | Art. 14 | LLM06 Excessive Agency |
| T2 skin in the game, in USDC | Manage | n/a | LLM06 (excessive autonomy) |
| T3 approvals are for the irreversible | Govern, Manage | Art. 14(4)(d)(e) | LLM06 (human approval) |
| T4 a receipt you don't have to trust us for | Measure, Manage | Art. 12 | n/a |
| T5 payment is a protocol, not a wallet we hold | Map | n/a | n/a |
| T6 one workflow, measured | Map, Measure | n/a | n/a |
04 · Where the controls stop
via negativa, applied to ourselves
- Spend caps fail open
If the budget store is down, paid calls go through. We chose availability over the cap. The action policy makes the opposite choice and holds the call.
- Four routes are behind the action policy
Email send, SMS send, voice call and commerce buy. Every other paid tool is bounded by its quote and the spend caps, not by policy.
- Prompt injection is not solved
Policy limits what a hijacked agent can do. It does not stop the agent from being hijacked.
- A receipt proves what happened, not that it was wise
Signatures settle disputes about facts. They don't settle disputes about judgment.
- No ROI number
We don't claim a lift in reply rate, revenue or pilot success. When we have one we'll say so here, with the method.
How sure to be of the evidence
- Gartner's figure is a forecast published in June 2025, not a count.
- The MIT NANDA figures are cited from press coverage. The report's method has been publicly disputed.
- τ-bench and AgentDojo measured models available in 2024. Newer models score higher. The gap between pass@1 and pass^k is the finding we rely on, not any one score.
- AI Act articles are quoted from an unofficial consolidated text. Obligations for high-risk systems depend on classification and have been subject to postponement. This page is not legal advice.
We can't tell you what agents will be doing in five years. Nobody can, and we'd avoid anyone who says otherwise. We can tell you what ours are not allowed to do today, and how to check.
References
References
- [1]Yao et al., τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (2024) Primary source · checked Oct 4, 2026
- [2]Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023) Primary source · checked Oct 4, 2026
- [3]Debenedetti et al., AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (2024) Primary source · checked Oct 4, 2026
- [4]Willison, The lethal trifecta for AI agents (16 Jun 2025) Primary source · checked Oct 4, 2026
- [5]Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (25 Jun 2025) Primary source · checked Oct 4, 2026
- [6]OWASP GenAI Security Project, LLM06:2025 Excessive Agency Primary source · checked Oct 4, 2026
- [7]NIST NCCoE, Accelerating the Adoption of Software and AI Agent Identity and Authorization, concept paper (5 Feb 2026) Primary source · checked Oct 4, 2026
- [8]x402: an open, neutral standard for internet-native payments Primary source · checked Oct 4, 2026
- [9]Regulation (EU) 2024/1689 (AI Act), Article 14: Human oversight Secondary source · checked Oct 4, 2026
- [10]Regulation (EU) 2024/1689 (AI Act), Article 12: Record-keeping Secondary source · checked Oct 4, 2026
- [11]Stripe, Developing an open standard for agentic commerce (29 Sep 2025) Primary source · checked Oct 4, 2026
- [12]Anthropic, Building effective agents (19 Dec 2024) Primary source · checked Oct 4, 2026
- [13]MIT NANDA, The GenAI Divide: State of AI in Business 2025, as reported by Virtualization Review (19 Aug 2025) Secondary source · checked Oct 4, 2026

