Fund your first agent-purchasing pilot with a number you'd sign off on losing completely, and refuse to top it up for thirty days. That one decision sets the worst case before the first purchase rather than after the statement lands, and it lets the pilot fail without anyone explaining a loss.

Here's the four-week version, sized for one agent, one job and one approver. It ends with a document your finance lead can read on their own.

Why bound the pilot before it starts?

A bound you set in advance is a number you already accepted, and an accepted number isn't a loss anyone has to justify. Monitoring after the fact gives you the same information a week late, once the money has moved. Set the ceiling, the categories in scope and the approval threshold on day zero, and the pilot's downside stops being an open question.

The difference shows up in who you talk to when something odd happens. A purchase blocked or escalated for approval while the merchant waits is a conversation with your approver. The same purchase found in next month's reconciliation is a conversation with finance.

Decided before the pilot startsDiscovered during it
The ceiling on total spendWhat the agent buys that you didn't predict
Merchant categories in scopeHow often a person overrules it
The week 2 approval thresholdHow long an approval sits unanswered
Who approves, and their backupWhether the job was scoped narrowly enough

The left column is arithmetic you do once. The right column is why you're running this.

Week 1: what does "working" mean, in numbers?

Pick one agent and one purchasing job narrow enough to describe in a sentence, then write down the numbers that would make you continue. Write them before the agent spends anything, because a target chosen in week 4 will be the target your results already cleared. A research agent buying dataset access and API credits from vendors you've used before is a good first job.

You won't find an industry figure to set that target against, and one borrowed from someone else's write-up won't survive your finance review. Set your own baseline from what the job costs today: the spend when a person does it by hand, and the hours they book to it.

Week 2: what does a low approval threshold buy you?

It shows you where the agent wants to spend before that costs you anything. Set the threshold low enough that a person sees most purchases in week 2, and treat every approval as a data point rather than an interruption. What arrives in that queue is the finding, and no policy document written in advance would have told you.

Two things go in your notes each time: whether the approver said yes, and how long they took. A queue where every request clears in under a minute means the threshold sits too low. Three requests sitting overnight means a staffing problem the rollout will inherit.

Expect purchases you didn't predict. Agents find merchants you never listed, which is why an agent buying something that doesn't exist belongs in the budget conversation.

Week 3: where do you loosen, and where do you leave it tight?

Loosen only where week 2 gave you evidence. A category where the approver said yes every single time is a category where the threshold sat too low, so raise it there and watch what changes. A category that produced a disagreement stays where it is until you know why. Loosening everything at once turns week 3 into a second week 1.

Change one thing at a time. An approver who says no because the amount was wrong is telling you about your threshold. One who says no because the purchase sat outside the job is telling you the job was scoped wrong, and that's the more expensive finding.

Test the exit this week too. Revoke the agent's card and confirm the next attempt stops, because a spending kill switch you've never used is a claim rather than a control.

Week 4: what artifact decides the rollout?

A per-agent record covering what the agent attempted and what a person stopped, with the elapsed time on each approval. One page per agent, readable by your finance lead without you translating it. If the decision to expand rests on a dashboard screenshot or your account of how the month went, the pilot produced a story instead of evidence.

Three numbers carry that page.

How often the agent tried something outside the range you expected. Count attempts rather than losses, because a blocked or escalated attempt never became a loss and still tells you what the agent wanted.

How often a person disagreed with it. That's the no rate in your approval queue, and your closest measure of whether your policy and your agent read the job the same way.

Approval time, median and worst case. The worst case decides whether this survives a hundred agents. Put the total purchase count beside all three, or none of the rates are readable.

Shatale gives each agent its own scoped virtual card with your policy enforced at the authorization moment, purchases outside policy blocked or escalated for approval, and an immutable per-agent record of every decision. In a pilot, the ceiling and the threshold sit on a card belonging to one agent, and the week-4 document comes out of that record.

What to ask

FAQ

How long should an AI agent purchasing pilot run?

Thirty days covers a full billing cycle and is short enough that the team still remembers why it started. The binding constraint is purchase volume: if the job produces four purchases a month, extend the window or pick a busier one.

What budget should a first agent purchasing pilot have?

Whatever you'd sign off on losing entirely, with no top-up path while the pilot runs. There's no industry figure to copy, and a number from someone else's write-up won't survive your finance review. Start from what the same job costs your team today, and cap the pilot below that.

What should you measure during an agent purchasing pilot?

Count attempts outside the range you expected and disagreements in the approval queue, then time each approval from request to decision. Put the total purchase count beside all three so the rates are readable. Set your own baseline for each before week 1.

Can you run an agent purchasing pilot on an existing corporate card?

You can, and you'll lose most of what the pilot was for. A shared card produces one statement covering the agent and everyone else holding that number, so an attempt or an approval can't be attributed to the agent afterwards. The output of the pilot is a per-agent record.


If the pilot clears, moving the rest of your agents off shared cards is the next decision, and a larger one. Early access is free for publishers.

Shatale is the control layer for AI-agent payments. Its authorization architecture is the subject of European patent application EP26194994.5 (filed; priority 28 July 2026).