AI Agent Pilot Costs: Measure the Work You Can Accept

Measure AI agent pilot costs per accepted outcome, including retries, human review and unfinished work. A practical worksheet for startup founders.

Ernest Bursa

Ernest Bursa

Founder · · 13 min read
A male founder and an older female colleague count cartons beside a notebook outside a small San Francisco shop.

To measure AI agent pilot cost, divide the resources spent on a defined batch of work by the outcomes independently accepted before its deadline. Include software, retries, human review and recovery. Show completion, quality and cash spending beside the result so a cheap unfinished task cannot pass for useful output.

Use that number to decide whether to extend a pilot. A model bill alone cannot tell you whether the workflow is affordable. You also need to know what reached the required standard, what remains in the queue and who helped it get there.

What does Pion tell you about agent economics?

Pion makes real-world agent experiments more accessible, but its launch does not establish the economics of your business. Treat it as a reason to measure a bounded workflow carefully.

On September 14, Andon Labs announced Pion as a research preview for running autonomous-business experiments. Its launch drew an active Hacker News discussion. Andon reported progress in its vending experiments while saying its cafe and retail shop were still unprofitable. Those are developer-reported results from particular settings, not an independent audit or a forecast for other companies. Andon Labs’ launch report.

The distinction matters because several economic questions can sit behind the word works. Does the system complete a task? Does a sale cover the goods sold? Does the operation cover its ongoing expenses? Does the investment pay back its setup cost? You need to choose the question before collecting numbers.

Andon’s earlier cafe report illustrates another distinction: ingredient and packaging costs, rent and wages, inventory value, and cash paid to suppliers describe different parts of the business. Unsold stock can retain value while cash becomes scarce. Neither view alone explains the whole operation. Andon Cafe report.

For your pilot, write the decision at the top of the worksheet: “Should we run this workflow again under these conditions?” Research can be worthwhile even when an experiment loses money. Commercial expansion needs its own evidence. Keep the learning budget visible, then assess ongoing delivery costs separately.

What counts as an accepted outcome?

An accepted outcome is a result that meets a written standard and has been checked independently of the agent’s own completion claim. Define that standard before the pilot starts.

Anthropic’s guide to agent evaluations distinguishes a record of the agent’s actions from the resulting state of the system. It also treats repeated attempts as separate trials. That is useful measurement discipline: the agent saying a task is done does not prove the required result exists. Anthropic’s evaluation guide.

Here is an original operational example. Suppose your pilot prepares supplier records from documents you already hold. An accepted record must identify the correct supplier, populate the required fields, attach the supporting document and flag conflicting information. A reviewer checks those conditions before marking it ready for the next stage. This preparation pilot does not authorize payments or change bank details.

The unit is one usable supplier record. Document pages, tool calls and drafts are activity. You can track them to diagnose cost, but they do not increase the number of accepted outcomes.

Freeze the batch and the deadline

Give every assigned case an identifier and record when it entered the pilot. Keep the original batch intact, including the awkward cases that fail. If you exclude cases before starting, document the eligibility rule so the eventual result describes the work you actually tested.

Choose a cutoff that reflects when the work is needed. A record accepted after that cutoff can become useful later, but it was not ready for the original service deadline. Preserve the original snapshot and add the later result. Otherwise, each extra day quietly improves the headline without showing the cost of waiting.

Use three outcome states at the cutoff: accepted, rejected and unfinished. A rejected result failed the standard; an unfinished case has no final accepted result yet. Keep a reason beside each so you can distinguish poor output from a missing input or a reviewer backlog.

What belongs in the pilot cost ledger?

Record resources against the whole assigned batch, including work that never gets accepted. Separate money paid from the value of staff time allocated to the pilot.

You can start with a worksheet and links to existing records. Reconcile what happened using the records you already have.

Record What to capture What it helps you check
Assigned work Case identifier, arrival time, task type Whether the batch changed
Acceptance Decision, timestamp, reviewer, evidence Whether useful output exists
Software Provider usage, tool fees, allocated subscription cost Resources used across all attempts
Human effort Setup, review and recovery time Work outside the agent’s own activity
Unfinished work Current state, reason, next owner Obligations still waiting
Cash Invoice, payment, credit or subsidy What actually left the business
Configuration Model, instructions and relevant tool changes Which setup produced the result

Keep ordinary review separate from recovery. Reading a prepared record and checking its evidence is review. Finding the correct document after a wrong match, repairing fields and restarting the task is recovery. Both count; the distinction helps you decide what to improve.

Human help does not invalidate the experiment. Project Vend’s second phase reported improved performance alongside continued human purchase checks and other support. Its public account does not provide a full cost reconciliation, so you cannot infer comprehensive unit economics from it or assume a particular cost was excluded. Project Vend, phase two.

Capture interruptions without turning work into surveillance

Ask participants to record time by activity and batch, with enough detail to explain repeated problems. A brief note such as “resolved supplier identity conflict” is more useful than a pile of screenshots. Avoid collecting private communications or unrelated staff activity just to make the worksheet look complete.

Make uncertainty visible. If recovery took approximately an hour, mark it as an estimate. If work was shared across several pilots, write down the allocation rule. Precise-looking totals built from unexplained guesses make the next decision harder.

How do you calculate cost per accepted outcome?

Add the operating resources used for the batch, then divide by the number of outcomes accepted at the cutoff. Report setup separately so readers can see both recurring operation and the first batch’s full resource requirement.

The following example is entirely hypothetical. The dollar amounts, time, workload and acceptance counts are invented to demonstrate the calculation. They are not Pion results, Kit customer data or an industry benchmark. All amounts are in US dollars for illustration.

Your supplier-record pilot receives 100 cases. At the cutoff, 80 are accepted, 12 are unfinished and 8 are rejected. AI and software operating costs total $200. Staff spend 6 hours reviewing and 2 hours recovering failed work, using an assumed loaded labor rate of $50 per hour.

Here, the loaded rate is an allocation for the cost of staff time, including the employment costs you choose to include. State your own definition when replacing these assumptions with actual records.

Hypothetical operating input Calculation Allocated cost
AI and software Batch total $200
Human review 6 hours × $50 $300
Recovery 2 hours × $50 $100
Total operating resources $200 + $300 + $100 $600

Operating cost per accepted outcome: $600 ÷ 80 = $7.50.

If no outcomes are accepted, the ratio is undefined. Report the resources spent and zero accepted results; do not display a zero cost or hide the failed batch.

Dividing the software bill by assigned cases gives $200 ÷ 100 = $2. That is software cost per assigned case. It leaves out staff effort and uses a denominator that includes unfinished and rejected work. Label it correctly if you report it, but do not use it to describe the cost of accepted output.

Show setup and cash in separate views

Suppose initial setup takes 10 hours at $50, or $500. Including it makes the first batch’s resource total $1,100, giving $1,100 ÷ 80 = $13.75 per accepted outcome. That tells you what the first batch consumed, while $7.50 describes its operating resources before setup.

For a later planning estimate, you could allocate setup across an expected volume. Show the volume assumption and what happens if the pilot stops early. Do not charge the entire setup amount and an amortized share of that same amount in one total.

These labor allocations are not necessarily new cash payments. Staff on existing salaries still consume time, but the pilot may not change payroll. Keep a cash column for actual charges and a resource column for the broader comparison. Neither should quietly substitute for the other.

Pion’s product page describes seed tokens for selected ideas and an expected revenue-share model, without a universal numeric fee. Record any support your own pilot actually receives, including its expiry or limits. Do not assume a temporary credit represents the price of ongoing operation. Pion product page.

How do you compare the pilot with today’s workflow?

Compare work of similar difficulty under the same acceptance standard and deadline. Put incomplete work beside the cost comparison so better unit cost does not conceal worse delivery.

In the hypothetical comparator, the current workflow accepts 100 records in 20 hours at $50 per hour. Its allocated resources are $1,000, or $10 per accepted outcome. The pilot’s operating figure is lower at $7.50, but it accepted only 80 records and its first-batch figure was $13.75 with setup.

You therefore have several findings, not a verdict that the workflows are equivalent. The pilot used fewer allocated operating resources and delivered fewer accepted results by the cutoff. You still need a plan for the unfinished and rejected cases. A hiring or savings claim would require more evidence than this worksheet provides.

Make the comparison symmetrical

Include review and corrections in the current workflow just as you do for the pilot. If a person checks another person’s work today, keep that effort in the baseline. If human mistakes are discovered later, record them rather than holding the agent to a standard the comparator never faced.

Choose comparable case groups before looking at outcomes. A pilot that processes clean documents while staff handle missing attachments has tested a narrower workload. That can be a sensible first experiment, provided the conclusion stays within that boundary.

Also record which people supplied the hours. Less junior processing time alongside more scarce specialist review may be a poor trade for your team. The average cost can hide a bottleneck that prevents the next batch from finishing.

Test whether the apparent saving survives ordinary changes

Keep a separate scenario column for assumptions that might change: support ends, review takes longer, the case mix becomes harder or setup must be repeated. Change one assumption at a time so you can see what drives the decision. Label every projected figure; do not mix scenarios into the observed pilot result.

Hours released are capacity you might use elsewhere. They become a cash saving only when an actual expense changes. For broader staffing choices, use that evidence alongside headcount planning for people and AI agents, rather than converting a time estimate directly into a vacancy decision.

When should you extend, change or stop the pilot?

Decide in advance what evidence will justify another batch. Cost, acceptance, timeliness and unacceptable failures need separate conditions because one average cannot answer all of them.

For the supplier-record example, a wrong bank detail escaping into a payment process would require a stop and investigation, even if the average cost looked attractive. Serious errors should have explicit handling rules. Pricing them into an average does not make the underlying exposure acceptable.

Write a short decision record at the cutoff:

  1. What was tested? State the batch, configuration, eligibility rules and deadline.
  2. What was delivered? List accepted, rejected and unfinished cases, with evidence and reasons.
  3. What did it consume? Show software, human effort, setup and actual cash spending.
  4. What remains? Name who will finish or repair outstanding work and where its later costs will be recorded.
  5. What happens next? Extend the same setup, change a specific part, or stop, with a reason and a review date.

If you change the model, instructions or tools, record the new configuration before the next batch. Better results after several changes tell you about the revised system. They do not isolate which change caused the improvement.

Missing evidence is also a result. If nobody can reconstruct review time or verify acceptance, improve the measurement before making a large expansion decision. You can still report what was learned, along with the spending used to learn it.

As a pilot grows, reviewer availability becomes an operational dependency. If that is your constraint, the separate guide to building a human review team covers staffing continuity. In the cost worksheet, retain the simpler question: how much qualified human effort did this batch require?

Where can Kit contribute to the cost evidence?

Usage records can supply part of your pilot ledger, but accepted outcomes and staff effort need their own evidence. Start by identifying exactly which calls a product meters.

Kit records estimated cost and usage metadata for covered in-app AI calls, with member and monthly views. Member rows exclude unattributed background calls that monthly aggregates include. Its limits check usage before new metered calls; they are not a universal exact ceiling on provider invoices. These records cover part of operating cost. Kit’s AI providers and spend controls guide.

Keep the coverage limits together when using those records: Hiring Live Review and Compensation Research background enrichment sit outside the ledger and caps, and embedding metering is incomplete. External MCP clients’ model bills are separate, though an MCP action can trigger metered generation inside Kit. Kit does not provide a verified Pion integration, staff-minute measurement or an all-business profit model. Its usage estimates alone cannot establish economic savings or reconcile every provider bill.

Bring the covered usage records together with invoices, acceptance evidence and your own time allocations. Keep the original batch and cutoff visible. You can then explain the decision: what became usable, what remained unfinished and what resources the work consumed.

Put the next pilot on a measurable footing. If your workflow runs through Kit, use its covered AI usage records as one input to your worksheet. Start a free trial.

Related articles

Ready to hire smarter?

Start free for 30 days. Cancel before it ends and you pay nothing. Set up your first hiring pipeline in minutes.

Start hiring free