A Worksheet for Fully Loaded Agent Costs
Updated October 8, 2026 · first published October 8, 2026
A model token total is not a fully loaded agent cost. A production run can also consume hosted search, code execution, storage, retries, fallback models, and human review. A useful worksheet connects each charge to a run ID and keeps measured provider charges separate from allocated overhead and labor estimates.
This worksheet complements cost per successful task: first assemble the complete cost ledger, then divide it by outcomes that passed the stated success test. It also extends separate tool billing into a repeatable finance record.
Start with one row per event
Use a stable workflow ID for the user intent and a run ID for each attempt. Record timestamp, environment, team or tenant, workflow, run, provider, service, resource ID, quantity, unit, rate-card version, currency, and invoice or metering source. Keep one event per model call, tool call, compute job, or review item. This makes retries visible and lets finance reconcile totals without guessing which nested operation created them.
Worksheet columns
- Model: provider, model, input/output/cached/reasoning units where exposed, measured rate, and direct charge.
- Tools: search, retrieval, browser, code execution, and other metered services, with call count or runtime and direct charge.
- Compute and storage: sandbox or container runtime, accelerator or CPU time, persistent storage, and network charges if billed.
- Reliability work: retries, fallback calls, abandoned branches, and repair runs. Keep them in total run cost even when they produced no useful output.
- Verification: evaluator calls and human review minutes, with labor cost marked as an internal estimate rather than a provider invoice charge.
- Shared platform: orchestration, observability, reserved capacity, or shared storage allocations, each with its allocation driver and period.
Calculate direct and fully loaded cost separately
Direct run cost is the sum of metered model, tool, compute, and storage charges attributed to that run. Fully loaded cost adds the declared share of platform overhead and review labor. Use a stable driver for shared costs, such as measured runtime, reserved capacity, or a documented per-run allocation. Publish the driver and its time window; do not distribute fixed costs by token count unless tokens actually explain the resource consumption.
A small hypothetical example can make the formula clear: if one run has $0.40 in model charges, $0.10 in tools, $0.20 in sandbox compute, and $0.30 of allocated review and platform cost, report $0.70 direct and $1.00 fully loaded. These figures illustrate the arithmetic only; they are not a market price or benchmark.
Keep attempts and outcomes distinct
Count every attempt in spend, then label the final workflow as successful, failed, abandoned, or requiring human completion under an explicit rule. Report cost per completed workflow alongside success rate, retry rate, and review effort. For a successful-task metric, use total eligible spend in the period divided by the number of workflows that passed the success test. Do not drop failed runs from the numerator or count intermediate agent steps as separate successes.
Reconcile and govern the worksheet
Reconcile event totals to provider invoices and cloud billing exports after usage data settles. Explain credits, discounts, minimums, late events, and currency conversion separately. Store the price source and effective date so historical costs can be recalculated without applying today's rate to last month's usage.
Assign an owner to each workflow and define a per-run budget, an aggregate period budget, and an escalation path. When a workflow crosses its budget, record the reason and whether it was allowed to continue. For runtime enforcement patterns, see the three budgets for autonomous agents.
FAQ
What belongs in an agent cost worksheet?
Record model usage, billed tools, compute, storage, retries, fallback calls, verification, and separately identified shared platform allocations and labor estimates.
Should failed agent runs be included in cost per task?
Include their spend in the total. Report the count of successful workflows using a declared success rule so failures remain visible in the unit economics.
How do I avoid double-counting shared platform costs?
Keep direct charges attached to their event, allocate shared costs in a separate ledger using one published driver, and reconcile the allocated total to the platform cost pool.
Related
Related
Want this applied to your stack? Bring the provider bills, gateway logs, and top workflows; we will map the cost drivers and savings path. Book a free audit →