One AI-agent task cost $0.43 at the median. Another cost $6.05. Same production system. That is a roughly 14.1× gap.
Important: that 14.1× gap comes from one operator’s self-reported alert-investigation workload. It is an example of how much runs can vary inside one system — not a benchmark for what AI agents generally cost.
The lesson is not that every agent task is expensive. It is that an AI agent cost per task can be very different from the price of one prompt or one million tokens. One “investigate this alert” request may become a loop of model calls, tool calls, log reading, retries, and verification.
PEEK POINT
$0.43 median → $6.05 max
One task. 14.1× apart.
Quick answer
- One operator publicly reported 7,364 AI-agent alert investigations and $4,458 of token spending over three months.
- Its median run cost $0.43; the most expensive reported run cost $6.05.
- NUTPEEK calculation: $6.05 ÷ $0.43 = 14.1×.
- This is self-reported production data from one workload—not an audited benchmark or the average cost of AI agents.
AI agent cost per task: $0.43 vs. $6.05
The operator described an agent that reads logs and traces to investigate monitoring alerts, then posts a root-cause write-up. Here is the reported three-month snapshot:
| Reported metric | Value |
|---|---|
| Alert investigations | 7,364 |
| Total token spend | $4,458 |
| Median run | $0.43 |
| p90 run | $1.29 |
| Most expensive run | $6.05 |
| Average tool calls | 16.1 |
| Expensive runs | 50–128 tool calls |
Data label: SELF-REPORTED PRODUCTION DATA. The figures come from a Reddit post by one operator. They are useful because they show the shape of a real workload, but they should not be used as a universal AI-agent price list.
What the math says
The same report allows two simple calculations:
- Maximum versus median: $6.05 ÷ $0.43 = 14.1×.
- Total spend per investigation: $4,458 ÷ 7,364 ≈ $0.61.
Why is the $0.61 average higher than the $0.43 median? The high-cost tail pulls the average up. In other words, many ordinary runs may be inexpensive while a smaller number of long, tool-heavy runs change the economics.
Why one task can cost 14× more
To a user, an agent task can look like one request: “Investigate this alert.” Inside the system, it can look more like this:
Ask → plan → fetch logs → read result → call a tool → read more → retry → verify → write answer.
Each step can add model input, tool output, and another decision. The reported expensive runs used 50 to 128 tool calls, compared with an average of 16.1. More tools are not automatically wasteful; some problems are genuinely harder. But each extra round creates another chance to add context and cost.
The hidden context problem
The report’s five most expensive runs had an input-to-output ratio of about 99:1: roughly 33.6 million input tokens versus 338,000 output tokens. That is a clue that the expensive part was not simply a long final answer.
It can be the material the agent has to read before it answers: log lines, traces, tool responses, prior attempts, and context carried into the next call. If a tool returns a large raw payload and that payload stays in the working context, later model calls may need to process it again.
That does not prove every long run is inefficient, and this article does not assume a fixed formula for how cost grows. It does show why “one task” is often too simple a unit for an agent system.
Why cost per task beats a token-price headline
API pricing pages usually quote input and output token rates. Those rates matter, but they cannot answer the business question by themselves: what does one successful, completed task cost?
That number includes model choice, tool loops, retries, context size, caching behavior, and the route a task takes through the system. Two agents using the same model can have very different cost per task because their tools and stopping rules are different.
PEEK INSIGHT
The expensive part is not always the answer. It may be the loop before the answer.
What to measure
Track AI agent cost per task alongside the metrics below.
- cost per completed task, not just total token spend;
- median, p90, and maximum—not only the average;
- tool calls per run and which tools create the largest outputs;
- retries, repeated queries, and context that remains after it is no longer needed; and
- quality alongside cost, because the cheapest run is not useful if it fails the task.
Calculate Your Own Cost per Completed Task
Start with the simplest number you can reproduce: total agent spend ÷ successfully completed tasks.
Example: if an agent costs $120 over a period and successfully completes 300 tasks, the average cost per completed task is $120 ÷ 300 = $0.40. Track median, p90 and maximum alongside the average when your system exposes them, because a small number of expensive runs can pull the average upward.
Bottom line
For AI agent cost per task, one prompt is not always one call. For agents, the real unit is the whole task loop. The reported $0.43-to-$6.05 gap is a reminder to measure the cost of completed work—and to inspect the tail, where an apparently small number of expensive runs can reshape the average.
Sources and related reading
- Self-reported production AI-agent cost data (Reddit)
- Is Paid AI Worth It? Count the Time You Spend Fixing It
Educational information only. Agent costs vary by model, workload, tools, prompts, context handling, and provider pricing.
