Start from the question a receipt has to answer
A useful way to design or evaluate an AI agent receipt is to imagine asking, a week after a task ran, “what happened here, and why?” If the receipt can’t answer that without you having to guess, reconstruct, or take it on faith that it went fine, it’s not doing its job. This sounds obvious, but most activity logs fail this test, because they’re built to record system events, not to answer a person’s question later.
The rest of this post walks through what belongs in a receipt by tracing one example task from start to finish, including a step that didn’t go cleanly, since that’s exactly the kind of moment a weak receipt tends to gloss over.
The worked example: an overnight vendor research task
Say the request, sent Sunday night, was: “Compare pricing and support quality for five project management tools we’re considering switching to, and have a summary ready by Monday morning.” Here’s what a receipt for that task should show, section by section.
The original request, in the words it was given, not a paraphrase. This is the baseline everything else gets judged against: did the result actually answer what was asked?
What was accessed and where information came from. For this task: each vendor’s public pricing page, a handful of third-party review sites, and, for two vendors, a support chat transcript used to test response time. Link each source so a claim like “vendor B’s support responded in under ten minutes” can be checked against the transcript.
What didn’t go as planned. For this task, imagine one vendor’s pricing page required a login to see current rates, so the agent used a cached price from a review site instead and flagged it as unverified rather than presenting it with the same confidence as the other four. This is the single most important thing a weak receipt tends to omit, because it’s the part most likely to matter if the comparison turns out to be wrong later.
Any approval points. For a research-only task like this, there may be none, since nothing was sent, paid or booked. The receipt should state that no approval was required, so you can confirm the agent did not take an action that needed your review.
The final result, with a clear link back to the underlying comparison, and a note on what’s solid versus what carries a caveat, like that one unverified price.
Time and cost, where relevant. For a task like this, how long it took and, if different models or paid lookups were used along the way, roughly what it cost to run. This helps with recurring tasks, where costs compound, and gives useful context for a one-off.
What a weak version of this looks like
Contrast that with a receipt that only says: “Completed: compared 5 project management tools, summary attached.” That’s not wrong, exactly, but it hides the one detail, the unverified price, that would matter most if the comparison were later used to make a real decision. It also gives you nothing to check if a colleague asks where a specific number came from. The difference between the two versions isn’t length for its own sake, it’s whether the uncertain parts are visible or smoothed over.
A checklist for red flags in a receipt
- No links to sources, only conclusions stated as fact.
- No distinction between what succeeded and what didn’t, everything presented with the same confidence.
- No mention of approvals, or where they happened, even for tasks that clearly touched something consequential.
- Written for a machine, not a person: raw timestamps and system events with no plain-language summary.
- No sense of cost or time, especially for a task that runs repeatedly.
If a receipt has more than one or two of these, it’s closer to a system log than something you can rely on to check the work later.
The stakes rise as tasks get less supervised
The less you watch a task happen in real time, which is the whole point of delegating recurring or overnight work to an AI operator, the more the receipt is doing the job your own attention used to do. A receipt that hides uncertainty or skips failed steps isn’t just an inconvenience; it’s the exact place where a real mistake would go unnoticed.
OperatorNest’s receipts record the request, sources accessed, steps that went wrong, approvals and decisions, and the final result. A receipt records what the operator did and why; it cannot establish whether a source was correct. Check a flagged, unverified figure like the one above before relying on it. See also what an AI agent receipt is for the shorter definition this post expands on.