An agent should leave you with proof
Our testing method: a real task, explicit permissions, a result you can inspect, and a record of the work still left to you.
An agent saying “done” is the beginning of verification. For a document task, we need the document. For a code change, we need the change and a relevant check. For a booking, we need the confirmation from the service that accepted it.
Agent Report will evaluate agents through tasks with inspectable outcomes. This is our proposed method. It is not a claim that we have already run a benchmark or tested every product we cover.
Give the task a finish line
Before the run, write down what success would look like. “Research a trip” leaves too much room for interpretation. “Find three options for these dates, check the total prices and cancellation terms, and link the booking pages” has a result a reader can inspect.
The task should represent something a person actually needs. We will publish its constraints with the result so readers can judge whether the trial resembles their own work.
Record the environment
Each hands-on report should name the product, trial date, plan, relevant version, device and connected services. We will state when a feature is unavailable to us. If a model or setting can affect the result, that belongs in the report too.
Availability changes. A successful run in one account is evidence about that run, not a promise that every reader has the same access.
Separate setup from execution
Connecting accounts, configuring permissions and learning the interface all take time. We will record those steps separately from the agent’s running time, and distinguish active supervision from waiting.
A ten-minute task that needs twenty minutes of preparation can still be useful repeatedly. Readers need both numbers to decide.
Count the help
We will record corrections, repeated instructions, manual recovery and steps completed outside the product. If the tester rescues a run, the rescue stays in the story.
We will also identify the point at which the user must make a decision: approving a recipient, a purchase, a file change or an external action. Approval should describe what will happen and keep the scope understandable.
Inspect the result where it lives
A summary of an action is not its receipt. We will verify the actual file, application, destination or service when the task permits it. If we cannot inspect the outcome, we will mark it unverified.
A failed step is useful evidence. We want to know whether the agent detects the failure, explains the remaining work and leaves the user able to recover.
Repeat before generalizing
A single trial can establish that something happened once. It cannot establish a reliability rate. We will disclose trial counts and avoid turning one successful demonstration into a numerical score for the whole product.
When a comparison is practical, we will keep task conditions aligned and explain differences in access, setup or capabilities. We will publish observations before awarding rankings.
Make the evidence easy to read
Our articles will distinguish four kinds of information:
- Vendor claim: a capability described by the company.
- Source review: our reading of linked documentation or announcements.
- Hands-on observation: an outcome we directly inspected during a disclosed trial.
- Analysis: our interpretation, including uncertainty and questions for further testing.
Every article should say which kind of work it contains. Material corrections should carry a dated note. Relationships to products we cover should be disclosed where they affect the article.
Our job is to help you decide what is worth trying, what deserves access, and what still needs your attention.
Published by Agent Report. Sources reviewed October 7, 2026.
Browse the other reports