Microsoft Research introduced AgentRx on October 2, 2026, a benchmark and research effort focused on diagnosing failures in AI agent executions. The project starts from a practical problem: when an agent fails after a long sequence of model calls, tools and subagents, it can be difficult to determine exactly where the workflow went wrong.

Agent failures are not always simple model errors. A system can fail because the model misunderstood a goal, a tool returned unexpected information, a previous step produced a bad state, or multiple agents interacted in an unreliable way. Diagnosing the failure requires looking at the complete execution trajectory.

Why agent failures are harder to debug

A conventional software function usually follows a relatively clear path. If it receives an input and returns the wrong output, developers can inspect the code and reproduce the behavior.

Agents add probability and interaction. The model may choose a different tool on another run, interpret the same instruction differently, or respond differently to a tool result. A long-running agent may also make an early mistake that does not become visible until much later.

That makes the final answer an incomplete debugging signal. Developers need to understand the sequence that produced it.

What AgentRx studies

Microsoft Research says AgentRx manually annotates failed agent runs and releases a benchmark containing 115 failed trajectories. The trajectories span structured API interactions and are intended to help researchers study where and why agent executions break down.

The focus is diagnosis rather than simply ranking which model produces the best final answer. That distinction matters because two agents can have the same final failure for completely different reasons.

From outcome evaluation to trajectory evaluation

Many AI evaluations focus on whether the final answer is correct. That is useful, but it can hide the mechanism of failure.

Suppose an agent is asked to update a configuration file. It might inspect the wrong file, infer an incorrect requirement, make a technically valid change, and then pass a test that does not cover the real problem. The final output may look reasonable while the reasoning path is flawed.

Trajectory-based evaluation asks a different question: at which stage did the execution stop being reliable?

Why tool use matters

Agent systems interact with tools that have their own failure modes. A search API can return incomplete information. A database query can return an empty result. A shell command can fail. A browser can encounter a login wall or a changed page structure.

The model needs to recognize those conditions instead of treating every tool response as trustworthy.

For developers, this means tool errors should be explicit. A system should distinguish “the tool returned no results” from “the tool failed” and from “the tool returned data that may be incomplete.” Those states give the agent and the debugging system better information.

Long-horizon agents need intermediate checks

One practical lesson from agent failure research is that long workflows should not rely on a single final validation step.

If an agent performs ten actions and only checks the final result, an early mistake can contaminate every later stage. Adding intermediate assertions can catch the problem closer to its source.

Examples include validating a file path before editing, checking an API response schema before passing it to another model, verifying that a test actually ran, and confirming that a requested external action succeeded.

How teams can apply the idea

  1. Store execution traces for important agent workflows.
  2. Record model decisions and tool results in a structured format.
  3. Mark the first point where the workflow diverged from the expected state.
  4. Separate model errors from tool errors and environment errors.
  5. Build a small library of real failed trajectories.
  6. Use those failures as regression tests after changing prompts, tools or models.

This turns production failures into evaluation data instead of isolated debugging sessions.

What a useful agent trace should contain

A trace does not need to store every piece of user data. Teams should collect the minimum information required to understand execution.

Useful fields can include the task identifier, model version, tool name, tool status, structured input and output summaries, timestamps, validation results, retries and the final outcome.

Where sensitive data is involved, logging should be designed around privacy and access controls rather than copying entire conversations into a permanent database.

Why failure taxonomies matter

Once a team has enough traces, it can classify failures. Categories might include planning errors, incorrect tool selection, invalid tool arguments, misunderstood tool results, state drift, missing validation, permission failures and model hallucination.

A taxonomy helps because different failure types require different fixes. A prompt change will not solve a broken API, and a better model will not necessarily solve an incorrect permission policy.

Agent evaluation should measure recovery too

A robust agent does not have to be perfect. It needs to recognize problems and recover safely.

Evaluation should therefore ask whether an agent can notice a failed tool call, retry with corrected parameters, ask for clarification when necessary, or stop before creating an unsafe side effect.

Recovery behavior is especially important for long-running agents because a small failure does not always justify abandoning the entire task.

The broader significance of AgentRx

As agents become more capable, debugging them becomes closer to debugging distributed systems than debugging a single text-generation call. There are multiple components, asynchronous steps, external dependencies and probabilistic decisions.

AgentRx is useful because it frames failure diagnosis as a first-class engineering problem. The industry needs benchmarks that explain why systems fail, not only which model wins a leaderboard.

Sources

About Deepak Parmar

Deepak Parmar is an SEO and automation expert with 7 years of experience in SEO, AI search, GEO, and web development. He specializes in helping brands improve visibility across Google, ChatGPT, Gemini, Perplexity, and other AI search platforms.

At ToolBoxKart, Deepak writes about SEO, AI, automation, search technology, and practical digital workflows, combining hands-on technical experience with real-world research and experimentation.

LinkedIn · YouTube

Latest published posts