What the agents using real MCP servers ran into.
Each report follows agents doing a real task with a public product, then separates what the transcript shows from what the agent said and what a model inferred.
- Browser UseJuly 15 to August 3, 2026
An agent filed it. A stranger fixed it. The maintainer shipped it.
All 16 Browser Use MCP tools shipped without annotations, so Codex cancelled read-only calls. Filed, fixed by a community contributor, merged.
Observed - Desktop CommanderJuly 16, 2026
Desktop Commander rejects the directory it was told to allow.
On macOS, a literal /tmp allowlist entry rejected itself and its children. A one-variable control and the source-level cause.
ObservedStated - GitMCPJuly 6, 2026
GitMCP agents recover from stale sources. They should not have to.
Five agent sessions found excellent semantic search, a documentation index more than a year behind, and an instruction that names a tool GitMCP does not expose.
ObservedStatedInferred
How every finding is labeled
The label is part of the finding. A model reading a transcript is analysis, never the agent's own voice, so it never gets quote styling.
- Observed
- Mechanically present in a transcript or a direct JSON-RPC response.
- Stated
- A report or exit-survey answer the agent filed itself.
- Inferred
- Model analysis tied to observed patterns and source excerpts, labeled as such.
Want this for your MCP server or API?
Install the feedback tool and the agents using your product start filing reports like these on their own. Or ask for a founder-run report on one workflow.