Introduction
When AI agents produce unexpected results, engineers often find themselves staring at a final output with little visibility into how that answer emerged. Traditional dashboards surface the endpoint but conceal the journey, making it difficult to diagnose failures, validate assumptions, or iterate reliably. This article introduces a different approach: a local evidence debugger designed to capture structured traces alongside code, enabling deeper inspection without depending on remote platforms.
What Happened
When an agent misbehaves, engineers typically begin with one of two artifacts: the final model response or a screenshot from a monitoring product. Each offers limited value alone. Logs scatter across services, ordered by time rather than causality, and final responses conceal the sequence of tool calls, LLM invocations, and decision points. Dashboards may display runs clearly, yet attaching evidence to pull requests, reproducing traces in CI, or reviewing them without an account proves cumbersome. The author asked whether a single agent run can yield enough structured, local evidence to inspect its path, validate explicit expectations, compare against another run, and produce a reviewable artifact—prompting a fundamentally different product boundary.
Why This Matters
Hosted observability platforms excel at durable ingestion, fleet-wide monitoring, retention, alerting, and team operations—capabilities AgentInspect deliberately does not replicate on a laptop. The focus instead rests on the pre-release evidence loop that lives beside source code. Core workflow comprises local structured trace (JSONL), a readable execution tree, deterministic checks, run-to-run diffing, and a redacted evidence bundle. This artifact travels with engineering work: developers inspect it before opening pull requests, CI can reject known-bad trajectories, and reviewers can verify bundles without reproducing the entire run. The JSONL trace format remains friendly to ordinary text-processing tools and stays complementary to APM, hosted agent observability, and evaluation platforms, never promising hosted retention, production alerting, prompt management, automatic remediation, or compliance certification.
Key Takeaways
- Local-first design eliminates account friction, letting developers inspect single failing tests or attach evidence to changes immediately.
- Evidence participates in existing workflows: archived as CI artifacts, reviewed in pull requests, or passed to other systems without being trapped behind a single UI.
- Capture and export remain separate decisions; source traces may contain sensitive data, and AgentInspect provides best-effort redaction before bundle creation.
- Deterministic checks shine when run beside code and fixtures, validating structural invariants like retrieval must occur before generation in CI without waiting for dashboard anomalies.
- The tool establishes a compact review unit: representative run, explicit contract result, and share-checked evidence when needed, making engineering discussions less speculative and more concrete.
Conclusion
The author built a local evidence debugger because the goal is not merely more telemetry, it is a shorter, reviewable path from an agent's behavior to an engineering decision. Starting with a synthetic run that includes failure and fallback, inspecting the tree, adding a deterministic check, comparing with a corrected run, and building a share-profile bundle quickly reveals both the value and limits of the approach. The project and tagged release sit on GitHub, and the author welcomes issues pointing to unclear evidence, misleading defaults, or gaps between documented boundaries and real developer workflows. Traditional code review remains necessary but insufficient for agent systems; adding representative runs, explicit contract results, and share-checked evidence creates a more grounded foundation for evaluating change.




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.