Introduction

Software engineering has undergone a radical shift in the last year. AI-powered coding agents now handle substantial portions of code generation, yet the verification lag remains. When agents produce features at speed, human engineers still face a tedious QA loop--manually clicking through interfaces, confirming every interaction works as intended. This article explores how one team built a navigable map of their product, enabling AI agents to test features autonomously and dramatically reduce verification time.

What Happened

Vincent noticed that even with powerful agents like Claude Code, a persistent bottleneck emerged: agents spent excessive time rediscovering how product pages function. The typical cycle--screenshot, interpret, act--consumed tokens and slowed delivery. Seeking a better path, Vincent encountered Lauren Tan's work on agentic software engineering, where verification takes center stage. Feature maps, described as markdown files detailing product workflows, showed promise: test time dropped from 25 to 15 minutes per feature. However, flat feature lists struggled with nested overlays, cross-surface journeys, and the multi-persona complexity of a live marketplace. The need grew for a structured, navigable map that could guide agents without redundant rediscovery.

Why This Matters

The solution took shape as a product-surface graph--an organized map where each node represents a user-occupiable surface: a page, modal, sheet, drawer, or region. Unlike simple page-level maps, this graph captures in-page transitions, modal hierarchies, and tab regions through edges defined by clicks, links, and submissions. The schema distinguishes surface kinds and ties each node to its underlying React components via paths and export names. By indexing surfaces recursively and organizing them into manageable domains, the graph avoids context rot while preserving the product's navigable structure. The result is a machine-readable map that lets agents navigate directly to intended workflows instead of stumbling through screenshots.

Key Takeaways

  • Subagents can claim domains, explore React trees, and spawn grandchild agents when component counts exceed thresholds, keeping indexing scoped and efficient.
  • Coverage tracking links each graph node to the source files that built it, using content hashes to detect drift after every change.
  • psg stale identifies surfaces needing refresh by comparing current file hashes against stored receipts, enabling targeted reindexing instead of full rewalks.
  • In practice, agents equipped with the graph reduced testing token usage by roughly threefold and cut execution time by half compared to screenshot-first workflows.
  • Maintaining the map requires a disciplined loop: update nodes when associated code changes, enforce claim locks to prevent collisions, and validate coverage after every PR.

Conclusion

AI coding agents can ship features faster than ever, but without a reliable map, verification becomes the bottleneck. A product-surface graph bridges that gap by giving agents a persistent, navigable understanding of where things live, how to reach them, and what to test. When kept in sync through automated stale detection and recursive indexing, the map becomes a living artifact that pays dividends every time a new feature lands. For teams building AI-native development lifecycles, the graph isn't just a nice-to-have--it's the foundation of trustworthy, scalable QA.