Postman and AWS have published the architecture behind Agent Mode, an AI-native way to work across API testing, documentation, discovery and implementation. The system runs on Amazon Bedrock for a community of 40 million developers, illustrating the engineering gap between an agent demonstration and a global production service.
Postman expected model quality and prompt design to be the hardest problems. In practice, the deeper challenges were tool sprawl, product context and assumptions embedded in an interface developed over 11 years. The resulting design offers useful lessons for teams adding agents to mature software.
More tools made the agent less reliable
Early versions exposed many small, precise tools for actions such as opening a request, updating a field or retrieving metadata. Long workflows then required repeated model turns, making the experience slow. Postman also found that tool-selection errors rose when the visible set exceeded about 40 tools.
The production architecture uses embeddings to narrow more than 170 tools to roughly 15 relevant choices for a request. Those tools are handed to a context-isolated subagent. A smaller menu improves selection and reduces the chance that the model invents a tool or chooses one whose name seems suitable but whose behaviour is wrong.
Application tools had to escape the interface
Some internal APIs were coupled to visible interface state. A tool might need a request tab to be open or create a new tab as a side effect, forcing the agent to imitate clicks rather than work directly with data. Postman is decoupling these functions so operations can run safely in the background.
This is a common legacy-product problem. Human interfaces often carry state implicitly through the screen, while an agent needs explicit identifiers, schemas and permissions. Product teams should treat agent access as an API design exercise, not simply attach model calls to existing interface automation.
Context was a bigger constraint than capability
Missing or incomplete context caused more failures than missing tools. An agent needs to know which workspace, collection, request and environment are active, along with what the user has selected and what earlier steps established. Supplying the application’s rendering data model did not solve the problem because it contained detail shaped for display rather than reasoning.
Postman built dedicated context handlers that distil each entity into relevant information. It separates broad, shallow background context from deep context selected by the user. The design treats the context window as a scarce budget where irrelevant data can crowd out important evidence long before the technical limit is reached.
Knowledge is retrieved when the task needs it
The product surface includes many protocols, mock servers, monitors, documentation tools, governance controls and request settings. Encoding all of that knowledge in a static system prompt would be expensive and mostly irrelevant to any single question.
Postman seeded a retrieval system with concise, feature-specific articles derived from its Learning Center. At runtime, Agent Mode selects material according to the query and active context. Documentation can evolve with the product, giving the agent targeted knowledge without expanding every prompt.
Human approval remains in the action path
Agent Mode requires approval before actions that modify application state. Tools are scoped to the task, and Amazon Bedrock Guardrails can redact personally identifiable information before content reaches the underlying model. Enterprise administrators can control that protection through guardrail settings.
These controls acknowledge that relevance and safety are separate concerns. An agent may correctly understand a request while selecting an action the user did not intend. Confirmation should therefore describe the exact object and change, not present a generic approval prompt after a long chain of reasoning.
Bedrock handles models, geography and traffic
Postman uses Bedrock to select among supported Claude models according to latency, cost and reasoning needs. Cross-Region inference helps absorb bursty developer traffic, while geographically scoped profiles can limit processing to supported regions within areas such as the United States or European Union.
Geographic controls still require careful configuration of identity policies, service controls and quotas for every destination region. Postman also configures zero data retention for supported models, but notes that availability and behaviour are model-dependent and must be checked for each production choice.
Prompt caching reduces repeated work
Agent conversations repeatedly send stable material such as system instructions, core tools, knowledge and prior context. Amazon Bedrock prompt caching lets Postman reuse stable prefixes rather than process them from scratch on every turn, reducing both latency and inference cost.
Caching strategy must remain compatible with privacy and correctness. Applications should know which content is shared, user-specific or likely to change, and invalidate cached segments when instructions or permissions are updated.
A blueprint for agents inside mature products
Postman’s experience shows that production agents depend less on a giant tool catalogue than on good abstraction. Purpose-shaped context, dynamic tool selection, explicit approvals and a versioned knowledge layer can make an existing product legible to a model without reproducing every interface gesture.
Amazon Bedrock provides the scalable inference and governance foundation, but the product architecture remains decisive. Teams building similar systems should measure tool-selection errors, context omissions, approval outcomes and end-to-end task success. Those signals reveal whether an agent is becoming genuinely useful, rather than merely more capable in isolation.