Postman’s Agent Mode enables developers to interact with the Postman product in an AI-native way, streamlining tasks across API testing, documentation, discovery, and implementation. This involves re-engineering how agents interpret and interact with a product that has evolved over 11 years with a user interface-driven approach.
Building and operating an AI agent for a global community of 40 million developers presents unique engineering challenges. Postman addressed these by integrating Agent Mode with Amazon Bedrock, which provides managed access to foundation models. This strategic choice allows Postman to handle variable, latency-sensitive demand and sharp traffic bursts without managing its own model-serving infrastructure, while maintaining flexibility in model selection, control over throughput, geographic processing, and cost.
Overcoming Tool Sprawl for Agent Efficiency
One significant challenge Postman encountered was 'tool sprawl.' Initially, the team designed highly atomic tools for small, precise actions, such as opening a request or updating a single field. While this approach offered correctness and control, it led to long sequences of tool calls for common workflows, resulting in slow user experiences. Users observed the agent performing actions step-by-step that they mentally grouped as a single operation.
Testing revealed that agent effectiveness diminished once the visible toolset exceeded approximately 40 tools. Agents would make errors like calling nonexistent tools, passing incorrect arguments despite valid schemas, or selecting contextually inappropriate tools. Even with larger or newer models, this behavior persisted. To mitigate this, Postman’s current architecture dynamically scopes the tools exposed to the model per task, isolating individual execution threads. This ensures the model only sees the tools relevant to the current task, improving efficiency and reducing errors.
Furthermore, many client APIs were implicitly tied to interface state. Tools that modified requests often required specific elements to be open or created new tabs as side effects. This forced the agent to mimic user interface interactions, such as opening a request tab to read it, rather than reasoning directly about data. Postman is actively decoupling tools from tabs, allowing Agent Mode to send requests in the background without an open tab, though user approval is still required for actions that modify application state.
Leveraging Schema-Based Reads for Data Access
For products like the API Catalog, which expose structured data such as service uptime, test results, and endpoint response times across many services, Postman consolidated multiple narrow views into a single query tool. By providing agents with the schemas of underlying ClickHouse tables, the agent can generate complex queries involving joins and WHERE clauses. This significantly reduces the number of distinct tools needed to answer analytical questions.
This approach shifts the engineering effort from building a specific tool for every question to modeling the data effectively once. The agent can then generate a far broader variety of queries than could be enumerated as individual tools. For instance, an agent can generate a SQL query to retrieve service IDs, total requests, errors, error rates, average latency, and P95 latency for services within a seven-day period, filtering for specific latency and request conditions.
Context as the Primary Bottleneck
Postman initially anticipated that a lack of tools would be the primary impediment to agent performance. However, practical experience revealed that the real bottleneck was often missing or insufficient context. Agents require relevant information to reason effectively and perform tasks accurately. Addressing this involved refining how context is provided to the models, ensuring that agents have access to the necessary data without being overwhelmed by irrelevant information. This also includes using responsible AI controls like Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the large language model.