A technical post published by AWS and Postman on October 9, 2026, explains how Postman built Agent Mode for a large developer audience using Amazon Bedrock. The main lesson is not that one model solves every problem: production agents need a carefully limited tool set, task-specific context, approval before state-changing actions, and deliberate choices about model routing, data retention and prompt caching.
What the case study covers
Postman Agent Mode lets users work across API testing, documentation, discovery and implementation using natural language. The AWS Machine Learning Blog describes how Postman runs it on Amazon Bedrock rather than managing its own model-serving infrastructure. The article is a technical account from the companies, not an independent benchmark or a public reference implementation; Postman says its production implementation is proprietary.
That distinction matters when reading the reported scale and test results. The article says Postman serves a community of about 40 million developers and describes internal engineering observations. Those details are useful examples, but they are not a guarantee that another product will see the same results.
1. Expose only the tools a task needs
Postman began with many small, precise tools, such as opening a request, updating a field or reading metadata. That made individual actions easier to control, but long workflows required many sequential tool calls. The team also found that the model sometimes chose a wrong or nonexistent tool when the visible catalog became too large.
In Postman's testing, tool-selection errors increased once the visible tool set exceeded roughly 40 tools. Its described architecture uses a root agent to search a catalog of more than 170 tools and narrow the list to about 15 relevant tools for a request, then hands those tools to a context-isolated sub-agent. These figures describe Postman's own experience, not a universal threshold for all models or applications.
The practical lesson is to treat tool descriptions and tool count as part of the context budget. Give an agent the smallest set of tools that can complete the task. Use clear names and argument schemas, separate read operations from write operations, and avoid exposing dangerous actions unless they are needed.
2. Use schema-aware reads instead of a tool for every question
For structured data, Postman combined several narrow views into a query tool that works with known table schemas. The agent can then compose queries that aggregate requests, errors and latency rather than calling a separate tool for every metric. This reduces the number of actions the model must choose from, while shifting more responsibility to data modeling and query validation.
A similar pattern can work in analytics systems: expose a safe, read-only query interface with an explicit schema, permitted tables, row limits and query budgets. Validate generated queries before execution. Read access is not automatically harmless; a query can still expose sensitive records or create a large bill if permissions and cost controls are weak.
3. Build context for reasoning, not for screen rendering
Postman found that missing context caused more problems than missing capabilities. Its product had accumulated years of UI assumptions: users knew which tab to open or where to find a request, but an agent needed structured information about the active entity, its state and the user's goal.
The team built context handlers that produce the information an agent needs for a particular entity instead of serializing the entire interface data model. It also uses a knowledge base with retrieval-augmented generation (RAG), selecting feature-specific documentation when the current task needs it.
For builders, this means separating broad background context from focused task context. Provide only the relevant repository, request, schema, log or document. Keep long user-generated fields from filling the context window, and retrieve more detail when a task requires it. A larger context window does not remove the need to decide what the model should see.
4. Use model routing and regional controls deliberately
The case study says Postman can use supported Anthropic Claude models through Amazon Bedrock and select a model according to workload. A faster model may suit frequent, latency-sensitive interactions, while a larger model may be reserved for tasks that need more reasoning. The article describes model selection as a configuration decision, but actual model availability still depends on Bedrock's supported models and AWS Regions.
Postman also describes using Bedrock cross-Region inference profiles to route requests among allowed destination Regions. Geographic profiles restrict routing to a defined geography; global profiles can route among supported destinations worldwide for more throughput. Those are different trade-offs. A team with a residency requirement should select a profile whose documented boundary meets that requirement and ensure its IAM and service control policies permit only the intended destinations.
Do not assume that choosing a geographic profile means inference runs inside Postman's own AWS environment. The article says processing occurs in eligible AWS Regions for the selected profile. Confirm the current model, profile, supported Regions and contractual requirements before using sensitive data.
5. Check data retention and PII controls per model
The AWS post says Amazon Bedrock does not use prompts and completions to train AWS models or distribute them to third parties. It also says Postman configured zero data retention for supported Agent Mode models using the relevant retention setting. The qualification is important: retention behavior is model-dependent, so teams should check the current Bedrock data-protection documentation for every model they plan to use rather than generalizing one configuration to the whole model family.
Postman also uses Amazon Bedrock Guardrails to redact personally identifiable information before it reaches the underlying model; enterprise administrators can enable this in Agent Mode settings. Redaction and retention controls reduce some risks, but they do not replace access controls, least-privilege permissions, logging, incident response or careful testing of prompts and tools.
6. Cache stable context to reduce repeated work
Agent conversations often resend the same system instructions, core tool definitions and product knowledge. The AWS article says Postman uses two prompt-cache checkpoints: a one-hour checkpoint for stable core context and a five-minute checkpoint for more variable context. The longer-lived cache can spread its write cost over more reads, while the shorter checkpoint suits content that changes more often.
Supported cache time-to-live options and benefits depend on the selected model. Teams should measure cache reads, cache writes, time to first token and cost for their own traffic. Caching is not automatically cheaper for every workload, especially if the prompt prefix changes frequently or the application has few repeated requests.
A practical checklist for production agents
- Map each task to a small set of read and write tools.
- Use schemas and validation for queries and tool arguments.
- Build task-specific context handlers instead of passing raw UI state.
- Require explicit user approval before actions that modify important state.
- Choose models and inference profiles based on task quality, latency, cost and permitted processing geography.
- Verify retention behavior for each model and enable PII guardrails where they fit the workflow.
- Measure prompt-cache hit rates, latency, token use and failed tool selections under real traffic.
- Keep a human review path for uncertain or high-impact actions.
For a practical approval framework, see our AI agent approval policy template. For API workflows, our guide to safe retries and fallbacks covers failure handling around model and tool calls.
Frequently asked questions
Did AWS announce a new Amazon Bedrock model in this post?
No. The October 9 article is a technical case study about how Postman uses Bedrock for Agent Mode. It focuses on architecture, routing, retention and caching.
Is Postman's production implementation open source?
No. The article says the production implementation is proprietary and is not available as a public sample repository.
Does zero data retention apply to every Claude model?
The article says the configuration is model-dependent. Confirm the current retention behavior for each model and inference configuration you plan to use.
Sources
- AWS Machine Learning Blog: How Postman runs Agent Mode for 40 million developers on Amazon Bedrock (October 9, 2026).
- Postman Docs: About Agent Mode.