Cut Amazon Bedrock inference cost (2026)

Akal Cloud Updated 9 min read

Quick answer

Three documented levers. Batch inference is priced 50% below on-demand but supports no tool calling and no structured output. Prompt caching bills cache reads at a per-model rate: AWS states 75% below on-demand input for Amazon Nova and a 90% discount for GPT-5.6, and cache writes can cost more than uncached input. Provisioned Throughput buys model units by the hour, with 1-month and 6-month commitment rates where published. AWS states prompt caching is not supported with the batch inference API, so agentic workloads must choose caching.

There are three documented ways to pay less per token on Amazon Bedrock, they save different amounts, and two of them cannot be used together. The combination that looks best on paper, batch plus prompt caching, is explicitly unsupported.

How is Amazon Bedrock inference priced per token?

On-demand inference bills per million input tokens and per million output tokens, at a rate published per model on the Amazon Bedrock pricing page. Read the rates there rather than from any article: they change per model release, and a stale token price is worse than no price. What does not change weekly is the structure of the discounts, which is what this post is about.

LeverDocumented savingWhat you give up
Batch inference "50% lower price compared to on-demand inference pricing" Synchronous responses, tool calling, structured output, prompt caching
Prompt caching A per-model cache-read rate: 75% below on-demand input for Amazon Nova, a 90% discount for GPT-5.6 Nothing structural, but cache writes can cost more than uncached input
Provisioned Throughput Hourly rate per model unit, with 1-month and 6-month commitment rates for the models that publish them Elasticity. You pay for the units whether or not you use them

When is Bedrock batch inference worth it?

The discount is the largest single lever available and the pricing page states it flatly: Bedrock "offers select foundation models (FMs) from leading AI providers ... for batch inference at a 50% lower price compared to on-demand inference pricing". Note "select": it is not every model, and the pricing page carries the model list.

The batch inference documentation describes the shape of the job. You write prompts into files, upload them to S3, submit one request, and collect output files from S3 when it finishes. That is a different program, not a flag.

Three limitations decide most cases, the first two stated on that page and the third on the prompt caching page:

  • No tool calling and no structured output. "Batch inference does not support tool calling (function calling) or structured output (response_format). Each record in the input JSONL file is processed independently without multi-turn interaction." That rules out every agentic workload, since an agent loop exists to call tools.
  • No provisioned models. "Batch inference isn't supported for provisioned models." The two commitment mechanisms do not stack.
  • No prompt caching. Covered below, and it is the one that changes the arithmetic.

Where batch does fit: classification passes, bulk summarisation, embedding generation, evaluation runs, anything you would happily run overnight. If the work is a queue rather than a conversation, it is halfway to being a batch job already.

Two operational notes from the same documentation. Quotas are separate from on-demand ones and live in the Bedrock endpoints and quotas reference, so a large job can queue behind its own account limits. And you do not have to poll for completion: AWS supports EventBridge notifications on batch job state changes, which is the difference between a scheduled pipeline and a script that sleeps.

How much does Bedrock prompt caching save?

It depends on the model, and there is no single ratio to quote. The pricing page's footnote "Cache read input tokens will be 75% less than on-demand input token price" appears under the Amazon Nova tables and nowhere else. The Anthropic tables instead list their own "Price per 1M input tokens (5m cache write)", "Price per 1M input tokens (1h cache write)" and "Price per 1M input tokens (cache read)" columns, rendered per model. The prompt caching documentation states one model family's rates in prose: for GPT-5.6, "Tokens written to cache are billed at 1.25× the uncached input token rate. Cache reads are billed at a 90% discount compared to uncached input tokens." For OpenAI models before GPT-5.6 it states "No cache write fee".

The mechanism is the same everywhere. AWS describes caching as a way "to reduce inference response latency and input token costs", and the Bedrock prompt caching product page explains why it is cheaper: "This cache lets the model skip recomputation of matching prefixes." That page puts the ceiling at "up to 90%" on cost. Read it as a ceiling: Amazon Nova's stated cache-read discount is 75%.

The documentation also states the cost that competitors' summaries tend to omit: "Depending on the model, tokens written to cache can be billed at a rate that is higher than the standard input token rate." On GPT-5.6's documented rates, a prefix written once and never read costs 25% more than not caching it. Written once and read once, the two uses cost 1.35 times the uncached input rate instead of 2 times, so a single cache hit already pays. The losing case is not low reuse. It is no reuse: a cache that expires before its first hit.

Reuse is bounded by the TTL: "The cache has a Time To Live (TTL), which resets with each successful cache hit ... If no cache hits occur within the TTL window, your cache expires. Many models support a 5-minute TTL." A prefix reused every four minutes stays warm indefinitely, and for Anthropic models AWS notes the 5-minute cache "will continue to be refreshed at no additional charge". One reused every twenty minutes is rewritten every time on a 5-minute TTL. The documented fix is a longer TTL where the model offers one: the model table lists 1 hour for most current Claude models and 30 minutes for GPT-5.6, and on the Anthropic tables the 1-hour write is its own price column.

The silent failure mode

Cache checkpoints have a minimum size that varies by model, and missing it does not raise an error. From the documentation: "You can only create a cache checkpoint if your total prompt prefix meets the minimum number of tokens ... If you add a cache checkpoint before meeting the minimum number of tokens, your inference still succeeds, but your prefix isn't cached."

So a checkpoint set below the minimum produces a working application with no caching and no signal that caching is off. The minimum is per model and the spread is wide. AWS's own example: "Claude Opus 5 requires at least 512 tokens per cache checkpoint, Claude Sonnet 5 requires at least 1,024 tokens per cache checkpoint, and Claude Haiku 4.5 requires at least 4,096 tokens per cache checkpoint." Across the model table the minimums run from 512 to 4,096, an eightfold range.

Two consequences worth designing around. Changing model changes the minimum, so a model swap can silently disable caching that was working. And a successful call proves nothing about caching. AWS says so directly: "Support for prompt caching doesn't guarantee a cache hit for any request. Check the cache usage fields in the model response to determine whether tokens were read from or written to cache." The same page calls Implicit Prompt Caching, the mode that needs no checkpoints, "best effort". Check the usage fields, and check them again after any prompt or model change.

Where you set the checkpoint depends on the API. Caching works through the Converse and ConverseStream APIs, whose response carries cacheReadInputTokens and cacheWriteInputTokens, through InvokeModel, and through Bedrock Prompt management, where it is a console choice of what to cache (None, Tools, system instructions, messages) rather than a request field. The APIs give the finer control: AWS notes that "For models that support Explicit Prompt Caching, the APIs provide granular control over the prompt cache."

Caching and cross-Region inference

These interact, and not in your favour. The documentation notes that cross-Region inference "automatically selects the optimal AWS Region within your geography to serve your inference request", then adds: "At times of high demand, these optimizations may lead to increased cache writes." Cache writes are the expensive half. Under load, the mechanism that improves availability degrades the caching discount.

Why can't you combine Bedrock batch and prompt caching?

Because AWS says so, in one sentence on the prompt caching page: "Prompt caching is only supported for on-demand inference endpoints. It is not supported with the batch inference API."

That forces a real choice on any workload with a large shared prefix, which is most retrieval-augmented and document-analysis work: a 50% cut on input and output, against a cache-read discount on the cached portion of input only (75% for Amazon Nova, 90% for GPT-5.6).

Workload shapeChooseWhy
Large shared prefix, many queries against it, latency matters Prompt caching Discounted cache reads on the reused input, and the responses stay synchronous
Every prompt different, no deadline, no tools Batch 50% off input and output, and there is nothing to cache anyway
Agentic loop with tool calls Prompt caching Batch does not support tool calling at all, so the choice is made for you
Steady, predictable, high-volume traffic on one model Provisioned Throughput Commitment pricing, though it excludes batch entirely

Any agent that calls tools lives in the third row: batch is unavailable to it at any price. That is worth saying out loud, because the "50% off with batch" advice is everywhere and it is inapplicable to a large share of the workloads reading it.

The prompt caching page muddies this in one place. It recommends the 1-hour TTL for "longer-running sessions or batch processing scenarios", on the same page that says caching is not supported with the batch inference API. Read "batch processing" there as a long run of on-demand calls, not a batch inference job.

There is one more option that reads like a fourth lever and is not: the OpenAI-compatible Batch API Bedrock also exposes for its OpenAI models. AWS documents it as a way to "run a batch inference job", filed under batch inference, so it is a different interface to the same mechanism. Nothing on that page exempts it from the batch limitations above, and the OpenAI section of the pricing page carries the same footnote as other providers: "Flex tier and Batch pricing is at 50% discount to Standard tier pricing."

When does Bedrock Provisioned Throughput beat on-demand?

Provisioned Throughput is bought in model units billed per hour. For most providers that list it, the pricing page publishes three columns per model: "Price per hour per model unit with no commitment", "for 1-month commitment", and "for 6-month commitment". Not every provider publishes them: the Anthropic section says "For Provisioned Throughput pricing, please reach out to your account team."

Where the rates are printed, longer commitment means a lower hourly rate. Cohere Command is listed at "$49.50", "$39.60" and "$23.77" per model unit hour, so the 1-month term is 20% below no commitment and the 6-month term 52% below. The usual consequence follows: the units bill whether or not traffic arrives.

It is a capacity decision rather than a discount decision, and it is the only option for some custom models. The pricing page notes that "For Full Rank fine-tuned models, Provisioned Throughput (PT) is the available option", while parameter-efficient fine-tunes can choose either. Where you have that choice, the arithmetic is the same one as any commitment: measure sustained utilisation first, because unused model units are the AI equivalent of an over-provisioned Reserved Instance.

How do you attribute the remaining Bedrock spend?

Cutting the rate is half the job. The other half is knowing which team spent it, and CloudWatch cannot tell you: ModelId is the only dimension on its Bedrock metrics. The two mechanisms that can are IAM principal cost allocation and application inference profiles, set out in Amazon Bedrock cost per team. If the agent calling the model is hosted on AgentCore, its compute and memory are a separate bill from these tokens, covered in Amazon Bedrock AgentCore pricing.

Once the data is in the CUR, the columns behave like any other service, with the same choice of cost basis described in which CUR cost column you should be summing. Provisioned Throughput commitments amortize; on-demand tokens do not. And because a bad prompt change can multiply token spend overnight with no infrastructure change to blame, anomaly detection rather than a budget is the alarm that catches it, subject to the detection lag documented there.

Share LinkedIn X Hacker News Reddit

See this on your own bill

Akal Cloud connects in about two minutes and shows the same numbers against your real AWS accounts.

Get started on AWS Marketplace

Related reading