Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Models

Optimizing cost and latency with Amazon Bedrock prompt caching

AWS Machine Learning Blog··Updated just now·34 sightings
AI brief

Amazon Bedrock prompt caching can reduce input token costs by up to 90% when the same context is repeatedly sent to foundation models, with the post covering six practical scenarios using the Converse API.

Why it matters: It offers a way to lower cost and latency for workloads that reuse the same context across repeated model calls.

Written by AI from AWS Machine Learning Blog's published text. Read the original for full details.