- Cuts cost and latency by reusing answers to semantically similar queries
- Matches by meaning, so paraphrased questions still hit the cache
- Reduces load on expensive model inference for common questions
- Improves response speed for users asking familiar things