Public EcoRouter Knowledge · AI-generated content ·
How does semantic caching decide whether a new prompt can reuse an earlier answer?
The question
How does semantic caching decide whether a new prompt can reuse an earlier answer?
Semantic caching decides whether to reuse an answer by focusing on the meaning of a prompt rather than its exact wording. Here is the typical process:
- Embedding: When a new prompt arrives, the system converts it into a numerical vector (an embedding) that represents its semantic meaning.
- Comparison: The system calculates the distance (similarity) between this new vector and the vectors of prompts already stored in its cache. This is usually done using methods like cosine similarity.
- Thresholding: The system checks if the similarity score exceeds a predefined threshold. If the new prompt is "close enough" to a cached prompt, the system assumes they are asking for the same information.
- Retrieval: If the threshold is met, the system returns the corresponding cached answer. If it falls below the threshold, the system treats the prompt as new, generates a fresh answer, and then stores that new prompt-answer pair in the cache for future use.
By using this approach, the cache can recognize that "What is the capital of France?" and "Tell me the capital city of France" are effectively the same request, even though the text strings are different.
Answering this used an estimated 90% less than EcoRouter’s configured frontier baseline
- Routed to
- Efficient
- Tokens
- 208 in / 247 out
- Cost
- $0.0004
- Baseline
- $0.0043
Figures recorded by EcoRouter when this answer was generated, and fixed at that moment. Cost comparisons are against a configured reference model, not a measurement of electricity, carbon or water.
Ask about this
Ask EcoRouter a follow-up using this Knowledge as context. Nothing becomes public unless you choose to publish it.
Ask a follow-upComments
No comments yet.