RAG & Search13.5ms Reflex$0 / Apache 2.0
22,400 active downloads
Semantic Cache Lookup
Sub-15ms deterministic similarity cache for repeated agent prompts
Author: CacheReflex
GitHub Repository →jevproxy // semantic-cache-lookup (reflex kernel)
$npx jevproxy run semantic-cache
JevProxy Intercept Resolved(Mode: DIRECT_REFLEX)
13.5ms|Cost: $0.0001|0 reasoning tokens burned
STDOUT • Tool Call Output:Status: 200 OK
agent.dispatch(query_semantic_cache)
payload: { "prompt": { "type": "string", "description": "Prompt string to match" }, "similarity_thresh...
✔ Decision short-circuited in 13.5ms without roundtrip to frontier LLM.
Upstream token bill saved: $0.0140 on this turn.
Semantic Cache Lookup
Compatible with Cursor, Claude Code, Windsurf, OpenCode
TRADITIONAL LLM CALL:UNOPTIMIZED
• Median Latency: 1320 ms
• Cost per Turn: $0.0140
• Mode: Full KV Cache Reload & TTFT Prefill
JEVPROXY REFLEX KERNEL:77x FASTER
• Median Latency: 13.5 ms
• Cost per Turn: $0.0001
• Accuracy: 99.1% deterministic
1-Click CLI Execution
Run this tool accelerated through the JevProxy gateway without manual wiring:
npx jevproxy run semantic-cacheOpenAI / Anthropic Tool Schema
JSON SpecificationPaste this schema into your agent tools definition or Cursor extensions:
{
"type": "function",
"function": {
"name": "query_semantic_cache",
"description": "Lookup cached prompt result by embedding similarity",
"parameters": {
"type": "object",
"properties": {
"prompt": {
"type": "string",
"description": "Prompt string to match"
},
"similarity_threshold": {
"type": "number",
"description": "Cosine similarity bound (default 0.96)"
}
},
"required": [
"prompt"
]
}
}
}Technical Architecture & Usage
13ms Response Caching
Over 30% of agent turns in team codebases ask identical questions about internal APIs. Semantic Cache serves responses in 13.5ms.
Accelerate Semantic Cache Lookup with JevProxy
Get 5,000,000 free decision tokens. Eliminate the 3-second tool freeze in Cursor and Claude Code in under 60 seconds.