JEV REGISTRY • VERIFIED TOOL|RAG & SEARCH| 18MS EDGE REFLEX
Home/Hub/RAG Relevance Scorer
RAG & Search18ms ReflexFreemium / API Key
17,200 active downloads

RAG Relevance Scorer

Non-autoregressive relevance scorer for retrieved context documents

jevproxy // rag-chunk-filter (reflex kernel)
$npx jevproxy run rag-filter
JevProxy Intercept Resolved(Mode: DIRECT_REFLEX)
18ms|Cost: $0.0001|0 reasoning tokens burned
STDOUT • Tool Call Output:Status: 200 OK
agent.dispatch(filter_rag_chunks)
payload: { "query": { "type": "string", "description": "User search query" }, "chunks": { "type": "a...
✔ Decision short-circuited in 18ms without roundtrip to frontier LLM.
Upstream token bill saved: $0.0180 on this turn.
RAG Relevance Scorer
Compatible with Cursor, Claude Code, Windsurf, OpenCode
TRADITIONAL LLM CALL:UNOPTIMIZED
• Median Latency: 1620 ms
• Cost per Turn: $0.0180
• Mode: Full KV Cache Reload & TTFT Prefill
JEVPROXY REFLEX KERNEL:77x FASTER
• Median Latency: 18 ms
• Cost per Turn: $0.0001
• Accuracy: 98.5% deterministic

1-Click CLI Execution

Run this tool accelerated through the JevProxy gateway without manual wiring:

npx jevproxy run rag-filter

OpenAI / Anthropic Tool Schema

JSON Specification

Paste this schema into your agent tools definition or Cursor extensions:

{
  "type": "function",
  "function": {
    "name": "filter_rag_chunks",
    "description": "Score retrieved documents and prune irrelevant context chunks",
    "parameters": {
      "type": "object",
      "properties": {
        "query": {
          "type": "string",
          "description": "User search query"
        },
        "chunks": {
          "type": "array",
          "items": {
            "type": "string"
          },
          "description": "Retrieved text candidates"
        },
        "top_k": {
          "type": "number",
          "description": "Number of winning chunks to preserve"
        }
      },
      "required": [
        "query",
        "chunks"
      ]
    }
  }
}

Technical Architecture & Usage

Cutting Context Bloat Before Generation

Injecting 10 vector chunks into an agent prompt often wastes 8,000 tokens on irrelevant text. RAG Relevance Scorer filters down to the top 2 in 18.0ms.

Accelerate RAG Relevance Scorer with JevProxy

Get 5,000,000 free decision tokens. Eliminate the 3-second tool freeze in Cursor and Claude Code in under 60 seconds.