Workflows & Automation21ms Reflex$0 / MIT License
8,700 active downloads
Ollaya
Fast-path non-autoregressive classifier bridge for Ollama and Llama 3
jevproxy // ollaya (reflex kernel)
$npx jevproxy run ollaya
JevProxy Intercept Resolved(Mode: DIRECT_REFLEX)
21ms|Cost: $0.0001|0 reasoning tokens burned
STDOUT • Tool Call Output:Status: 200 OK
agent.dispatch(ollaya_route)
payload: { "prompt": { "type": "string", "description": "Input text" }, "local_model": { "type": "st...
✔ Decision short-circuited in 21ms without roundtrip to frontier LLM.
Upstream token bill saved: $0.0120 on this turn.
Ollaya
Compatible with Cursor, Claude Code, Windsurf, OpenCode
TRADITIONAL LLM CALL:UNOPTIMIZED
• Median Latency: 1350 ms
• Cost per Turn: $0.0120
• Mode: Full KV Cache Reload & TTFT Prefill
JEVPROXY REFLEX KERNEL:77x FASTER
• Median Latency: 21 ms
• Cost per Turn: $0.0001
• Accuracy: 97.1% deterministic
1-Click CLI Execution
Run this tool accelerated through the JevProxy gateway without manual wiring:
npx jevproxy run ollayaOpenAI / Anthropic Tool Schema
JSON SpecificationPaste this schema into your agent tools definition or Cursor extensions:
{
"type": "function",
"function": {
"name": "ollaya_route",
"description": "Classify request and decide whether to handle locally or forward to edge",
"parameters": {
"type": "object",
"properties": {
"prompt": {
"type": "string",
"description": "Input text"
},
"local_model": {
"type": "string",
"description": "Configured Ollama model tag"
}
},
"required": [
"prompt"
]
}
}
}Technical Architecture & Usage
Hybrid Local and Edge Architecture
Running local models on laptops or workstations often introduces thermal throttling during long tool evaluation loops. Ollaya offloads decision classification to JevProxy in 21ms.
Accelerate Ollaya with JevProxy
Get 5,000,000 free decision tokens. Eliminate the 3-second tool freeze in Cursor and Claude Code in under 60 seconds.