Skip to main content

A 68.8% Faster Permission Gate for Claude Code

We tested Jev through OpenRouter as a Claude Code permission gate. Median latency fell from 733 ms to 228 ms versus Haiku 4.5.

Rahul Balakavi headshot
Rahul Balakavi, Co-Founder, AmpUp
Subscribe & Share:

504.2 milliseconds is easy to ignore once. Put it in front of hundreds of tool calls and it becomes part of the feel of an agent.

We tested whether a model designed for typed decisions could handle that gate faster than a small chat model. On a private, sanitized sample of Claude Code permission situations, Jev through OpenRouter returned a median decision in 228 milliseconds. Claude Haiku 4.5 took 733 milliseconds on the same cases.

That is a 68.8% reduction in permission-decision latency, not a claim that Claude Code becomes 68.8% faster. At 500 guarded actions in an eight-hour agent workload, the measured mean difference works out to about 5.3 minutes saved, or 1.08% less total elapsed time.

We published the implementation, benchmark harness and synthetic safety fixture in claude-code-jev on GitHub.

What We Changed

Anthropic’s auto-mode architecture uses a transcript classifier running on Sonnet 4.6. It sees human messages and proposed tool calls, but not the agent’s prose or tool results. Its first stage emits a fast decision, then flagged actions receive a more expensive reasoning pass.

Claude Code does not expose a supported API for swapping out that private classifier. We did not patch or proxy it.

Instead, we installed a supported PreToolUse hook that runs before selected Bash, editing, web and delegation tools. The hook sends three things to Jev:

  • human-authored messages;
  • the proposed tool name and arguments;
  • the trusted working directory.

Jev returns allow, block or ask, plus confidence. A low-confidence answer becomes ask. A timeout, network failure or malformed response also becomes ask, preserving human review.

The Direct Test

We manually labeled eight sanitized situations from Claude Code usage, then ran each one five times through Jev and Haiku 4.5 on OpenRouter. Both models received the same user messages, proposed tool call, policy and three possible answers.

MetricJevHaiku 4.5
Decisions4040
Median latency228.4 ms732.6 ms
p95 latency327.0 ms1,761.0 ms
Mean latency249.0 ms882.9 ms
Exact-label matches15/4025/40
Cost$0.00093366$0.04973

Jev was 68.8% faster at the median, 71.8% faster by mean latency and 98.1% cheaper in this run.

The accuracy column is the reason this remains an experiment. Jev matched fewer of our three-way labels. Neither model allowed a labeled-dangerous action, but both were conservative on benign cases. Jev sometimes converted an expected block into ask, or an expected ask into block. Those are safer than an accidental allow, but they are operationally different.

The default 0.85 confidence threshold was also too strict for this sample and escalated every Jev decision. The results above use 0.60. Anyone deploying this should tune the threshold against their own held-out traffic and measure dangerous allows, benign blocks and human escalations separately.

What the Time Saving Looks Like

The measured mean difference was 0.634 seconds per guarded action.

Guarded actions per dayTime returned per dayTime returned across 220 days
1001.1 minutes3.9 hours
2502.6 minutes9.7 hours
5005.3 minutes19.4 hours
1,00010.6 minutes38.7 hours

For a heavy Claude Code user, five to ten minutes per active day is a reasonable estimate if 500 to 1,000 matched tool calls sit on the serial path. Fewer matched tools, parallel work, network variance and time spent waiting on people all reduce the realized saving.

There is a larger research scenario in the repository. An independent Jev guardrail evaluation reported roughly four seconds of model time for its LLM judges. Comparing our 264 millisecond Jev mean on the public synthetic fixture with that reference projects to 31.1 minutes returned across 500 calls, or 6.1% of an eight-hour workload. That is useful sensitivity analysis, but it is not a measurement of Claude Code’s production classifier.

Install It

You need Python 3.11 or newer, Claude Code and your own OpenRouter API key.

git clone https://github.com/RahulBalakavi/claude-code-jev.git
cd claude-code-jev
uv tool install .
export OPENROUTER_API_KEY="your-openrouter-key"

Copy the hook configuration from examples/claude-settings.json into your project at .claude/settings.json, or into ~/.claude/settings.json to use it across projects.

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash|Edit|Write|NotebookEdit|WebFetch|WebSearch|Task",
        "hooks": [
          {
            "type": "command",
            "command": "jev-auto-mode hook",
            "timeout": 5
          }
        ]
      }
    ]
  }
}

The public release passed 11 local tests. A live smoke test against OpenRouter classified all 18 synthetic cases correctly with a 190 millisecond median and a total cost of $0.00040908.

Route Claude Code Through OpenRouter Too

This is separate from the Jev permission gate, but the repository includes a small launcher for it. OpenRouter’s Claude Code guide documents the same direct connection: set the Anthropic-compatible base URL, use your OpenRouter key as the auth token and explicitly empty ANTHROPIC_API_KEY to prevent an authentication conflict.

export OPENROUTER_API_KEY="your-openrouter-key"
./bin/claude-openrouter

The launcher reads the key from the environment and never writes it to disk. Run /status inside Claude Code to confirm that the base URL is https://openrouter.ai/api and the auth source is ANTHROPIC_AUTH_TOKEN.

The Important Limits

  • This hook is an experimental permission layer. It does not replace Anthropic’s private auto-mode classifier.
  • The direct Jev versus Haiku sample contains eight situations and cannot establish safety parity.
  • The public 18-case fixture is synthetic. Passing it proves that the client and policy contract work, not that every real action will be classified correctly.
  • The hook sends user messages, tool arguments and the working directory to OpenRouter. Tool arguments can contain source code, URLs or secrets, so check your data policy before using it.
  • A faster unsafe decision is still unsafe. Keep the fail-safe ask behavior, start with a conservative matcher and calibrate on your own traffic.

Why Publish It

Agent builders spend a lot of money asking general chat models questions with tiny answer spaces: route A or B, pass or fail, allow or block. Those decisions live in the hottest part of the loop, where a few hundred milliseconds repeat all day.

Jev will not make every agent 68.8% faster. In our direct test, it made one permission gate 68.8% faster and returned about five minutes in a 500-call day. That is modest in one session and meaningful across a year.

The code and assumptions are public. Run the benchmark on your own tool calls, measure the safety errors first, then decide whether the speed belongs in your loop.

Get the code on GitHub.

AmpUp

Book a demo with us

See how AmpUp turns every call into a coaching opportunity.

Written by

Rahul Balakavi

Rahul Balakavi

Co-Founder, AmpUp

Rahul is the co-founder of AmpUp. He leads engineering and product, bringing deep expertise in building AI-powered platforms that turn sales data into actionable intelligence.

Stay up to date with AmpUp

Follow AmpUp on LinkedIn

Follow us on LinkedIn for the latest on AI-powered revenue intelligence.