Ben Newton - Commerce Frontend Specialist

My Claude Code mod lets Jev choose the agent and model for every call

Jev picks the model for every subagent call in about a third of a second. On lookups it cut the cost by 51% and every fact still checked out.

My Claude Code mod lets Jev choose the agent and model for every call

I was watching Claude Code work through my iOS repo and it spun up a subagent to count the Swift files. On Opus. Then it sent another one off to find a single config line, also on Opus. Subagents just inherit whatever model the main session is on unless Claude says otherwise, so little errands like that end up on the expensive model.

So I built a mod for it. Before every agent call it asks Jev, TypeSafe's decision model, which agent type fits the task, which model is the cheapest one that'll still get it right, and how much effort it needs. All three come back in one request in about a third of a second, each with a confidence score. If Jev is sure it's a read-only lookup Haiku can handle, the mod switches the call to the Explore agent on Haiku. If it isn't sure, the call runs the way Claude set it up. There's a little band above the prompt that shows what it picked, mostly because I wanted to see it working.

The mod in about a minute: the hook, Jev's pick, and the band above the prompt.

How the mod works

It's a Claude Code mod (a plugin of function hooks that runs inside the session). The hook that matters sits on every Agent tool call. Before the subagent starts, the mod sends Jev the task description, up to 4,000 characters of the prompt, the agent types Claude Code has on offer with their descriptions, and a one-line description of each model. Haiku is "simple lookups, file counts, grep-style searches." Fable is "open-ended work with no clear path." That kind of thing.

Jev sends back three typed answers. A choice of agent type with a probability for every option, a choice of model the same way, and an effort score. No paragraph to parse. That's the whole reason I used it for this.

1 · Claude Code

Claude calls the Agent tool

With a task description, a prompt, and usually no model, so the subagent inherits the main session's model (Opus for me).

→

2 · The mod's hook

Intercepts the call before it runs

Sends the task, the agent types on offer, and the four models to Jev. Up to 4,000 characters of the prompt.

→

3 · Jev, one request

Three typed answers

Agent typeExplore 0.96
Cheapest modelHaiku 0.80
Effortlow

195 to 350 ms, about 760 input tokens.

Gate: agent == Explore && agent.conf ≥ 0.8 && model == haiku && model.conf ≥ 0.8

Passes

Rewritten to Explore on Haiku

Read-only agent, cheapest model. The only change the mod is allowed to make.

Anything else

Runs exactly as Claude set it up

Jev's pick still shows in the band above the prompt, it just isn't applied.

What happens on every subagent call. Numbers in step 3 are the real pick for "find where HealthKit authorization is requested."

The gate in the hook is four conditions and it's the only thing that can change a call:

const isLookup = mode === 'apply' && answers.agent.choice === 'Explore' && answers.agent.confidence >= 0.8 && answers.model.choice === 'haiku' && answers.model.confidence >= 0.8 && (requestedTier < 0 || requestedTier > tierOf('haiku')) if (!isLookup) return next(e) return next({ ...e, subagent_type: 'Explore', model: 'haiku' })

If Jev errors or times out, the call just runs the way Claude set it up. There's also /jev advise, which shows the pick in the band without applying anything, and /jev off. Every pick and every run gets logged to a JSON file so I could check its work afterwards.

It didn't start this narrow. The first version applied everything Jev said.

How I tested it

I ran every task twice on my Spotter iOS repo, once the default way (Opus, general-purpose agent, Claude's own settings) and once with the mod. Then a blind judge (Opus, with the repo open) got both answers labeled X and Y in random order. It checked every file path, line number and config value against the code and picked a winner or called a tie.

The tasks covered the range of stuff I actually send to subagents. A file count, a few "where is X configured" lookups, explaining the fastlane lanes, a two-device sync risk assessment, a security review, an offline-first sync architecture design, and inventing new coaching features. Round 2 added four more lookups. Round 3 used 12 brand new tasks the mod had never seen.

Round 1: let Jev move anything

In the first round I let Jev move calls up or down, applying agent, model and effort whenever it was at least 60% sure. Here's what it picked:

TaskAgent pickModel pickEffort
Count Swift files
trivial
Explore0.94
Haiku1.00
low
Find CloudKit container config
lookup
Explore0.98
Haiku0.83
medium
List SwiftData models
lookup
Explore0.45
Sonnet0.56
high
Explain fastlane lanes
explain
general0.50
Sonnet0.46
medium
Assess two-device sync risk
analysis
general0.58
Sonnet0.83
high
Security review of the app
review
general0.67
Opus0.60
high
Design offline-first sync
design
Plan0.99
Fable0.73
xhigh
Invent new coaching features
ideation
claude0.41
Fable0.34
high
Round 1: Jev's raw picks on the first eight tasks, with its confidence in each. Orange effort means it asked for more than the default. Its model calls look sensible. Its effort calls are what lost the round.

Looking at the model column, most of those are reasonable. Haiku for the file count, Opus for the security review, Fable for designing a new sync architecture. The effort column is where it went wrong. It asked for high or xhigh effort on five of the eight tasks, and more effort just meant longer answers. Several blew past the word limit in the prompt and the judge marked them down for it.

It lost 6 of 8 to the default, won 1, tied 1. The architecture task went to Plan on Fable at xhigh, cost 2.5 times as much, and still lost.

Round 2: only move down

So I took away the ability to raise anything. Jev could only move a call to a cheaper model, and effort stayed at the default. It moved 7 of 12 calls. Six went to Haiku, one went to Sonnet.

The cost on those 7 calls dropped from $0.60 to $0.31. Across all 12 tasks it was $1.27 down to $0.94.

The judge still preferred the default on 4 of the 7 and tied the other 3. The one that actually worried me was the Sonnet call. It was the two-device sync risk task, which is code analysis and not a lookup, and Sonnet got a real fact wrong. It said workout entries get deleted when they actually just get detached from the workout (the relationship uses .nullify). Opus had it right. If I'd been acting on that answer I'd have been chasing a bug that doesn't exist.

The six Haiku answers were a different story. Every fact checked out.

What it saved on lookups

The 7 lookup calls it sent to Haiku across rounds 2 and 3 went from $0.52 to $0.25. That's 51% cheaper, somewhere between 36% and 68% depending on the call.

−51%
$0.52 → $0.25 across the 7 lookups sent to Haiku
Opus (default)Explore on Haiku
HealthKit authorizationround 3 · $0.079 → $0.025
68%
Count Swift files$0.029 → $0.013
57%
CloudKit container config$0.082 → $0.035
57%
Keychain accessibility value$0.050 → $0.022
56%
Bundle ID + deployment target$0.052 → $0.023
56%
List the app's view files$0.077 → $0.041
46%
Every /api/chat call site$0.149 → $0.095
36%
Cost per lookup call, default versus routed. Every file path, line number and config value in the Haiku answers checked out. Costs are steady-state per call; six are from round 2, HealthKit is the one round 3 lookup that cleared the 0.8 bar.

Every fact in those answers checked out, line numbers and config values included. Finding where HealthKit permission gets requested is a grep and a careful read. I was paying Opus prices for that.

Haiku's answers were leaner though. On 3 of the 6 round 2 pairs the judge liked the Opus answer better because it added extra context, like where a setting can be overridden. The other 3 were ties. I can live with that, and if I want the extra context I can ask for it.

Round 3: lookups only, on tasks it had never seen

So now it only moves lookups, only down, and only when Jev is at least 80% sure about both the task type and the model. I set that rule before running it on 12 fresh tasks: 8 lookups and 4 analysis tasks like reviewing a function for multi-device bugs and writing a test plan for the server's quota system.

Round 1 · 8 tasks

Up or down

Apply agent, model and effort whenever Jev is 60%+ sure, in either direction.

611
Total cost (est.)$2.47 → $2.33

More effort just made answers longer and blew the word limits. The Fable pick cost 2.5x and still lost.

Round 2 · 12 tasks

Down only

Jev can only move a call to a cheaper model. 7 of 12 calls got moved.

43
Moved calls$0.60 → $0.31

Sonnet said workout entries get deleted. They get detached (.nullify). Opus had it right.

Round 3 · 12 fresh tasks

Lookups only

Explore 0.8+ and Haiku 0.8+, rule fixed before the run. 8 lookups, 4 analysis.

1 routed, correct
Routed call$0.079 → $0.025

Left all 4 analysis tasks alone. Only routed 1 of the 8 lookups, which is the next problem.

Blind judge preferred default (Opus)Preferred Jev's routingTie
Three rounds on the Spotter iOS repo, each judged blind by Opus with the repo open. Round 2 and 3 verdicts count only the calls Jev actually changed.

It left all four analysis tasks alone, which is what I wanted after the Sonnet mistake. The one lookup it routed (finding the HealthKit authorization request) came back right at a third of the cost.

The problem is it only routed 1 of the 8 fresh lookups. Total cost for round 3 was basically a wash, $1.17 default against $1.18 with the mod, because nearly everything ran on Opus anyway.

Jev was sure these were lookups. It was a lot less sure Haiku was enough.

Round 3, the 8 fresh lookup tasks. Every one was picked as Explore and Haiku; the question was how confident.

routes today (0.8 / 0.8) would route at Haiku ≥ 0.6 0.500.600.700.800.901.00 0.500.600.700.800.901.00 Confidence it's an Explore (read-only lookup) task → Confidence Haiku is enough → L1 L2 L3 L4 L5 L6 L7 L8
L1 HealthKit authorization0.96 / 0.80 · routed
L2 Free-tier turn limit0.95 / 0.52
L3 Default backend URL0.96 / 0.55
L4 Server model per tier0.90 / 0.48
L5 List the ADRs0.75 / 0.99
L6 Gap reminder notifications0.98 / 0.64
L7 Coach tool names0.95 / 0.63
L8 Coach memory fields0.88 / 0.73
Explore confidence / Haiku confidence for each lookup. L5 is the odd one: Jev was 99% sure Haiku could list a folder of ADRs, but only 75% sure it was an Explore task, so it stayed on Opus.

Jev was sure about the task type on almost all of them. It said Explore with 0.88 to 0.98 confidence on seven of the eight. It was a lot less sure Haiku was enough, mostly between 0.48 and 0.73. The weird one is L5, listing the architecture decision records. Jev was 99% sure Haiku could do it and only 75% sure it was an Explore task, so it fell through.

If I kept the Explore bar at 0.8 and dropped the Haiku bar to 0.6, three more lookups would have routed (gap reminder notifications, coach tool names, coach memory fields), so 4 of 8 instead of 1. I'm not changing the threshold based on the same tasks I'd be grading it on though. That's how you end up fooling yourself. I'm going to try the lower bar on another fresh set before I trust it.

Why Jev for this

The routing call has to cost less than it saves, and it can't slow me down. Jev reads about 760 tokens per call and answered in 195 to 350 ms in my runs. It hands back typed answers with probabilities, so "only route above 80%" is a couple of lines of code instead of me asking a model for a paragraph and parsing it.

The probabilities are also what made the round 3 chart possible. I can see exactly how close each call came to routing, which tells me what to test next instead of guessing.

Caveats

One codebase, one run per task, and the judge is a model. I also picked which tasks counted as lookups. How much it saves you depends on how many of your agent calls are lookups. It doesn't save money overall in round 3 and I'm not going to pretend it does. What it does right now is make the cheap calls cheap without touching the ones that need Opus.

The security review in the benchmark also found a production secret I'd committed back in July. Both runs caught it.

Join the conversation

Have thoughts about this post? Reply on 𝕏 — I read every one.

Discuss on 𝕏
Loading tweet...

I wrote this post inside BlackOps, my content operating system for thinking, drafting, and refining ideas — with AI assistance.

If you want the behind-the-scenes updates and weekly insights, subscribe to the newsletter.

Related Posts