My Claude Code mod lets Jev choose the agent and model for every call
Jev picks the model for every subagent call in about a third of a second. On lookups it cut the cost by 51% and every fact still checked out.

I was watching Claude Code work through my iOS repo and it spun up a subagent to count the Swift files. On Opus. Then it sent another one off to find a single config line, also on Opus. Subagents just inherit whatever model the main session is on unless Claude says otherwise, so little errands like that end up on the expensive model.
So I built a mod for it. Before every agent call it asks Jev, TypeSafe's decision model, which agent type fits the task, which model is the cheapest one that'll still get it right, and how much effort it needs. All three come back in one request in about a third of a second, each with a confidence score. If Jev is sure it's a read-only lookup Haiku can handle, the mod switches the call to the Explore agent on Haiku. If it isn't sure, the call runs the way Claude set it up. There's a little band above the prompt that shows what it picked, mostly because I wanted to see it working.
How the mod works
It's a Claude Code mod (a plugin of function hooks that runs inside the session). The hook that matters sits on every Agent tool call. Before the subagent starts, the mod sends Jev the task description, up to 4,000 characters of the prompt, the agent types Claude Code has on offer with their descriptions, and a one-line description of each model. Haiku is "simple lookups, file counts, grep-style searches." Fable is "open-ended work with no clear path." That kind of thing.
Jev sends back three typed answers. A choice of agent type with a probability for every option, a choice of model the same way, and an effort score. No paragraph to parse. That's the whole reason I used it for this.
1 · Claude Code
Claude calls the Agent tool
With a task description, a prompt, and usually no model, so the subagent inherits the main session's model (Opus for me).
2 · The mod's hook
Intercepts the call before it runs
Sends the task, the agent types on offer, and the four models to Jev. Up to 4,000 characters of the prompt.
3 · Jev, one request
Three typed answers
195 to 350 ms, about 760 input tokens.
agent == Explore && agent.conf ≥ 0.8 && model == haiku && model.conf ≥ 0.8Passes
Rewritten to Explore on Haiku
Read-only agent, cheapest model. The only change the mod is allowed to make.
Anything else
Runs exactly as Claude set it up
Jev's pick still shows in the band above the prompt, it just isn't applied.
The gate in the hook is four conditions and it's the only thing that can change a call:
const isLookup = mode === 'apply' && answers.agent.choice === 'Explore' && answers.agent.confidence >= 0.8 && answers.model.choice === 'haiku' && answers.model.confidence >= 0.8 && (requestedTier < 0 || requestedTier > tierOf('haiku')) if (!isLookup) return next(e) return next({ ...e, subagent_type: 'Explore', model: 'haiku' })
If Jev errors or times out, the call just runs the way Claude set it up. There's also /jev advise, which shows the pick in the band without applying anything, and /jev off. Every pick and every run gets logged to a JSON file so I could check its work afterwards.
It didn't start this narrow. The first version applied everything Jev said.
How I tested it
I ran every task twice on my Spotter iOS repo, once the default way (Opus, general-purpose agent, Claude's own settings) and once with the mod. Then a blind judge (Opus, with the repo open) got both answers labeled X and Y in random order. It checked every file path, line number and config value against the code and picked a winner or called a tie.
The tasks covered the range of stuff I actually send to subagents. A file count, a few "where is X configured" lookups, explaining the fastlane lanes, a two-device sync risk assessment, a security review, an offline-first sync architecture design, and inventing new coaching features. Round 2 added four more lookups. Round 3 used 12 brand new tasks the mod had never seen.
Round 1: let Jev move anything
In the first round I let Jev move calls up or down, applying agent, model and effort whenever it was at least 60% sure. Here's what it picked:
| Task | Agent pick | Model pick | Effort |
|---|---|---|---|
| Count Swift files trivial | Explore0.94 | Haiku1.00 | low |
| Find CloudKit container config lookup | Explore0.98 | Haiku0.83 | medium |
| List SwiftData models lookup | Explore0.45 | Sonnet0.56 | high |
| Explain fastlane lanes explain | general0.50 | Sonnet0.46 | medium |
| Assess two-device sync risk analysis | general0.58 | Sonnet0.83 | high |
| Security review of the app review | general0.67 | Opus0.60 | high |
| Design offline-first sync design | Plan0.99 | Fable0.73 | xhigh |
| Invent new coaching features ideation | claude0.41 | Fable0.34 | high |
Looking at the model column, most of those are reasonable. Haiku for the file count, Opus for the security review, Fable for designing a new sync architecture. The effort column is where it went wrong. It asked for high or xhigh effort on five of the eight tasks, and more effort just meant longer answers. Several blew past the word limit in the prompt and the judge marked them down for it.
It lost 6 of 8 to the default, won 1, tied 1. The architecture task went to Plan on Fable at xhigh, cost 2.5 times as much, and still lost.
Round 2: only move down
So I took away the ability to raise anything. Jev could only move a call to a cheaper model, and effort stayed at the default. It moved 7 of 12 calls. Six went to Haiku, one went to Sonnet.
The cost on those 7 calls dropped from $0.60 to $0.31. Across all 12 tasks it was $1.27 down to $0.94.
The judge still preferred the default on 4 of the 7 and tied the other 3. The one that actually worried me was the Sonnet call. It was the two-device sync risk task, which is code analysis and not a lookup, and Sonnet got a real fact wrong. It said workout entries get deleted when they actually just get detached from the workout (the relationship uses .nullify). Opus had it right. If I'd been acting on that answer I'd have been chasing a bug that doesn't exist.
The six Haiku answers were a different story. Every fact checked out.
What it saved on lookups
The 7 lookup calls it sent to Haiku across rounds 2 and 3 went from $0.52 to $0.25. That's 51% cheaper, somewhere between 36% and 68% depending on the call.
Every fact in those answers checked out, line numbers and config values included. Finding where HealthKit permission gets requested is a grep and a careful read. I was paying Opus prices for that.
Haiku's answers were leaner though. On 3 of the 6 round 2 pairs the judge liked the Opus answer better because it added extra context, like where a setting can be overridden. The other 3 were ties. I can live with that, and if I want the extra context I can ask for it.
Round 3: lookups only, on tasks it had never seen
So now it only moves lookups, only down, and only when Jev is at least 80% sure about both the task type and the model. I set that rule before running it on 12 fresh tasks: 8 lookups and 4 analysis tasks like reviewing a function for multi-device bugs and writing a test plan for the server's quota system.
Round 1 · 8 tasks
Up or down
Apply agent, model and effort whenever Jev is 60%+ sure, in either direction.
More effort just made answers longer and blew the word limits. The Fable pick cost 2.5x and still lost.
Round 2 · 12 tasks
Down only
Jev can only move a call to a cheaper model. 7 of 12 calls got moved.
Sonnet said workout entries get deleted. They get detached (.nullify). Opus had it right.
Round 3 · 12 fresh tasks
Lookups only
Explore 0.8+ and Haiku 0.8+, rule fixed before the run. 8 lookups, 4 analysis.
Left all 4 analysis tasks alone. Only routed 1 of the 8 lookups, which is the next problem.
It left all four analysis tasks alone, which is what I wanted after the Sonnet mistake. The one lookup it routed (finding the HealthKit authorization request) came back right at a third of the cost.
The problem is it only routed 1 of the 8 fresh lookups. Total cost for round 3 was basically a wash, $1.17 default against $1.18 with the mod, because nearly everything ran on Opus anyway.
Jev was sure these were lookups. It was a lot less sure Haiku was enough.
Round 3, the 8 fresh lookup tasks. Every one was picked as Explore and Haiku; the question was how confident.
Jev was sure about the task type on almost all of them. It said Explore with 0.88 to 0.98 confidence on seven of the eight. It was a lot less sure Haiku was enough, mostly between 0.48 and 0.73. The weird one is L5, listing the architecture decision records. Jev was 99% sure Haiku could do it and only 75% sure it was an Explore task, so it fell through.
If I kept the Explore bar at 0.8 and dropped the Haiku bar to 0.6, three more lookups would have routed (gap reminder notifications, coach tool names, coach memory fields), so 4 of 8 instead of 1. I'm not changing the threshold based on the same tasks I'd be grading it on though. That's how you end up fooling yourself. I'm going to try the lower bar on another fresh set before I trust it.
Why Jev for this
The routing call has to cost less than it saves, and it can't slow me down. Jev reads about 760 tokens per call and answered in 195 to 350 ms in my runs. It hands back typed answers with probabilities, so "only route above 80%" is a couple of lines of code instead of me asking a model for a paragraph and parsing it.
The probabilities are also what made the round 3 chart possible. I can see exactly how close each call came to routing, which tells me what to test next instead of guessing.
Caveats
One codebase, one run per task, and the judge is a model. I also picked which tasks counted as lookups. How much it saves you depends on how many of your agent calls are lookups. It doesn't save money overall in round 3 and I'm not going to pretend it does. What it does right now is make the cheap calls cheap without touching the ones that need Opus.
The security review in the benchmark also found a production secret I'd committed back in July. Both runs caught it.
I wrote this post inside BlackOps, my content operating system for thinking, drafting, and refining ideas — with AI assistance.
If you want the behind-the-scenes updates and weekly insights, subscribe to the newsletter.
Related Posts

Sorties: Fire Real Actions Straight From Your AI Chat
A Sortie turns a sentence in your AI chat into a real action. Fix a bug, message your team, fire an automation. Set it up once, fire it by name from anywhere, and your keys never touch the conversation.

Claude Code Is Not Your Project Manager. The Artifacts Are.
A year ago I gave Claude a ROADMAP.md and called it my project manager. The post still ranks. The post is also out of date. Here is what Claude actually manages now.

Personal Second Brains Are Step One. Operational Second Brains Are The Actual Destination.
Every few months a new "second brain" build hits X. The pattern is real, the folder taxonomies are getting sharper. But every one of them stops in the same place: a markdown file. Here's the wall.