Ben Newton - Commerce Frontend Specialist

Jev now makes the decisions in BlackOps that I used to hand to an LLM

A decision model that answers with numbers instead of prose. It scores my reply targets, checks voice on every publish, filters my feeds and reads replies to brain tasks.

Jev now makes the decisions in BlackOps that I used to hand to an LLM
•18 min read

Every product with AI in it has spots where something needs deciding. Is this post worth replying to. Does this draft sound like me. Is this article worth reading, or is it another funding announcement. For most of BlackOps' life I handed those questions to a chat model, asked it nicely to answer in JSON, and parsed whatever came back.

Over the last couple of weeks I've been pulling those calls out one at a time and giving them to Jev. It now runs in four places inside BlackOps. Two of them were LLM calls I ripped out. The other two are features I built on Jev from the start, because once it was in the codebase there was no reason to reach for a chat model to make a yes/no call. And through a Sortie I can carry Jev into anything else I build, which is how three of my own tools dropped their chat model steps too.

What Jev is

Jev is TypeSafe's System One model, and it doesn't generate text. You send it some state (a post, an email, a draft) and a set of questions, and it sends back typed answers with calibrated probabilities, all answered in parallel in one pass.

Three kinds of question:

  • a yes/no condition (TypeSafe calls it a noul), and how likely it is to hold
  • a pick from a list, with the full distribution behind the pick
  • a score on a ladder of ordered levels

The possible answers are fixed before the call. Jev can't hand back something outside the set I defined, so there's no parsing step, no retry when the JSON comes back malformed, and no page of prompt begging a general model to act like a classifier. I get a number and I draw a line under it.

The same input and the same question give the same answer every time. A chat model doesn't do that, and a threshold on a number that moves between runs isn't much of a threshold.

Interactive explainer

One decision, two paths

A feature needs a verdict. Before, a chat model wrote an answer and my code parsed it. Now Jev returns typed answers and my code applies the thresholds.

How a Jev call works
  1. 1InputThe state (a post, a draft, an email) and the questions. I set the possible answers first.
  2. 2JevAnswers all the questions in parallel, in one pass. It writes no text.
  3. 3VerdictTyped answers with probabilities. Same input, same answer.
  4. 4Feature actsCode compares the numbers to thresholds. Then it queues, blocks, archives or closes.
Pick an example
InputA post on X asks how people ship an MCP server.

Before: LLM call

gpt-5-mini, asked for JSON

  1. Send the post with about 2,000 tokens of calibration prose.
  2. The model writes text.
  3. Code pulls the JSON out with a regex.
  4. Scans run in batches of ten. The same post can land on either side of the threshold on two runs.
Result: a score, maybe. No clear reason for a skip.

Now: Jev

one call, six questions

  1. Prefilter in code: too old? Already replied in that thread? If yes, stop. No call.
  2. Ask in_expertise, has_standing, post_type, avoid, reply_priority, hunt.
  3. Verdict: avoidlow post_typequestion reply_priorityhigh
  4. Code: no veto, type is allowed, so queue by priority.
Feature acts: the post goes in my reply queue. The verdict is saved.
InputA politically charged pile-on that matches one of my hunts.

Before: LLM call

gpt-5-mini, asked for JSON

  1. Send the post with the calibration prose.
  2. The model writes text. Code parses a score out of it.
  3. One number holds everything. A topic match can push it over the line.
Result: no separate veto, no saved reason.

Now: Jev

same call, same six questions

  1. The post passes the prefilter.
  2. Verdict: huntmatch avoidhigh
  3. Code: avoid is a veto. Nothing else can override it.
Feature acts: the post is dropped. The saved verdict shows the real reason.
InputA tweet I schedule through BlackOps. One style rule says: no em dashes.

Before: LLM critic

gpt-5-mini, asked for a JSON verdict

  1. Send the draft and the voice rules.
  2. The model writes a verdict. Code parses it.
  3. If the call errored, it recorded a flat 50 with no violations.
  4. The gate blocked the post with nothing to fix. One tweet got "50/70" three times.
Result: a refusal with no reason.

Now: Jev

one call per piece of content

  1. Ask one voice score on a five-level ladder (15, 40, 60, 80 or 95).
  2. Ask one probability per style rule: does the text clearly break it?
  3. Code: a rule at 0.7 or above is a violation. A major break caps the score at 49. The pass mark is 70.
  4. If Jev times out, the post goes through marked degraded.
Feature acts: pass, or refuse and name the broken rules, worst first.
InputAn RSS item about a startup funding round. It is on topic.

Before: topic match

no model, no judgment step

  1. The item matches a topic I follow.
  2. It goes straight to the inbox.
  3. About 70% of captured items got deleted on sight.
Result: on topic, and useless.

Now: Jev

built on Jev from the start

  1. Score the item on arrival, before any inbox.
  2. Ask item_type, usable_angle, novel.
  3. Verdict: item_typefunding news
  4. Code: that type is on the kill list.
Feature acts: archived with its scores, not deleted. It purges after 30 days.
InputAn email reply to a brain task reminder: "Here is the photo."

Before: nothing

this feature is new

  1. It never had an LLM in it.
  2. Jev was already my default for this kind of question.
Result: no old path to compare.

Now: Jev

reads the reply before anything closes

  1. Ask: does the reply resolve the ask? Is it a question back? What does it try to do?
  2. Verdict: resolves the askclear
  3. Code: only a clear resolution closes the task. Ambiguous, or a Jev error, leaves it open.
Feature acts: the task closes. I can reopen it if it closed wrong.
InputA bug typed into chat: "the tag thing is broken again."

Before: hardcoded

my bug logger

  1. Severity was a hardcoded default.
  2. An agent could go into the repo to chase a vague report.
Result: every bug looked the same.

Now: the jev Sortie

any tool fires it by name; the key stays on the server

  1. Ask: how severe? What kind of record? A duplicate of an open one? Enough detail to act on?
  2. Verdict: enough detailunder 0.5
  3. Code: under 0.5 means no agent goes out.
Feature acts: filed for me to look at.
Side by side
LLM callJev
OutputText that code must parseTyped answers with probabilities
Possible answersAnything the model writesOnly the set I defined
Same input twiceCan changeSame answer
Parse and retryYes, when the JSON breaksNone
PricePays for generated text$0.042 per million input tokens, output free
Reply targetingBatches of ten$0.000074 per post, six answers in 350ms in one live scan

Answers show as high or low for illustration. Thresholds are the real ones. Writing (drafts, replies, rewrites) stays with the LLM.

What it costs

TypeSafe charges $0.042 per million input tokens, and output is free. TypeSafe's own claim is 40 to 400 times cheaper and 40 to 200 times faster than frontier models on comparable tasks. Those are their numbers, and they say themselves it's probably the high end, so here are mine.

On reply targeting, after I trimmed the payload, scoring one post costs $0.000074. That works out to about $2.22 a month for someone scoring a thousand posts a day. In one live scan Jev came back in 350ms with all six answers filled in. At that price I stopped rationing decisions. I can put one on every post, every feed item and every publish.

None of this means the LLM is gone from BlackOps. Drafting a reply, revising a post, writing anything at all, that's still a writing model. Jev took over the parts where I was paying a writing model to hand me a verdict.

Reply targeting in the Chrome extension

As I work through X search and explore, the BlackOps Chrome extension flags posts worth replying to, based on my hunts (saved descriptions of the conversations I want to be in). It used to ask gpt-5-mini for prose and regex the JSON back out. So I was paying for text generation to get a score, scans got rationed into batches of ten, and the same post could land on either side of the threshold on two runs. The old prompt carried about 2,000 tokens of calibration prose just to make a general model behave like a classifier.

Now each post goes through a deterministic prefilter first (too old, already replied in that thread), and the survivors get one Jev call with six questions:

  • in_expertise: does this touch something I have hands-on experience with
  • has_standing: do I have a first-hand decision, mistake or number to add, beyond agreeing
  • post_type: question, hot take, announcement, promo, engagement bait or personal news
  • avoid: is this a pile-on, politically charged, or otherwise a bad place to show up
  • reply_priority: how worth replying it is, on a four-level ladder
  • hunt: which of my hunts it belongs to, or none

The routing is arithmetic in code, never another model's opinion. avoid is a veto no matter what else scores. Promo, engagement bait and personal news get dropped. Everything else queues by priority. Every verdict is saved, so I can tune the thresholds against what I actually replied to, and a skipped post tells me the real reason it was skipped.

Two fixes came out of the first days. Jev was judging my expertise without being told anything about me, so an MCP announcement scored 0.38 on in_expertise for someone who has shipped an MCP server and written three posts about it. The payload now carries what I know, pulled from the names of the brains (sets of markdown notes BlackOps keeps for me) attached to each hunt. Then I measured the payload itself. It was 3,248 tokens per post and the post was 61 of them, because I was sending each hunt brief twice. Cutting the duplicate and capping the reply intent brought it to 1,759 tokens, 46% less.

Brand voice check on every publish

Every time I publish or schedule anything through BlackOps (tweets, threads, blog posts, X articles, LinkedIn, Threads, TikTok) it gets checked against the brand voice I set up for that site. That check used to be a gpt-5-mini critic. When the critic's call errored, it recorded a flat 50 with no violations and blocked the post with nothing to fix. One of my tweets got refused three times in a row with the same bare "50/70".

Jev runs the check now, one call per piece of content:

  • one score for overall voice on a five-level ladder, from off-voice to on-voice throughout, mapped to 15, 40, 60, 80 or 95
  • one probability per style rule, asking whether the text clearly breaks it

A rule at 0.7 or above counts as a violation. At 0.9 it's a major one and caps the score at 49. The pass mark is 70. When something fails, the refusal names the rules that broke, worst first, so I know what to fix. The verdict also comes back in the publish response with the engine, model, score and threshold, so a real pass never looks the same as a check that got skipped.

It judges each piece as the platform it's going to, so a short reply on X isn't marked down for being short. If Jev times out, the post goes through marked degraded instead of blocking me. If it fails fast, the old LLM critic steps in.

The blog auto-revision loop still uses the LLM critic, because that one feeds written suggestions into a rewrite. That's writing, and writing is still the LLM's job.

Content scoring on every feed item

BlackOps content monitoring pulls in RSS items and discovered content for the topics I care about. Roughly 70% of what it captured got deleted on sight. Most of it was on topic, and useless anyway: funding rounds, launch announcements, press releases with nothing to say. Asking whether something was on topic was the wrong filter.

Now Jev scores every item as it arrives, before it can reach an inbox, with three questions in one call:

  • item_type: funding news, product launch, benchmark, opinion, tutorial, research, press release or incident
  • usable_angle: could someone working in my domains say something about this beyond summarizing it
  • novel: is this new, or another outlet's version of a story already captured this week

Types on a kill list are suppressed. The rest rank by usable angle, with novel catching aggregator duplicates. Nothing gets deleted at ingest. A suppressed item is archived with its scores attached so a bad threshold stays visible, and archived items purge after 30 days. What survives becomes a ranked, capped digest, and if nothing clears the bar the digest says so instead of padding itself out.

The domains usable_angle judges against are set per site at /admin/content-scoring. A site without them doesn't ingest at all, because an empty domain list doesn't score neutral. It scores everything down and looks exactly like a quiet feed.

Reading replies to brain tasks

A brain can now keep track of things it's waiting on, like a photo or an answer only one person has, and email that person a reminder. When they reply, Jev reads the reply before anything closes. It asks whether the reply actually resolves the ask, whether it's a question back, and what the reply is trying to do. Only a clear resolution closes the task. Anything ambiguous, or any Jev error, leaves it open, and I can reopen a task that closed wrong.

This one never had an LLM in it. By the time I built it, Jev was already my default for this kind of question.

A portable Jev I can throw anything at

The four features above are Jev wired into BlackOps itself, and the Sortie is how I carry it into everything else.

BlackOps has Sorties, a saved endpoint that keeps the URL, the headers and an encrypted credential on the server. I put the TypeSafe key in one called jev. Anything that can fire a Sortie by name can use it: a script, a scheduled job, an agent in a chat, Claude or Cursor over MCP. It sends a state and some questions, and since firing a Sortie hands back the endpoint's full response, Jev's typed answers come straight back to whatever asked. None of those callers ever holds the key. It's decrypted on the server at fire time and never ends up in a conversation, a skill file or a repo. Every fire is rate limited and shows up in the Sortie's fire history.

So when something new needs a yes/no, a pick from a list or a score, there's no integration to build. I name the Sortie and write the questions. Three of my own tools moved over that way in one day.

My bug logger turns a description typed into chat into a structured record, then sends an unattended agent into the repo to write the fix. Severity used to be a hardcoded default. Now one Jev call asks how severe it is, what kind of record it really is (which catches the feature request I filed as a bug), whether it duplicates something already open, and whether there's enough detail to act on. Under 0.5 on that last one, it gets filed for me to look at and no agent gets sent off chasing "the tag thing is broken again."

My writing check reads drafts against my voice rules. Jev decides whether a pattern is there, and the writing model finds it and rewrites it. It was good on vocabulary right away. On rhythm, like the same list construction used four times in one draft, it wasn't, so those checks stayed with me.

My bill tracker reads my email for hosting charges, API bills and subscriptions and files each one into a spreadsheet. Working out what counts and what kind of expense it is comes down to a pick from a list, which is what Jev does best. The job already knew how to fire a Sortie, so wiring it up took about as long as making coffee.

Where it stands

It's live. Every BlackOps feature Jev runs in saves each verdict, so I can see what it decided on any post, draft, feed item or reply.

The thresholds are still mostly my guesses: 0.7 for a voice violation, 0.5 for an actionable bug, the floors on the feeds. There was no history to cut them against, which is why every one of these features records its verdicts from day one. They get tuned against what actually happened.

Every place Jev runs also has a way out when it's down. Reply targeting falls back to the old LLM post by post, the voice check lets content through marked degraded, feeds let items through flagged unscored, and brain tasks stay open. I found out why that matters right after launch. Every Jev call from the extension was failing on a missing field and quietly falling back to the LLM, and nothing looked wrong because the fallback worked. It logs an error now when a whole scan fails.

Almost every feature I've built has a step that quietly needs judgment, and I'd handed most of them to a writing model. Jev gives me the same answer every time, and a number I can set a line against instead of a paragraph I have to interpret. I keep finding more places to put it, and with the Sortie, any new place just has to fire it by name.

If you want this running on your own publishing, brains and feeds, BlackOps is at blackopscenter.com/start.

Where in what you've built is there a step that needs judgment, that you either hardcoded years ago or handed to a chat model and stopped looking at? I'd like to hear what you find.

šŸ’Œ Want more insights like this?

Subscribe to my newsletter for weekly deep dives into frontend development, AI, and productivity

I wrote this post inside BlackOps, my content operating system for thinking, drafting, and refining ideas — with AI assistance.

If you want the behind-the-scenes updates and weekly insights, subscribe to the newsletter.

Related Posts