A document lands in my production AI workflow, and four small decisions need an answer before a person ever looks at it: which model should handle this task, which reference pages are worth reading first, whether this file is actually a source or something that slipped in by mistake, and where it should be filed. None of those four needs a frontier model's full reasoning. All four still need to be right more often than a fixed rule gets them right.
Jev is a small decision model from TypeSafe, and I call it from code for exactly those narrow jobs. I send it a short piece of state and a few questions; for each one it returns a choice with a probability for every option, or a yes/no probability, in about 0.3 seconds. Code decides what happens with that answer. A person decides everything that matters.
What does Jev actually do, in plain terms?
Jev does not chat, write or search. It answers one bounded question at a time about something I already built: is this task mechanical or hard, is this reference page useful, does this file look like a draft conclusion instead of a source, where does a file like this usually get filed. My rule: code runs the flow, Jev answers the small question inside it, and the probability decides which branch runs next. It never authorises a step, and it is cheap enough to call on every task, though TypeSafe doesn't publish a per-call price.
The four places I actually use it
| What it decides | If Jev is unsure | Measured result |
|---|---|---|
| Which model handles a task | Confidence under 0.8 moves the task one tier up, never down | 38/46 correct with the first version of the questions, 43-44/46 after restructuring them, over two runs, 0 tasks routed too low |
| Which reference pages a reviewer reads first | Reranking fails, the plain keyword order runs instead | "Directly relevant" picks went from 33/45 to 43/45 for one blind judge and 35/45 to 44/45 for another; off-topic picks dropped from 3 to 0 |
| Whether a file is a misplaced draft, not a source | Only skipped if two separate answers agree, with high confidence; otherwise it stays in | Caught 10-11 of 12 test files over two runs; skipped 0 of 62 genuine input files |
| Where an incoming file belongs | It only suggests; nothing moves automatically | 48 of 52 suggestions matched how a person had already filed the same files |
Model choice: cheaper by default, stronger on purpose
Giving each tier a short definition, what it is not for, and examples is what moved the score in the table above. A fallback classifier takes over automatically if Jev is unavailable.
Reference pages: re-rank, don't replace
A keyword search still builds the shortlist; Jev only reorders it. Calls run in parallel, about 1.2 seconds a batch, so none of this adds waiting time, and a failed call just falls back to the plain order.
Catching a document that shouldn't be there
Sometimes a draft conclusion gets saved under an ordinary file name and ends up sitting among the source documents, which means the review would end up citing its own conclusions. Checking the file name alone caught none of a dozen test cases like that. Having Jev read the content and answer two separate questions, agreeing with high confidence before anything is skipped, is what got the result in the table. It's worth being honest about the limit: it still missed one or two, so I treat it as a safety net on top of the existing process, not a guarantee.
Filing: a suggestion, never a move
It only suggests a destination; a person decides and moves the file, every time.
Where I decided not to use it
I tried Jev on the job that matters most here: judging whether the evidence in the documents actually supports a given statement. In a pilot it ranked the most valuable findings, the things missing from the documents entirely, lowest. That's the opposite of useful for that job, so I kept my existing method there instead.
What changed my results
A short definition, what it is not for, and examples for each option moved model routing from 34/46 to 43/46 before I touched anything else. Never pulling examples from the test set, or the score is fake. A second question can rescue the first: the two-answer rule above removed every false alarm I was seeing with one question alone. And the provider's probabilities are rounded to two decimals, so a set of options doesn't always sum exactly to 1; the workflow accounts for that instead of treating it as an error.
How I test a change
Real cases from the workflow, labels checked by two models from different vendors, old behaviour against new on the same cases, two runs each. A change ships only if it measurably improves and nothing else gets worse. My samples are tens of cases at a time, so I treat the results as evidence for this workflow, not a general claim.
Frequently asked questions
What is Jev?
Jev is a small decision model from TypeSafe, called from code over HTTPS. You send it a short piece of state and a set of narrow questions, and for each one it returns a choice with a probability for every option, or a yes/no probability. It is not a chat model, and it does not write or search anything.
Does Jev ever decide anything on its own?
No. Jev answers a question with a probability; code decides what to do with that answer, and a person reviews the result. On my production workflow it never authorises a step, never writes a document and never searches for anything. It only answers the question it was asked.
What happens if Jev is unavailable?
Every use falls back to the earlier behaviour: a fixed classifier instead of Jev's routing, the plain keyword order instead of its reranking, no bad-input check instead of its file screen. The workflow keeps running. The risk I manage instead is a quiet one: a missing or revoked key can look healthy while quality drops, so the health check verifies the key and the run report says when a fallback was used.
Why not use Jev for every judgment in the workflow?
Because I measured it on the one job that matters most, judging whether the evidence actually supports a statement, and it was worse than my existing method there: in a pilot it ranked the most valuable findings lowest. I dropped it for that job. It stays on the four narrow jobs where it measurably helps, not everywhere a decision happens.
What does Jev see from the documents being reviewed?
What Jev may see is a deliberate decision I make per workflow, because it is an external service. I am not stating a general data policy here; it depends on the workflow and what is actually being sent, decided case by case, not assumed.
If something like this slows your team down
None of this is really about Jev specifically. It's about noticing which decisions inside a workflow are small, repeated and narrow enough that a classifier can answer them faster and more cheaply than a full chat model, while keeping the decisions that need judgment with a person. If you have a workflow where small decisions like these slow your team down or get decided inconsistently, a free 45-minute call is where that conversation starts, no cost and no commitment to it.
I wrote about using a model like this as a classifier across a whole workspace, not just one workflow, in one agent on my phone, five tools behind it. Where I decided a cheaper model wasn't good enough for a harder job, judging whether evidence actually supports a conclusion, is in how I evaluate a cheaper model for a second opinion. You can read more about how I scope this kind of work on /services.