All posts
private-aiagentsmodel-routingjevcost-controloperations

One Agent on My Phone, Five Tools Behind It

A one-day experiment on my laptop: one assistant on my phone takes the request, Jev sizes it, cheap models do the research, frontier models get the hard work, and every task ends with a written hand-off. What moves into my production workspace first is the discipline, not the router.

Nobody should choose a model to get work done. The expensive model should be a decision, not a habit.

Stéphane Lepain··Updated ·10 min read

What should one assistant change for the person using it?

I want one place to ask for work, shared instructions behind it and a written hand-off when the work is finished. Choosing tools and models should not be the user's job. This was a laptop experiment; I describe what broke and what I would carry into a working system, not a measured saving for a customer's business.

For months I had five AI tools on my laptop and I was the switchboard.

Each one knew a different slice of how I like to work. Each one had its own idea of where my projects live. Every morning I picked a tool, picked a model from a dropdown, explained the context again, and hoped it remembered the rule I had given it the week before. Inside my chat workspace I had already replaced the dropdown with automatic routing (I wrote about it here). That fixed chat. It did not fix me.

So I spent a day fixing it, on my laptop, as a test. The tools involved are my personal coding assistants, not the private AI workspace I build and run for real work. That workspace already routes chat by difficulty; what follows is the piece it did not have yet, tried out where a mistake costs nothing. This is what I built, what it does for me, and what I am taking from it into the production workspace, and what I am not.

What I wanted

Three things, written down before touching anything.

  1. One place to talk. I say "handle the invoicing app" or "what is the state of the website" and the right thing happens. From my phone when I am not at my desk.
  2. Cheap by default, expensive on purpose. Internet research, lookups and summaries should never run on a frontier model. Reviewing a firewall should.
  3. Never repeat myself. Correct the assistant once and every tool knows it. Stop a task on Tuesday and Wednesday's session picks it up where it stopped.

I walk through the laptop experiment in this video, using my AI likeness and voice; loading the YouTube player connects to Google.

The shape of it

   me (phone or terminal)
          │
          ▼
   ┌─────────────────────────────────────────────┐
   │  Foreman: the one assistant I talk to        │
   │  1. ask Jev how hard, what kind, how risky   │
   │  2. read the current state of that area      │
   │  3. hand the job to ONE worker, then wait    │
   │  4. check the result                         │
   │  5. write the hand-off, report back          │
   └───────┬─────────┬──────────┬────────┬───────┘
           │         │          │        │
        Tier 0    Tier 1     Tier 2   Ask me
        cheap     mid       frontier   first
           │         │          │
           └─────────┴──────────┴──▶ shared memory: rules · state · skills · hand-offs
                                     (one folder, every tool reads it)

A shared memory that every tool reads. One folder holds the standing rules (how I want work delivered, what is off limits, which machine is production), one short file per area with the current state of things, and a set of skills: written procedures for recurring work. Every AI tool on the machine reads the same folder. When I open any of them, it already knows.

One assistant in front. The one that already answers me on my phone became the foreman. It does not do the work. It takes the request, reads the state of that area, hands the job to a worker, checks the result, writes a short hand-off note and reports back with one line saying which worker it used.

Jev decides the lane. Jev is the model behind TypeSafe. It is not a chat model and it does not write anything. You send it a piece of state, in my case the request plus the list of available skills, and a small set of typed questions. It answers each one with a structured value and a probability, in one call, for a fraction of a cent. I ask it five things about every request:

QuestionTypeWhat it decides
How hard is this?score 0–3the tier
What kind of work is it?choice: research, implementation, operations, decisioncaps research at the mid tier, forces operations to the frontier tier
Which area of my work?choicewhich state file and skill the worker gets
Does it touch anything live, public, or hard to reverse?yes/no probabilityrisk gate
Is this a decision only I can make?yes/no probabilitythe foreman asks me first

Those numbers pick the lane. A language model would have guessed; Jev classifies. That difference is what makes the routing predictable enough to trust.

Three lanes of workers, by name. The foreman never chooses the expensive lane out of habit. It needs a reason, and the reason is in the hand-off.

LaneModelWhat goes thereWhat it costs me
Tier 0DeepSeek Flash (GLM 5.3 as alternate)lookups, status questions, summaries, internet researchmetered, cents
Tier 1GPT-5.6 Luna, high reasoningbounded implementation: scripts, documents, reviewsflat subscription
Tier 2GPT-5.6 Sol, then GPT-6 Astra on escalationproduction, network, security, architecture, ambiguous debuggingflat subscription
Tier 2, Claude laneClaude, through its own toolwork on my own machine and repositoriesflat subscription
Ask menoneanything Jev flags as my decisionmy time, once

Research never reaches a frontier model any more. That single rule is where most of the waste was.

A written hand-off after every task. What was asked, what was done, what was verified, what is open, and which lane did it. The next session, in any tool, starts by reading it.

What it looks like from the phone

I typed a question about which cars run well on E85 and whether that beats diesel over a few years. Jev sized it as research. The foreman sent it to a cheap worker, waited, checked the answer, saved the note and replied with the comparison and one line: "via Luna". No frontier model was involved. I did not open a laptop.

Later that afternoon I asked it to review the firewall on my home router. Same words, same phone. This time Jev returned a risk of 0.9 and "operations", so the request went to the frontier lane, and the foreman would have stopped to ask me before changing anything.

Same front door. Different cost. That is the whole point.

"which cars run on E85, and vs diesel?"
   Jev: hard 0.98 · research · risk 0.28 · not my decision
   → Tier 1, GPT-5.6 Luna                       cost: flat

"review the firewall rules on my router"
   Jev: hard 2.16 · operations · risk 0.90 · not my decision
   → Tier 2, frontier, would ask before changing anything

"delete the demo server, we don't need it"
   Jev: hard 2.47 · risk 0.90 · my decision 0.91
   → stop, ask me first

What broke, honestly

None of this worked on the first run, and I would rather tell you what went wrong than pretend otherwise.

  • The foreman gave its first cheap worker an enormous brief, a full live audit, and the worker ran out of time. New rule: a cheap brief is "read the state, answer the question", nothing wider. Live checks the foreman does itself, in one command.
  • The foreman started researching the same question in parallel while the worker was busy. Twice the tokens, one answer. New rule: while a worker runs, the foreman waits.
  • Jev is confident about area, risk and whether I need to be asked. It is less confident about difficulty when the request is short. So difficulty now comes from the kind of work as much as the score, and a cheap attempt that fails escalates one lane. Doubt does not.
  • Two of the frontier models hang at their highest reasoning setting when run without a screen. One notch lower, they answer in seconds.

Each of those is one line in a rules file now. That is the real product of the day: the rules, not the wiring.

Where this goes next

The laptop was the rehearsal. The private AI workspace I run in production, and build for customers, is a different animal: a self-hosted chat platform with its own document store, retrieval and agents, a gateway in front of every model that already routes chat by difficulty and keeps the spend logs, and an assistant of its own that drives those agents. It does not need a second router, and it will not get one just because the laptop has one.

What the test taught me is that the routing was not where the value was. The value came from four operating rules: a brief sized to the tier, one worker at a time, a check before anything is reported, and a written hand-off that the next session reads first. Those rules are plain text. They move into the production workspace first, without touching how it routes.

The routing question stays open and gets measured, not decided. For a few weeks the hand-offs on the laptop and the spend logs in the workspace will show whether Jev's sizing beats the difficulty scorer the workspace already has. If it does, the smallest possible change is to let Jev score inside the existing gateway. If it does not, nothing changes. Either way, that decision is made on numbers, and it is made carefully, because that workspace is where the real work happens.

What a customer gets out of this

I build private AI workspaces for companies, and this is the pattern I would now put in front of one.

  • One assistant for the team, reachable from a phone. Nobody chooses a tool. Nobody chooses a model. They say what they need, and the workspace they already have does the work.
  • A bill that follows the difficulty of the work. Routine questions, document lookups and research run on cheap or flat-rate models. The expensive ones are reserved, and every use of them is justified in writing.
  • Institutional memory the company owns. Rules, state and hand-offs are files in the company's own repository, readable by any tool it adopts later. Changing vendors does not mean starting over.
  • A safety gate. Anything that touches a live system, money or something public is flagged before it runs, and a decision that belongs to a human is put to a human, once.
  • A record. Every task leaves a note saying what was done, what was checked and what is open. That is what an auditor, a manager or the next engineer wants to read.

What it is not: certain. Jev returns a judgement with a probability attached, workers make mistakes, and the whole design exists because of that. Verification and escalation are the answer to a probabilistic system, not features added to a reliable one.

If you want the same thing

Start with the rules file, not the tools. Write down how you want work delivered, what is off limits and where things live, in a form any assistant can read. Then put one front door in front of the workspace you already have and let something like Jev size the requests for a month. The hand-off notes will tell you more than any benchmark.

I am happy to talk through what this would look like on your own systems. The conversation is where every engagement starts.