What should one assistant change for the person using it?
I want one place to ask for work, shared instructions behind it and a written hand-off when the work is finished. Choosing tools and models should not be the user's job. This was a laptop experiment; I describe what broke and what I would carry into a working system, not a measured saving for a customer's business.
For months I had five AI tools on my laptop and I was the switchboard.
Each one knew a different slice of how I like to work. Each one had its own idea of where my projects live. Every morning I picked a tool, picked a model from a dropdown, explained the context again, and hoped it remembered the rule I had given it the week before. Inside my chat workspace I had already replaced the dropdown with automatic routing (I wrote about it here). That fixed chat. It did not fix me.
So I spent a day fixing it, on my laptop, as a test. The tools involved are my personal coding assistants, not the private AI workspace I build and run for real work. That workspace already routes chat by difficulty; what follows is the piece it did not have yet, tried out where a mistake costs nothing. This is what I built, what it does for me, and what I am taking from it into the production workspace, and what I am not.
What I wanted
Three things, written down before touching anything.
- One place to talk. I say "handle the invoicing app" or "what is the state of the website" and the right thing happens. From my phone when I am not at my desk.
- Cheap by default, expensive on purpose. Internet research, lookups and summaries should never run on a frontier model. Reviewing a firewall should.
- Never repeat myself. Correct the assistant once and every tool knows it. Stop a task on Tuesday and Wednesday's session picks it up where it stopped.
I walk through the laptop experiment in this video, using my AI likeness and voice; loading the YouTube player connects to Google.
The shape of it
me (phone or terminal)
│
▼
┌─────────────────────────────────────────────┐
│ Foreman: the one assistant I talk to │
│ 1. ask Jev how hard, what kind, how risky │
│ 2. read the current state of that area │
│ 3. hand the job to ONE worker, then wait │
│ 4. check the result │
│ 5. write the hand-off, report back │
└───────┬─────────┬──────────┬────────┬───────┘
│ │ │ │
Tier 0 Tier 1 Tier 2 Ask me
cheap mid frontier first
│ │ │
└─────────┴──────────┴──▶ shared memory: rules · state · skills · hand-offs
(one folder, every tool reads it)
A shared memory that every tool reads. One folder holds the standing rules (how I want work delivered, what is off limits, which machine is production), one short file per area with the current state of things, and a set of skills: written procedures for recurring work. Every AI tool on the machine reads the same folder. When I open any of them, it already knows.
One assistant in front. The one that already answers me on my phone became the foreman. It does not do the work. It takes the request, reads the state of that area, hands the job to a worker, checks the result, writes a short hand-off note and reports back with one line saying which worker it used.
Jev decides the lane. Jev is the model behind TypeSafe. It is not a chat model and it does not write anything. You send it a piece of state, in my case the request plus the list of available skills, and a small set of typed questions. It answers each one with a structured value and a probability, in one call, for a fraction of a cent. I ask it five things about every request:
| Question | Type | What it decides |
|---|---|---|
| How hard is this? | score 0–3 | the tier |
| What kind of work is it? | choice: research, implementation, operations, decision | caps research at the mid tier, forces operations to the frontier tier |
| Which area of my work? | choice | which state file and skill the worker gets |
| Does it touch anything live, public, or hard to reverse? | yes/no probability | risk gate |
| Is this a decision only I can make? | yes/no probability | the foreman asks me first |
Those numbers pick the lane. A language model would have guessed; Jev classifies. That difference is what makes the routing predictable enough to trust.
Three lanes of workers, by name. The foreman never chooses the expensive lane out of habit. It needs a reason, and the reason is in the hand-off.
| Lane | Model | What goes there | What it costs me |
|---|---|---|---|
| Tier 0 | DeepSeek Flash (GLM 5.3 as alternate) | lookups, status questions, summaries, internet research | metered, cents |
| Tier 1 | GPT-5.6 Luna, high reasoning | bounded implementation: scripts, documents, reviews | flat subscription |
| Tier 2 | GPT-5.6 Sol, then GPT-6 Astra on escalation | production, network, security, architecture, ambiguous debugging | flat subscription |
| Tier 2, Claude lane | Claude, through its own tool | work on my own machine and repositories | flat subscription |
| Ask me | none | anything Jev flags as my decision | my time, once |
Research never reaches a frontier model any more. That single rule is where most of the waste was.
A written hand-off after every task. What was asked, what was done, what was verified, what is open, and which lane did it. The next session, in any tool, starts by reading it.
What it looks like from the phone
I typed a question about which cars run well on E85 and whether that beats diesel over a few years. Jev sized it as research. The foreman sent it to a cheap worker, waited, checked the answer, saved the note and replied with the comparison and one line: "via Luna". No frontier model was involved. I did not open a laptop.
Later that afternoon I asked it to review the firewall on my home router. Same words, same phone. This time Jev returned a risk of 0.9 and "operations", so the request went to the frontier lane, and the foreman would have stopped to ask me before changing anything.
Same front door. Different cost. That is the whole point.
"which cars run on E85, and vs diesel?"
Jev: hard 0.98 · research · risk 0.28 · not my decision
→ Tier 1, GPT-5.6 Luna cost: flat
"review the firewall rules on my router"
Jev: hard 2.16 · operations · risk 0.90 · not my decision
→ Tier 2, frontier, would ask before changing anything
"delete the demo server, we don't need it"
Jev: hard 2.47 · risk 0.90 · my decision 0.91
→ stop, ask me first
What broke, honestly
None of this worked on the first run, and I would rather tell you what went wrong than pretend otherwise.
- The foreman gave its first cheap worker an enormous brief, a full live audit, and the worker ran out of time. New rule: a cheap brief is "read the state, answer the question", nothing wider. Live checks the foreman does itself, in one command.
- The foreman started researching the same question in parallel while the worker was busy. Twice the tokens, one answer. New rule: while a worker runs, the foreman waits.
- Jev is confident about area, risk and whether I need to be asked. It is less confident about difficulty when the request is short. So difficulty now comes from the kind of work as much as the score, and a cheap attempt that fails escalates one lane. Doubt does not.
- Two of the frontier models hang at their highest reasoning setting when run without a screen. One notch lower, they answer in seconds.
Each of those is one line in a rules file now. That is the real product of the day: the rules, not the wiring.
Where this goes next
The laptop was the rehearsal. The private AI workspace I run in production, and build for customers, is a different animal: a self-hosted chat platform with its own document store, retrieval and agents, a gateway in front of every model that already routes chat by difficulty and keeps the spend logs, and an assistant of its own that drives those agents. It does not need a second router, and it will not get one just because the laptop has one.
What the test taught me is that the routing was not where the value was. The value came from four operating rules: a brief sized to the tier, one worker at a time, a check before anything is reported, and a written hand-off that the next session reads first. Those rules are plain text. They move into the production workspace first, without touching how it routes.
The routing question stays open and gets measured, not decided. For a few weeks the hand-offs on the laptop and the spend logs in the workspace will show whether Jev's sizing beats the difficulty scorer the workspace already has. If it does, the smallest possible change is to let Jev score inside the existing gateway. If it does not, nothing changes. Either way, that decision is made on numbers, and it is made carefully, because that workspace is where the real work happens.
What a customer gets out of this
I build private AI workspaces for companies, and this is the pattern I would now put in front of one.
- One assistant for the team, reachable from a phone. Nobody chooses a tool. Nobody chooses a model. They say what they need, and the workspace they already have does the work.
- A bill that follows the difficulty of the work. Routine questions, document lookups and research run on cheap or flat-rate models. The expensive ones are reserved, and every use of them is justified in writing.
- Institutional memory the company owns. Rules, state and hand-offs are files in the company's own repository, readable by any tool it adopts later. Changing vendors does not mean starting over.
- A safety gate. Anything that touches a live system, money or something public is flagged before it runs, and a decision that belongs to a human is put to a human, once.
- A record. Every task leaves a note saying what was done, what was checked and what is open. That is what an auditor, a manager or the next engineer wants to read.
What it is not: certain. Jev returns a judgement with a probability attached, workers make mistakes, and the whole design exists because of that. Verification and escalation are the answer to a probabilistic system, not features added to a reliable one.
If you want the same thing
Start with the rules file, not the tools. Write down how you want work delivered, what is off limits and where things live, in a form any assistant can read. Then put one front door in front of the workspace you already have and let something like Jev size the requests for a month. The hand-off notes will tell you more than any benchmark.
I am happy to talk through what this would look like on your own systems. The conversation is where every engagement starts.