Field Notes
Self-hosted AI, written down.
My technical notes on AI cost, document checks and operating choices. Read each article with its date, assumptions and limits in mind. These are not customer results or service guarantees.
One email when a new essay ships. No other use, unsubscribe by replying, or take the RSS feed.
What are you deciding?
Which AI Model Should Check Your AI's Work?
I sealed four AI models in a sandbox with no view of the answer key, then had each one review the same frozen code for planted defects. Two providers' own safety filters stopped their model from finishing the job before it reached my score sheet.
How Jev Speeds Up My Production AI Review Workflow
Jev is a small decision model I call from code for four narrow jobs on my production AI workflow: picking a model, ranking reference pages, catching bad inputs and suggesting where a file belongs. A person decides every result.
What Is an Automation Engineer, and Do You Need One?
An automation engineer embeds inside a company to turn AI into a working part of a real process, not a demo.
One Agent on My Phone, Five Tools Behind It
A one-day experiment on my laptop: one assistant on my phone takes the request, Jev sizes it, cheap models do the research, frontier models get the hard work, and every task ends with a written hand-off. What moves into my production workspace first is the discipline, not the router.
How to Scope an AI Trial Before Buying Tools
Choose one recurring task, define the data boundary and human checks, and write down when to continue or stop before committing to tools.
Acceptance Checks for an Internal Knowledge Assistant
Write down what a usable answer must do, which sources it may use and when a person must take over. A fictional handbook shows how to specify the checks.
Cheaper AI Tokens. A Bigger Bill.
Kimi K3 charges less per token than GPT-5.6 Sol. Yet one published benchmark reports a higher cost per task. I checked the prices, the token counts and what the comparison actually proves.
AI Pilot Cost: Measure Cost per Accepted Result
Define an accepted result, count failed attempts and human review, and compare the full cost of one recurring task before committing to a build.
Automatic Model Routing: What to Test Before Using It
I assess automatic model routing by task quality, full cost, fallback behaviour and data permissions. A complexity estimate must not decide where data may go.
The LLM Writes Data. A Script Writes the Document.
I separate AI findings from document rendering so the source data, review decisions and output can be checked without regenerating the reasoning.
Agent Finds, Human Checks: Using Gated Vendor Knowledge
I separate finding a reference from permission to retrieve or share it. Human access does not automatically authorise copying gated content into an AI system.
Prompt Caching: Verify Configuration Against Actual Usage
I verify cache settings through accepted configuration and provider usage. A misspelled or unsupported option can look correct in a file and never take effect.
How I Evaluate a Cheaper Model for a Second Opinion
A cheaper reviewer is useful only if it catches the errors that matter. Here is how I would test one without confusing agreement with correctness.
When an AI Validator Invents a Rule
I check AI rule citations against an agreed source. A plausible identifier is not evidence, and a correct reference does not prove the conclusion.
Why an Agent Can Use Less Than Its Model’s Context Window
I check the effective context limit, framework configuration and actual requests before relying on a model’s advertised window for long tasks.
An AI Coding App Indexed 627,652 Files On My Laptop
OpenAI's Codex desktop app registered my home directory as a workspace and kept scanning a project folder deleted three months earlier. The interesting part isn't the bug. It's that I could find it at all.
Named or Assumed: The First Question in Architecture Review
The weak point of most architecture dossiers is the sentence that says security is 'handled by the platform.' Here's the one distinction that separates auditable inheritance from a gap in disguise.
AI Router Health Checks: Count Cost and Detection Delay
I count scheduled API calls and compare their cost with the required failure-detection time before changing a model-router health-check interval.
Buying an AI Server vs Paying for OpenAI API Tokens: A TCO Comparison
Is buying an AI server cheaper than paying for OpenAI API tokens? Compare hardware, subscriptions and API usage with setup, power, operator time and result quality included.
3 Hidden Failure Modes of Self-Hosted LLMs Past 25 Users
KV cache fragmentation, the context-window I/O trap, and kernel scheduler thrashing — the three failure modes that crash self-hosted LLM stacks at scale, and how to fix them.
OCR and Scanned PDFs: Map the Whole Document Path
Before calling a document workflow private, I check where OCR, embeddings, logs and backups actually run. A local model alone does not settle it.
vLLM vs llama.cpp vs Ollama at 25 Concurrent Users
Single-user local inference is a solved problem. Production LLM serving at 25 concurrent users is a different engineering domain — and the published benchmarks all agree on the shape of the failure.
OpenAI DPA and GDPR Article 28: What to Check
A practical review of the contract, API retention, regional processing and the decisions your team still owns before sending personal data to an AI service.