Back to overview

Field Notes

Self-hosted AI, written down.

My technical notes on AI cost, document checks and operating choices. Read each article with its date, assumptions and limits in mind. These are not customer results or service guarantees.

One email when a new essay ships. No other use, unsubscribe by replying, or take the RSS feed.

What are you deciding?

agentsmodel-evaluationai-governance

Which AI Model Should Check Your AI's Work?

I sealed four AI models in a sandbox with no view of the answer key, then had each one review the same frozen code for planted defects. Two providers' own safety filters stopped their model from finishing the job before it reached my score sheet.

10 min read
jevmodel-routingprivate-ai

How Jev Speeds Up My Production AI Review Workflow

Jev is a small decision model I call from code for four narrow jobs on my production AI workflow: picking a model, ranking reference pages, catching bad inputs and suggesting where a file belongs. A person decides every result.

7 min read
automation-engineerai-automationsmall-business-ai

What Is an Automation Engineer, and Do You Need One?

An automation engineer embeds inside a company to turn AI into a working part of a real process, not a demo.

7 min read
private-aiagentsmodel-routing

One Agent on My Phone, Five Tools Behind It

A one-day experiment on my laptop: one assistant on my phone takes the request, Jev sizes it, cheap models do the research, frontier models get the hard work, and every task ends with a written hand-off. What moves into my production workspace first is the discipline, not the router.

10 min read
ai-trialrecurring-taskhuman-review

How to Scope an AI Trial Before Buying Tools

Choose one recurring task, define the data boundary and human checks, and write down when to continue or stop before committing to tools.

6 min read
internal-knowledge-assistantacceptance-checkshuman-review

Acceptance Checks for an Internal Knowledge Assistant

Write down what a usable answer must do, which sources it may use and when a person must take over. A fictional handbook shows how to specify the checks.

7 min read
llm-costtoken-pricingkimi-k3

Cheaper AI Tokens. A Bigger Bill.

Kimi K3 charges less per token than GPT-5.6 Sol. Yet one published benchmark reports a higher cost per task. I checked the prices, the token counts and what the comparison actually proves.

6 min read
ai-pilotai-roiprivate-ai

AI Pilot Cost: Measure Cost per Accepted Result

Define an accepted result, count failed attempts and human review, and compare the full cost of one recurring task before committing to a build.

6 min read
model-routinglitellmself-hosted-ai

Automatic Model Routing: What to Test Before Using It

I assess automatic model routing by task quality, full cost, fallback behaviour and data permissions. A complexity estimate must not decide where data may go.

4 min read
agentsai-governancellm-operations

The LLM Writes Data. A Script Writes the Document.

I separate AI findings from document rendering so the source data, review decisions and output can be checked without regenerating the reasoning.

3 min read
agentsai-governancecompliance

Agent Finds, Human Checks: Using Gated Vendor Knowledge

I separate finding a reference from permission to retrieve or share it. Human access does not automatically authorise copying gated content into an AI system.

3 min read
cost-controlllm-operationsprompt-caching

Prompt Caching: Verify Configuration Against Actual Usage

I verify cache settings through accepted configuration and provider usage. A misspelled or unsupported option can look correct in a file and never take effect.

3 min read
agentsmodel-evaluationcost-control

How I Evaluate a Cheaper Model for a Second Opinion

A cheaper reviewer is useful only if it catches the errors that matter. Here is how I would test one without confusing agreement with correctness.

3 min read
agentsai-governancecompliance

When an AI Validator Invents a Rule

I check AI rule citations against an agreed source. A plausible identifier is not evidence, and a correct reference does not prove the conclusion.

3 min read
agentslibrechatoperations

Why an Agent Can Use Less Than Its Model’s Context Window

I check the effective context limit, framework configuration and actual requests before relying on a model’s advertised window for long tasks.

3 min read
sovereign-aiself-hosted-aivendor-risk

An AI Coding App Indexed 627,652 Files On My Laptop

OpenAI's Codex desktop app registered my home directory as a workspace and kept scanning a project folder deleted three months earlier. The interesting part isn't the bug. It's that I could find it at all.

8 min read
architecture-reviewsecurityvendor-assessment

Named or Assumed: The First Question in Architecture Review

The weak point of most architecture dossiers is the sentence that says security is 'handled by the platform.' Here's the one distinction that separates auditable inheritance from a gap in disguise.

5 min read
cost-controlself-hosted-ailitellm

AI Router Health Checks: Count Cost and Detection Delay

I count scheduled API calls and compare their cost with the required failure-detection time before changing a model-router health-check interval.

3 min read
self-hosted-aitcoopenai

Buying an AI Server vs Paying for OpenAI API Tokens: A TCO Comparison

Is buying an AI server cheaper than paying for OpenAI API tokens? Compare hardware, subscriptions and API usage with setup, power, operator time and result quality included.

8 min read
self-hosted-aiinfrastructurevllm

3 Hidden Failure Modes of Self-Hosted LLMs Past 25 Users

KV cache fragmentation, the context-window I/O trap, and kernel scheduler thrashing — the three failure modes that crash self-hosted LLM stacks at scale, and how to fix them.

11 min read
ocrself-hosted-aicompliance

OCR and Scanned PDFs: Map the Whole Document Path

Before calling a document workflow private, I check where OCR, embeddings, logs and backups actually run. A local model alone does not settle it.

4 min read
vllmllamacppollama

vLLM vs llama.cpp vs Ollama at 25 Concurrent Users

Single-user local inference is a solved problem. Production LLM serving at 25 concurrent users is a different engineering domain — and the published benchmarks all agree on the shape of the failure.

7 min read
gdprcomplianceopenai

OpenAI DPA and GDPR Article 28: What to Check

A practical review of the contract, API retention, regional processing and the decisions your team still owns before sending personal data to an AI service.

7 min read