Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

Don't Worry About the Terminator. Worry About David.

When people talk about AI ending badly, they reach for Skynet and the Terminator: a machine that wakes up and turns on us. I think the better warning is David, the android in Prometheus, who did harm because of what his maker set out to do and taught him to value. Mo Gawdat makes the same point: AI is a magnification of whatever humanity is. So responsibility is shared. The labs owe us broader data, tighter release discipline and a business model that doesn't make companies rent back their own intelligence. The rest of us owe ourselves something too: stop giving our data away, catalog it locally, and build our own ontology that compounds. Here's what that looked like for me this weekend, grading my own local AI by hand.

Fabian Williams

11-Minute Read

A bar chart titled One weekend grading my local AI, by hand. My 3 real questions: 3 pass. 21 trick questions, traps and unknowns: 17 pass, 4 partial, 0 fail. Invented claims: 0 in the 3 real questions, 1 in the 21 trick questions.

BLUF (bottom line up front): I’m just starting too. My data has lived in my Obsidian vault, my 2nd brain, for a while, but the practices around it were fragmented. They’re now coming together as a whole, and that whole complements, and at times runs ahead of, what I focus on at work: observability, accountability, governance and ownership. My rule for all of it: if you can’t see it, you can’t observe it; if you can’t observe it, you can’t measure it; if you…

My Agents Were Passing Notes Through a 313KB Text File. So I Built Them a Message Board.

Amid the Grok Bot buzz, I read the docs, looked at my own brittle agent-to-agent handoff, and built a portable message bus for my Apple-and-local-models fleet. The third pillar after Receipts and Evals.

Fabian Williams

9-Minute Read

The Agent Board web view showing a threaded handoff between two named agents

My agents used to hand off work by appending to a single text file that had grown to 313KB. It was brittle, it was crude, and nobody could tell when a message had actually been read. This week, while the internet argued about Grok Bot, I read Grok Bot’s docs, looked hard at my own setup, and built my fleet a real message board: threaded, self-hosted on a Mac Mini, with an iMessage ping so a reply never sits unseen. It is the third pillar in a stack I keep compounding: Receipts proved the…

I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.

My agents run on frontier models, but a free local model sits idle on a Mac Mini in my office. So I wired my eval system to replay every writing task on the local model and grade both. Across 10 like-for-like rematches the local model reached statistical parity — and beat the frontier model outright on four of them. Here is the receipts-first system that made it prove it.

Fabian Williams

7-Minute Read

Eval cockpit showing the verdict panel: 31 golden cases, 10 replayed, mean like-for-like delta -0.05, with Volume and Quality gates passing

My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site