Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

Limen: I Put a Camera at My Window and Asked Who Owns What It Sees

This week is Fix, Hack, Learn at Microsoft, and my project is a local-first street-vision system I am calling Limen, the Latin word for threshold. An Arduino board watches my front window, a vision model on my own MacBook Pro M3 Max names what crosses (person, vehicle, wildlife), and a human confirms the label to build a training set I actually own. No cloud, no subscription, no frame leaving the house. Today I ran a scripted 90-second proof of concept: I walked up, drove in, parked, and walked to the door, and every one of 180 frames was auto-labeled locally in about 8.4 seconds each. Here is what worked, what did not (plates are still unreadable), and the one hard constraint that shaped the whole architecture: my evals rig is sacred and this project was never allowed to touch it.

Fabian Williams

9-Minute Read

A view from a front window: a man walking up the path toward the door holding a bag and a bottle, boxed on device with a green person label, with his vehicle boxed and a blue vehicle label, and a caption reading Limen local VLM auto-label on-device no cloud

This week is Fix, Hack, Learn (FHL) at work (Microsoft), our internal week to chase a build we care about. Last FHL I worked on my agents and agentic loops, helping me scale myself and my work as a product manager. This time my project is personal, and it starts with a question I could not stop turning over: when a camera watches the front of my house, who owns what it sees?

Doug Was Right: I Swapped In The MoE, And The Concurrency Math Changed

Last post I measured the concurrency ceiling on a 27B dense model and closed with a promise. Doug Ware, who builds applied AI systems, had handed me the one caveat I did not test: a dense model is the hard case, and a mixture-of-experts model with only about 3B active parameters per token should leave real headroom for a second slot to pay off. So I ran the exact same 36-load matrix on an MLX MoE, changed one variable, and let the numbers settle it. The dense model flatlines. The MoE keeps climbing. And at 8 concurrent agents the dense model makes you wait 32 seconds for a first token while the MoE answers in under 1.

Fabian Williams

9-Minute Read

A line chart showing aggregate throughput staying flat near 20 tokens per second on a dense model while a mixture-of-experts model climbs from 58 to 159 tokens per second as concurrent agents rise from one to eight

Two days ago I posted 36 barrier-synchronized loads against 1 local model and let the numbers answer a question Reddit handed me. That post ended with a promise. A reader named Doug Ware, who builds applied AI systems for a living, had replied with the one caveat I had not tested, and I said out loud that measuring it was the next thing I would do. This is that measurement. I did not wait, I did not hand-wave it, and I kept every honest asterisk in.

You Asked, So I Measured: What Concurrency Actually Costs On One Local Model

Reddit pushed back on my batching post with 3 sharp, testable claims: decode is memory-bandwidth-bound, prefill batches better than decode, and prompt length changes the whole story. So I built a measurement matrix, ran 36 barrier-synchronized loads against 1 local Qwen 3.8 on my Macbook Pro M3 Max, and let the numbers settle it. 2 of the 3 predictions held. 1 did not show up the way I expected, and I am keeping the correction in. Then a real agent handoff stalled on the exact wall these charts describe, and I show how a one-day-old runtime was already a first-class citizen in my receipts governance framework.

Fabian Williams

14-Minute Read

A line chart showing aggregate throughput flattening while per-agent decode rate collapses as concurrent agents rise from one to eight on one local MLX model

Over last week and these last 2 days in this week I posted measured answers to readers both in Twitter and Reddit who asked whether my experiments of 2 local agents on 1 Mac run in parallel or take turns. The answers was that they share one continuous batch, batching is real, and 3 quiet choices collapse it. I thought that was the end of it. It was not. The comments were better than my post, WHICH IS AWESOME, becaue this is the crowdsourcing of brain power I love. I not trying to be a KNOW IT…

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site