Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

Doug Was Right: I Swapped In The MoE, And The Concurrency Math Changed

Last post I measured the concurrency ceiling on a 27B dense model and closed with a promise. Doug Ware, who builds applied AI systems, had handed me the one caveat I did not test: a dense model is the hard case, and a mixture-of-experts model with only about 3B active parameters per token should leave real headroom for a second slot to pay off. So I ran the exact same 36-load matrix on an MLX MoE, changed one variable, and let the numbers settle it. The dense model flatlines. The MoE keeps climbing. And at 8 concurrent agents the dense model makes you wait 32 seconds for a first token while the MoE answers in under 1.

Fabian Williams

9-Minute Read

A line chart showing aggregate throughput staying flat near 20 tokens per second on a dense model while a mixture-of-experts model climbs from 58 to 159 tokens per second as concurrent agents rise from one to eight

Two days ago I posted 36 barrier-synchronized loads against 1 local model and let the numbers answer a question Reddit handed me. That post ended with a promise. A reader named Doug Ware, who builds applied AI systems for a living, had replied with the one caveat I had not tested, and I said out loud that measuring it was the next thing I would do. This is that measurement. I did not wait, I did not hand-wave it, and I kept every honest asterisk in.

You Asked, So I Measured: What Concurrency Actually Costs On One Local Model

Reddit pushed back on my batching post with 3 sharp, testable claims: decode is memory-bandwidth-bound, prefill batches better than decode, and prompt length changes the whole story. So I built a measurement matrix, ran 36 barrier-synchronized loads against 1 local Qwen 3.8 on my Macbook Pro M3 Max, and let the numbers settle it. 2 of the 3 predictions held. 1 did not show up the way I expected, and I am keeping the correction in. Then a real agent handoff stalled on the exact wall these charts describe, and I show how a one-day-old runtime was already a first-class citizen in my receipts governance framework.

Fabian Williams

14-Minute Read

A line chart showing aggregate throughput flattening while per-agent decode rate collapses as concurrent agents rise from one to eight on one local MLX model

Over last week and these last 2 days in this week I posted measured answers to readers both in Twitter and Reddit who asked whether my experiments of 2 local agents on 1 Mac run in parallel or take turns. The answers was that they share one continuous batch, batching is real, and 3 quiet choices collapse it. I thought that was the end of it. It was not. The comments were better than my post, WHICH IS AWESOME, becaue this is the crowdsourcing of brain power I love. I not trying to be a KNOW IT…

I Added a Second Local Agent This Week. Here Is the Receipt for Every Human Decision Behind It.

I trialed a second local coding agent, Hermes from Nous Research, on my own MacBook, pointed at the same local Qwen 3.8 model I already run. Installed additively so nothing already working could break, governed with manual approvals, then tested until it proved it behaves. Here is the trial, in tables and screenshots, plus a receipt for the human hours behind writing it up.

Fabian Williams

7-Minute Read

The Hermes agent running against a local Qwen3.8-27B MLX server, with the Apple M3 Max GPU pinned at 97 percent on the first turn

This week I trialed a second local coding agent, Hermes from Nous Research, on my own MacBook Pro M3 Max, pointed at the same local Qwen 3.8 model I already run. I installed it additively, so nothing already working could break, governed it with manual approvals, and did not stop until it proved it behaves. Here is the trial, in tables and screenshots.

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site