<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Local LLM on Fabian G. Williams</title>
    <link>https://www.fabswill.com/tags/local-llm/</link>
    <description>Recent content in Local LLM on Fabian G. Williams</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Sun, 02 Aug 2026 00:00:00 +0000</lastBuildDate>
    
	<atom:link href="https://www.fabswill.com/tags/local-llm/index.xml" rel="self" type="application/rss+xml" />
    
    
    <item>
      <title>I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.</title>
      <link>https://www.fabswill.com/blog/local-model-frontier-rematch-auto-replay-evals/</link>
      <pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate>
      
      <guid>https://www.fabswill.com/blog/local-model-frontier-rematch-auto-replay-evals/</guid>
      <description>TL;DR My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?
Today I answered it with a system instead of a guess. I taught my eval framework to automatically replay every writing task on the local model right after the frontier model runs it, then grade both with the same judge and record the gap.</description>
    </item>
    
  </channel>
</rss>