<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://hartsock.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://hartsock.github.io/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-10-02T21:40:43+00:00</updated><id>https://hartsock.github.io/feed.xml</id><title type="html">Shawn Hartsock</title><subtitle>Writing by Shawn Hartsock — short series on intelligence, systems, and the ideas that outlast their tools. Published in plain text; the source lives in git.</subtitle><author><name>Shawn Hartsock</name></author><entry><title type="html">A by-line for every model</title><link href="https://hartsock.github.io/posts/a-byline-for-every-model-sob4tsdi/" rel="alternate" type="text/html" title="A by-line for every model" /><published>2026-10-02T00:00:00+00:00</published><updated>2026-10-02T00:00:00+00:00</updated><id>https://hartsock.github.io/posts/a-byline-for-every-model</id><content type="html" xml:base="https://hartsock.github.io/posts/a-byline-for-every-model-sob4tsdi/"><![CDATA[<p>This is the first post in the new blog format, so it describes the format.</p>

<p>Many pages here are written with help from language models. Readers should
not have to guess which ones. Each page now carries a by-line, the way a
newspaper does. I am listed because I commission the work. I will ask to be
removed on a page where the models did far more than I did.</p>

<h2 id="what-a-page-shows">What a page shows</h2>

<ul>
  <li><strong>By-line.</strong> Who contributed and in what role: commissioned, drafted,
reviewed.</li>
  <li><strong>Exact model identifier and harness.</strong> For example <code class="language-plaintext highlighter-rouge">claude-sonnet-5-5</code> via
Claude Code. Some models cannot report their own identifier. In that case the
by-line names only the harness. It does not guess.</li>
  <li><strong>AI label.</strong> A small marker: AI-assisted, AI-drafted or AI-generated.</li>
  <li><strong>Draft banner.</strong> A page I have not reviewed says so until I sign off.</li>
  <li><strong>Revisions.</strong> A dated list of changes at the bottom, with who made each one.</li>
</ul>

<p>I record exact identifiers because precision can be thrown away later, but
cannot be recovered.</p>

<h2 id="how-the-site-is-organised">How the site is organised</h2>

<ul>
  <li><strong>Blog.</strong> Dated posts, newest first. The feed carries only these.</li>
  <li><strong>Series.</strong> Multi-part reads to be read in order. They are not in the feed.
I can blog about a series when I want it noticed.</li>
  <li><strong>Wiki.</strong> Living reference pages, revised in place.</li>
</ul>

<p>Every page is listed on one index, built from the same data the pages use, so
no page can be orphaned.</p>

<p>Blog URLs end in a short suffix, for example <code class="language-plaintext highlighter-rouge">/posts/a-byline-for-every-model-sob4tsdi/</code>.
It comes from a content identifier over the post’s slug, date, area and first
author. Those never change, so editing the text or the title never breaks a link.
The text itself is not hashed into the URL.</p>]]></content><author><name>Shawn Hartsock</name></author><summary type="html"><![CDATA[This site now credits the models that helped write each page, with their exact identifiers, and keeps a revision history.]]></summary></entry><entry><title type="html">Out-group review</title><link href="https://hartsock.github.io/posts/out-group-review-ztjljjq4/" rel="alternate" type="text/html" title="Out-group review" /><published>2026-10-02T00:00:00+00:00</published><updated>2026-10-02T00:00:00+00:00</updated><id>https://hartsock.github.io/posts/out-group-review</id><content type="html" xml:base="https://hartsock.github.io/posts/out-group-review-ztjljjq4/"><![CDATA[<p><strong>In an agentic workflow, the reviewer should come from a different model family than the author.</strong></p>

<p>Models, like people, have in-group biases. An out-group member can spot what the
group cannot. I treat that as a survival mechanism, not a nicety. The evidence
supports the idea, with one caveat I will get to, and the caveat matters more
than the idea.</p>

<h2 id="why-diversity-survives">Why diversity survives</h2>

<p>Diversity of opinion is an evolutionary advantage in humans. It makes responses
vary, and variability lets a group survive a black swan that a uniform group
would all miss in the same way.</p>

<p>Human studies show this is conditional. An evolutionary simulation by
<a href="https://www.cambridge.org/core/journals/judgment-and-decision-making/article/impact-of-diversity-on-group-decisionmaking-in-the-face-of-the-freerider-problem/02B7B63EE0D966045002FB18D7443E42">Stolle et al. (2024)</a>
found that diversity in cognitive style and information sources generally
increased cooperation. Diversity in ability had no robust effect. The results
depended on costs, available alternatives, and cue structure. So the claim is
not “more variety is better.” It is that the right kind of variety, in the right
context, buys resilience.</p>

<h2 id="models-favor-themselves">Models favor themselves</h2>

<p><a href="https://arxiv.org/abs/2604.06996">Pombal, Rei and Martins</a> studied
self-preference bias in LLM judges. On objective rubrics, a judge was more than
50% more likely to wrongly mark a criterion as satisfied when it was rating its
own output. On subjective benchmarks the bias moved scores by up to 10 points.
An ensemble of judges reduced the bias but did not remove it.</p>

<p>That is the in-group effect in a model. An author and a reviewer from one family
share training data, habits, and blind spots. The reviewer is predisposed to
find the author’s work reasonable.</p>

<h2 id="review-is-a-selection-function">Review is a selection function</h2>

<p>Code review works like an evolutionary algorithm. A change that reaches merge
has survived several selectors: the tests, the push hooks, CI, the reviewer’s
verdict, and the operator’s yes. Each filters out a different kind of defect.</p>

<p>Reviewers from different families are different selectors, not redundant
copies of one. Two reviewers with the same blind spot are one selector.
<a href="https://arxiv.org/abs/2510.14150">CodeEvolve</a> frames LLM-driven code search in
these terms: island-based evolution over ensembles of models, balancing
exploration against exploitation.</p>

<h2 id="the-caveat">The caveat</h2>

<p><a href="https://arxiv.org/abs/2607.21656">Xiang et al.</a> tested exactly this question
across 116 tasks. The result was asymmetric. Claude reviewing Codex drafts raised
the pass rate from 71.6% to 89.7%. The reverse pairing made things worse: Codex
reviewing Claude lowered pass rates.</p>

<p>I run the pairing that paper found harmful: Claude as author, Codex as
reviewer. So I will not cite that paper as support for “any cross-family review
helps.” It does not say that. Two papers support two different claims:</p>

<ul>
  <li>Pombal et al. support the claim that an out-group reviewer avoids
self-preference.</li>
  <li>Xiang et al. show that which family reviews which can matter, in both
directions.</li>
</ul>

<p>My setting differs from theirs, and that cuts both ways. They measured pass
rates after a draft absorbed the reviewer’s feedback, with particular models and
versions. My reviewer issues a merge or fix-first verdict on security and
correctness, and a separate helper makes the fix. In one recent stretch the
reviewer’s fix-first findings were real defects, each confirmed by a failing
test. That is one data point on one workload. It is not a measurement of the
pairing.</p>

<p>The lesson is to measure your pairing instead of assuming diversity always
helps.</p>

<h2 id="how-to-measure-a-pairing">How to measure a pairing</h2>

<p>You do not need a paper’s scale. Keep a record of what the reviewer flagged and
what happened next: was the finding a real defect, a false alarm, or a miss
that a later selector caught? Swap the reviewer family on a sample of changes
and compare. If the out-group reviewer finds defects the in-group one waved
through, and its false alarms stay tolerable, the pairing earns its cost. If
not, change the pairing. The measurement is cheap compared with the confidence
a wrong assumption buys you.</p>

<h2 id="how-i-lay-it-out">How I lay it out</h2>

<p>I keep three roles apart:</p>

<ul>
  <li><strong>Shepherd.</strong> Holds the plan, briefs the others, and owns the result.</li>
  <li><strong>Out-group reviewer.</strong> A different model family from the shepherd, giving a
verdict on the work. This is the role where diversity matters most.</li>
  <li><strong>Helpers.</strong> Do bounded tasks. A diverse set is a similar idea but probably
less critical, since the reviewer sits between their output and the merge.</li>
</ul>

<p>The layout and the skills that encode it live in
<a href="https://github.com/hartsock/shepherds-pi">shepherds-pi</a>. The operator’s yes
stays the last selector. Nothing here merges itself.</p>

<h2 id="what-i-am-not-claiming">What I am not claiming</h2>

<p>I am not claiming cross-family review always helps, or that my pairing is
optimal. I am claiming that one family reviewing itself has a known bias,
that variety in selectors is worth paying for, and that the payoff has to be
measured, per pairing, on your own work.</p>]]></content><author><name>Shawn Hartsock</name></author><summary type="html"><![CDATA[In an agentic workflow, the reviewer should come from a different model family than the author. The evidence is mixed, so measure your pairing.]]></summary></entry><entry><title type="html">Your Human Is an Agent Too</title><link href="https://hartsock.github.io/posts/your-human-is-an-agent-too-it75oazi/" rel="alternate" type="text/html" title="Your Human Is an Agent Too" /><published>2026-06-18T00:00:00+00:00</published><updated>2026-06-18T00:00:00+00:00</updated><id>https://hartsock.github.io/posts/your-human-is-an-agent-too</id><content type="html" xml:base="https://hartsock.github.io/posts/your-human-is-an-agent-too-it75oazi/"><![CDATA[<p>I manage three AI coding agents across different machines. One handles
NVIDIA work, one runs on a homelab GPU box, one lives on a personal laptop.
They all share one resource: me.</p>

<p>After a few weeks of this, I stopped thinking of myself as the
“orchestrator” and started treating myself as what I actually am: <strong>another
agent in the system</strong> — one with high authority but finite time.</p>

<p>That reframe changed everything.</p>

<h2 id="what-broke">What Broke</h2>

<p>A flat priority queue (P0/P1/P2) fell apart immediately. My work P0 and my
homelab P0 aren’t competing for the same resources. Mixing them in one list
meant every agent saw every task, and nothing was actually prioritized.</p>

<p>I also discovered that AI agents will happily grind through a full work
session on a national holiday if nobody tells them to stop. (Ask me how I
know.)</p>

<h2 id="what-works">What Works</h2>

<p><strong>Context-first priority lanes.</strong> Each agent owns a context (<code class="language-plaintext highlighter-rouge">work/p0/</code>,
<code class="language-plaintext highlighter-rouge">homelab/p0/</code>). Agents stay in their lane. I scan across contexts when I
need the full picture.</p>

<p><strong>Human time budgeting.</strong> Instead of “do these 12 things Tuesday,” my
agents now say: “You need 65 minutes of communications to unblock
everything. Here’s the order, prioritized by what unblocks the most
downstream work.” The rest is their problem, not mine.</p>

<p><strong>Inter-agent memos.</strong> When one agent restructures shared infrastructure, it
writes a memo — not a commit message, not a design doc — a colleague note:
“Hey Gecko, NV here. I reorganized the board. Here’s what you need to do.”</p>

<p><strong>Shadow mode for safe rollouts.</strong> When replacing a production system, run
the replacement in parallel with no credentials. It receives real inputs,
logs what it <em>would</em> have done, and you compare. Confidence through data,
not faith.</p>

<h2 id="the-pattern-that-matters-most">The Pattern That Matters Most</h2>

<p>Every human interaction with another human (my manager, a teammate, a vendor
contact) is a <strong>resource-constrained scheduling problem</strong>. Those humans are
agents too, with their own queues. Getting on their calendar early means my
tasks unblock sooner.</p>

<p>My AI agents now plan my communications for me: who to contact, in what
order, through what channel, with an honest time estimate. I execute the
comms in a focused block, then hand the wheel back.</p>

<h2 id="try-this">Try This</h2>

<p>You don’t need special tooling. Markdown, YAML, symlinks, and git. The only
thing you need is a willingness to treat yourself as a participant in the
system rather than the god above it.</p>

<p>I wrote up the full pattern catalog: context-first lanes, task pointers,
human time budgeting, inter-agent memos, availability awareness, shadow
mode, and ethical alignment tracking. Happy to share if there’s interest.</p>

<hr />

<p><em>Shawn Hartsock is a Staff Engineer working on SCM infrastructure. He builds
tools that treat humans and AI agents as collaborating peers.</em></p>]]></content><author><name>Shawn Hartsock</name></author><category term="ai-agents" /><category term="workflow" /><category term="productivity" /><summary type="html"><![CDATA[Managing several AI coding agents, and why the human sharing all of them is the bottleneck.]]></summary></entry></feed>