Panout review packOverviewFinal memoMVPLive outputsJudge roundsAppendicesDecisions

Source: docs/research/2026-09-03-defensibility/rounds/r5-judgment.md

Panout round 5 — investment judgment: "The rework layer for agent labor"

Judge: Garry Tan lens. Methodology: gstack office-hours startup mode (six forcing questions, anti-sycophancy rules, position on every answer, end with an assignment) plus plan-ceo-review cognitive patterns (inversion reflex, proxy skepticism, focus as subtraction, classification instinct, temporal depth). Date: 2026-09-03. Self-imposed dogfood gate: 2026-09-05 — reported by the founder as a failure, and re-verified independently below. Question scored: would I write the check today for a defensible, $1B-capable company, given this founder's constraints and this evidence? Calibration: 62/72, 66, 67, 64. Read for scale only. Round 4 pre-registered that the memo-only ceiling was ~70 conditional on appendix H coming back positive, and stated in terms that a negative H means "the honest score goes down, not up." H came back negative. I checked every number in the memo against the appendix it cites and re-derived three from the machine.


Score: 59/100

Verdict: the fifth honest memo in a row, the best single piece of analysis in the whole file, and the lowest score yet — because appendix H did not merely fail to rescue the map, it measured the map drifting, and the founder demoted the map without demoting the four other places in the document that depend on it.


Dimension table

#DimensionScoreReasoning
1Q1 Demand reality6Appendix E is unchanged and still the strongest artifact in the file: 249 upvotes on one VS Code issue, 78 on a Claude Code meta-issue whose title concedes the community is building workarounds, ~15k shell files invoking the skip-permissions flag, auto-approve hooks shipped to customers by Railway and Render, Anthropic's own 20%→40% auto-approve curve. Round 4 held this at 6 solely because of the sample-mismatched merge-latency claim; that is now fixed correctly and in the memo's own voice ("in popular repos the difference disappears... Codex-branch PRs at 28 minutes, slower than the control"). So the fix earned its point back. It is spent immediately on what the fix reveals: the population where agent output demonstrably outruns reading is the long tail of small and solo repos, and that is not the population your $1B denominator counts. You repaired the claim by shrinking the buyer, and the memo does not reconcile the two. Sixteen days after the assignment, still zero humans.
2Q2 Status quo6Positions A/B/C remain the cleanest competitive framing anyone has written for this space, and the vendor-doc sweep behind "learned from your own history and recorded durably" is still the narrowest defensible gap claim in five rounds. Held. Appendix H actually gives you something here that you did not take: with bots excluded and calendar spans matched, the pooled public gap under your own adopted metric is +7.7 points with a repo-clustered CI of [+1.3, +13.5], and 4 of 11 repos print positive with CI excluding zero against 0 negative. That is a better answer than round 4's +2.5-through-zero. You declined to claim it because the separation is carried by the "reworked by a different author" clause, which is an ownership proxy. That refusal is correct and I credit it under integrity — but it means the cost of Position B for a stranger is still, in the memo's own words, "an unknown that differs per team," and H has now added: an unknown that also moves. Selling a variance is hard. Selling a drifting variance is harder.
3Q3 Desperate specificity4Unchanged, and unchanged is the finding. One named human, and it is you: 99.9th percentile agent volume, everything in bypassPermissions, zero stashes machine-wide, 91.6% of sessions with no negative event. Amazon is still the best named external consequence and volunteering the dispute is still right. Sixteen days ago: ten operators, two numbers and a yes/no, nine hours. Twelve hours ago: eleven maintainers whose repos are on your disk, one message each. Neither exists. You cannot score above 4 here with an N of one, and the N has been one for five rounds.
4Q4 Narrowest wedge5Down one, and this is the most expensive line. Round 4 said your hook was a coin flip whose coin you had not tested. You tested it. It is a coin: readable per-class in 2 of 11 repos, class ordering across halves at rho −0.10 on k=2 classes where the statistic can only be ±1, and 7 of 11 repos indeterminate under R4-strict. You responded honestly — step 1 is demoted to a diagnostic and minute-one value is relocated to "contracts that print only failures, and the override record they generate... a linter with a memory on day one." Two problems, and they are structural, not editorial. First, that is not day-one value to the user. By your own sentence, "the customer's CI already runs the checks." On day one the free tier runs your checks and records that you overrode them; the user gets nothing back until day 30. That is deferred value requiring sustained installation — which is v1's inbox and v2's brief wearing a third costume. You assert the failure mode "cannot recur" because capture is passive; capture being passive fixes your data problem, not the user's reason to keep it installed. Second, step 3 survived unamended. "Contracts for the two worst classes" is still in the wedge, and H just told you the two worst classes are unreadable in 9 of 11 repos. You removed the input and kept the step.
5Q6 Future fit6Down one. The core argument still holds and is still argued rather than asserted: rework has a floor, not a zero; better models shift which task classes are safe rather than removing the need to know which ones. But H creates an objection to it that the memo does not answer and I think is serious. Your asset accumulates over months. H measured the target quantity drifting over calendar time faster than sampling noise can explain — observed split-half rho +0.164 against a random-re-split control whose median is +0.48, i.e. the 5th percentile of pure noise. The halves are two to three months apart. That is the same clock as a model generation. If what is safe changes on the same timescale as the measurement needed to establish what is safe, the map is permanently describing a world that has already moved, and the thing you are compounding depreciates as fast as it accrues. Round 3 moved the load-bearing quantity from a level to a variance; round 5's own evidence says the variance is non-stationary. Nobody has priced that yet, including you.
6Moat mechanism6Down one, and not because the moat sentence got worse — you adopted round 4's rewrite verbatim and it is the right sentence, deliberately independent of the rework gap. Layer 1 survives H completely and remains the best physical idea in the file: the override record exists in no vendor's log and no git history, and appendix G's trailer-vs-PR sign flip is genuine proof that post-hoc authorship cannot be recovered by anyone. Layer 3 arguably strengthens under H — if per-team estimates are noisy and drifting, the pooled cross-team prior is the only estimator with enough n, which is a real argument you did not make. Layer 2 is the problem, and Layer 2 is the switching cost. "Each task class carries a level from 0 to 4" is a per-class claim, and H says the class is below the resolution of the measurement at 100–200 agent commits. Your day-60 milestone is "first autonomy level earned in at least one task class." You have not shown that quantity is estimable at any sample size a real team produces in five months, and you now have direct evidence it is not estimable at the sample sizes eleven busy public repos produce in five. A moat whose unit is unmeasurable is a moat whose unit is wrong — see "strongest defensible version," where I think the fix is available and cheap.
7Incumbent response7Unchanged and still good. The table is honest where it costs: harness auto-approval is "the closest shipped surface and the memo does not claim it is empty," Entire is UNVERIFIED and the row says so, GitHub's approving-reviews preview is disclosed and answered with switching cost rather than neutrality. Nothing in H changes it, in either direction. The standing exposure is unchanged: "Panout consumes their decisions as evidence" makes you a dependent of six vendors' export surfaces, and the retention leg of your own moat sentence is a setting one of them controls.
8$1B credibility7Held. All three round-4 fixes landed and I checked them: the second and third denominators are there (Cursor ~7M MAU / 6 ≈ 1.2M teams; Copilot 26M / 10 = 2.6M) and land in the same band, so no single secondary-sourced number carries the market; and the meter/autonomy clause is in the $1B sentence in almost exactly the required form. Arithmetic re-checked: 2,000 × 500 × $0.33 = $3.96M, 50,000 → $99M, penetration 0.4%/2%/10%. That is +1 of work. It is offset exactly once: the merge-latency repair (dimension 1) locates your demonstrable buyer in the long tail of small repos, while the denominator counts Codex weekly actives and Copilot seats across all org sizes — populations where your own appendix B says the effect disappears and review culture was already thin. State which of those two companies you are, because the ramp is built on one and the evidence on the other.
9Founder fit4Down one, and I want to be precise about why, because the timeline is compressed and I am not punishing hours. Round 4's assignment was written specifically to pre-empt the substitution it expected: "The deliverable is eleven sent messages, not a twelfth appendix," with four hours of split-half work named as a prerequisite, not the deliverable. You did the four hours — superbly, in a day, with a noise-floor control most professionals would not think to run — and appendix H's own third line records: "no maintainer was contacted." That is not a scheduling artifact. The exact failure mode was named to your face in one sentence and executed anyway within hours. Fourth time. Meanwhile I re-verified the machine myself today: last commit c1a28dcf, 2026-08-23, eleven days; no .panout anywhere; no ledger on disk. Commit-time capture is the one clock in this company that cannot be compressed, round 4 called it a one-way door on the calendar with a forty-eight-hour deadline, and it has not started. On paper the fit is still the best version of this company — repo, terminal, commits, zero new habits, 5,294 founder-authored commits in 180 days as substrate. In revealed behavior, five rounds have produced five documents, eight appendices, four fourteen-day plans, one failed gate, zero commits and zero humans.
10Evidence integrity8Held at 8, and the split inside it is wider than ever: the measurement is a 9 and the propagation is a 7. Appendix H is the best analytical work in the file and one of the better things I have seen a solo founder produce. Two moves earn that: the random-re-split noise floor (A5), which converts "weak correlation" into the strictly worse finding "actively drifting," and running it anyway; and caveat 2, which names thin cells as the alternative explanation for your own headline and then shows the control already accounts for it. You also ran objection 3's idea-half without being asked twice, got 7-of-11 indeterminate, and reported it. And the memo's reporting of H is faithful — I checked every figure it lifts (0.164, [−0.62,+0.90], 7/11, 0.48 median, 5th percentile, +7.7 [1.3,13.5], 2-of-11 readable) and all of them tie. Zero misreporting of the appendix that hurts most. Four residues cost the ninth point, and two are load-bearing. (i) The closing section still says "The founder's 22-point gap is robust to every filter (appendix F)" — but the product adopts R4-strict, under which the founder gap is +11.2 [5.8, 16.9], and F's own §5 records a new reversal in dynorouter-gateway at −11.7 on 44 agent commits. Your body says "between 11 and 23 depending on definition." Your summary says 22. That is the same shading round 4 flagged, relocated one section down. (ii) The one-liner still reads "It shows a team which classes of agent work get silently fixed later and which do not" — the capability H just measured as unreadable in 9 of 11 repos. You demoted the map in the wedge and left it in the sentence that defines the company. (iii) Appendix D's own power note says detecting a 15pp rework difference at a 6% exposure rate needs 10³–10⁴ linked commits; your day-60 kill criterion specifies 300. Memo contradicts appendix by an order of magnitude — see the hardest push. (iv) Minor: "a 71-commit gateway service, 44 agent commits" — F measures 44 agent + 24 human = 68.

Raw: 59/100.

Gut adjustment: 0. Stated in both halves, because they cancel exactly. Up +2: you ran the single experiment that could kill the reframe you had just retreated to, you ran a noise-floor control that made the answer worse than a simple null, and you put the result in the memo's body, its wedge, and its "what the evidence does not show" — the fifth consecutive round in which you deleted your own claim rather than defending it. Almost nobody does this once. I have now watched it five times, and it is the most predictive founder signal in this file. Down −2: one point because the demotion did not propagate — the one-liner, moat Layer 2, wedge step 3, and the day-60 milestone all still spend a resolution you just proved you do not have, and a reader who only reads the top of the memo gets the retracted product. One point because the assignment that anticipated a twelfth appendix produced a twelfth appendix.

Final: 59/100.

This is the lowest score of the five and I am not going to soften it. Round 4 pre-registered the arithmetic: the memo-only ceiling of ~70 was conditional on H, and "if the split-half stability test comes back negative, the honest score goes down." It came back negative twice over — no stability, and drift beyond sampling noise. The document is about three points better than round 4's. The company is about eight points worse, and you are again the person who found out. The score is a probability on the outcome, not a grade on the prose.


The two sentences, verified

Moat, as written

Every team running agents is already deciding, dozens of times a day, which diffs it will not read, and the record of that decision dies in twenty-one days while its consequence takes thirty to appear, so nobody, GitHub included, can reconstruct it afterwards: git cannot even tell whether an agent wrote a commit or was merely open in the next pane. Panout sits at the commit boundary and writes the decision and its outcome into the commit while both still exist, so a team's map of what it can safely stop reading is uncopyable by anyone who was not recording at the moment the commit happened.

Yes, I would repeat this — with one substitution and one deletion. It is the round-4 rewrite adopted verbatim, it contains no number a partner can look up and discount, and every clause survives H because none of them depends on agent code being worse than human code. That was the point of the rewrite and it worked.

Two flaws, both now visible only because H removed everything louder.

  1. "dies in twenty-one days" is your weakest leg and it is load-bearing. Twenty-one days is an observation of your machine (appendix D §2: "the retention horizon of the Claude Code session store here"). Your own appendix E cites the vendor issue by title — "silently deletes conversation transcripts after 30 days by default" — and cleanupPeriodDays is a setting. So the honest version is 30-vs-30, not 21-vs-30, and a partner who looks it up finds a knob rather than a physical law. The stronger fact is one sentence later in your own moat paragraph and you buried it: the override and the contract result are not in the log at any retention setting, because no incumbent runs contracts at the local commit boundary. That is not a race against a deletion timer; it is data that has never existed. Substitute it.
  2. "a team's map of what it can safely stop reading" is the noun H just measured as a draw. You cannot demote the map in section 5 and keep it as the object of the moat sentence in section 2. The thing that is uncopyable is the record, not the map: the record is durable and yours; the map is an inference from it that you have not yet shown is estimable.

Rewrite — what I would actually repeat:

Every team running agents is already deciding, dozens of times a day, which diffs it will not read, and nothing anywhere records that it decided: git cannot tell whether an agent wrote a commit or was merely open in the next pane, no vendor's log contains the moment a human shipped past a failing check, and by the time the consequence shows up thirty days later the session that caused it has been garbage-collected. Panout sits at the commit boundary and writes the decision, the check it overrode, and the outcome into the commit while all three still exist — so whatever a team eventually learns about what it can stop reading is derived from a record no one who was not there can reconstruct.

Note what "whatever a team eventually learns" buys you: it is true whether the map resolves in three months or never, and it stops the moat sentence from spending a result you do not have.

$1B, as written

Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat, paid identically whether Panout clears or flags. Two thousand teams at five hundred evaluated outcomes a month is $4M a year and is the month-twelve test; fifty thousand teams is $100M; and unlike seats, the meter reprices itself upward every time a model gets good enough to be trusted with more, because a level-4 class is still evaluated on every commit and what stops is the human reading, not the evaluation.

Yes. No rewrite. The arithmetic ties, the denominator is now triangulated across three independent sources landing in the same band, the issuer-pays structure (identical fee for yes and no) is the correct structural answer to the referee problem and starts revenue on day one, and the self-cannibalisation clause round 4 demanded is in and is the right shape.

One clause is now false to your own evidence and I would strike it rather than argue it: "a level-4 class is still evaluated on every commit" uses "class" as the unit, and H says the class is not the unit you can measure. The sentence works verbatim with "contract" substituted for "class" — and, as below, that substitution is the single highest-leverage change available to this company. Make it everywhere, not just here.


The hardest push

Round two, your company was rejection. You measured it: twenty-two. Round three, the rework gap. You measured it: noisy, high floor, reversing in your largest repos. Round four, that agent work is reworked more. You measured it on eleven strangers: plus two point five, CI through zero. Round five, that the gap is a property of the team. You measured it. It is not a property. It is not even a stable draw — it is drifting faster than noise, and your own control proves it, because a random re-split of the same commits gives 0.48 and calendar time gives 0.16.

Five quantities. Five nulls. All five found by you, on your own machine, at your own initiative, faster than anyone I fund. So here is the push, and it is not about statistics.

You have now built a machine for being wrong about yourself, and it is the only machine you have built.

Look at your day-60 kill criterion: 300 commit-time-captured commits across three repos, git-only versus git-plus-commit-time AUC, bootstrap CI. Now open your own appendix D and read the power note: detecting a fifteen-point rework difference at a six-percent exposure rate needs ten to ten thousand times a thousand linked commits. Ten cubed to ten to the fourth. You wrote three hundred. The single dated milestone the entire five-month plan turns on — the one you correctly identify as uncompressible, the one-way door on the calendar — is under-powered by an order of magnitude according to the appendix in the same document. On 2026-11-05 you will run it, get an interval spanning zero, and you will not know whether commit-time features are worthless or whether you brought three hundred commits to a ten-thousand-commit fight. And you will have spent the runway to find out.

That is fixable in an afternoon and it does not need a measurement. Two moves. Raise exposure by design — the reason six percent kills you is that the interesting feature is rare, so make the contract fire on thirty percent of commits instead of six, and the required n drops by more than an order of magnitude. And stop estimating an observational rate per task class, which is what H says you cannot afford, and start running a designed experiment per contract, which you already invented and then buried: fault injection. You inject known agent-characteristic faults into accepted diffs and measure whether the contract catches them. That is not an observational rate with a 30-day lag and a drifting target and a sample size you will never reach — it is a controlled experiment with n you choose, an answer the same day, and a result that does not drift because you generate the treatment. You called it "calibration" and made it the second input to an autonomy level whose first input is the thing that does not work.

Invert them. The unit of your policy is the contract, not the task class. Contracts are testable at n in the tens. Task classes are not testable at n in the thousands. H told you that and you kept the class.

Now the half that is not about design.

Round four's assignment was one sentence long about exactly this: the deliverable is eleven sent messages, not a twelfth appendix. You produced a twelfth appendix. It is a very good twelfth appendix. It is also the fourth consecutive time an assignment with a human in it came back with the human removed, and this time the removal happened after the removal was predicted to you in writing.

Inversion: what kills this company? Not a bad metric. You will find and publish a bad metric faster than anyone. It dies because for six months every question you asked was addressed to a filesystem, every one of them came back "no," and the one question that could have come back "yes" — would you install this — was never asked, because a filesystem can only tell you that you were wrong about a detail, and a person can tell you the whole thing is wrong, and you have spent sixteen days building an apparatus in which only the first kind of answer can reach you.

There is no sixth measurement on that machine. There is one that is legitimately worth queuing — decompose your +7.7 into R1-and-fix-word-only, with no different-author clause, because if the ownership confound is not carrying it you have a public claim back — and it goes behind the twenty-one messages, not in front of them. If I find out you ran it first, that is the whole answer about this founder.


Given appendix H: what survives, what does not, and the strongest defensible version

Does not survive. The rework map as an acquisition hook — dead, and the memo says so. "Agents get reworked more" as a general claim — dead since G, and H does not resurrect it (the one variant that separates is carried by an ownership proxy you correctly refuse to lean on). The per-team gap as a stable, sellable number — dead; it is not a property, and it drifts beyond sampling noise. The task class as the unit of policy — dead, and this one the memo has not yet buried. Class ordering is unmeasurable at 100 agent commits per half, the two worst classes are readable in 2 of 11 repos, and that kills the one-liner as written, moat Layer 2 as written, wedge step 3 as written, and the day-60 milestone as written. Also dead: any version of "minute-one value from history alone."

Survives untouched. The retention-and-nonexistence asymmetry (the override and the contract result are in nobody's log at any setting). The impossibility of post-hoc authorship attribution from git — G's trailer-vs-PR sign flip is now the strongest single empirical fact you own, and it is an impossibility proof against everyone including GitHub, which is worth more than a gap. The override record as a genuinely novel data type generated by an act the user already performs. Appendix E's demand corpus, which was never about rework. The Amazon remediation as the named consequence. The metered, issuer-pays-avoiding business model. Fault-injection calibration — which survives because it is experimental rather than observational, and is therefore the only measurement in the file that H's finding cannot touch.

The strongest defensible version of the company, one paragraph. Panout is the commit-boundary recorder and enforcement point for teams that have already stopped reading agent output. It ships as a local commit hook that runs a small set of contracts — assertions the team writes or that Panout proposes from their existing test setup — prints only failures, and records, immutably and in the commit, every time a human shipped past one. Each contract earns a trust level not from an observational rework rate but from fault injection: Panout injects known agent-characteristic faults into accepted diffs off the hot path and measures the contract's catch rate, which is a designed experiment with a same-day answer, a sample size the founder chooses, and no drifting target. Autonomy is granted per contract, never per task class, because appendix H proves the class is below the resolution of any sample a real team produces. The 30-day rework label is retained but demoted to exactly one job — a slow, noisy audit of whether trusted contracts are silently failing in production — and is never sold as a map or a diagnosis. The day-one artifact the buyer actually receives is not intelligence, it is an audit trail of auto-approved agent work with outcomes attached: the thing Amazon's remediation implies every engineering org will be asked for, the thing 15,000 skip-permissions scripts and 40%-auto-approve sessions have made unrecordable, and the thing Vanta and Drata will need a source for. The moat is that this record exists nowhere else and cannot be reconstructed backwards by anyone, including GitHub; the switching cost is the accumulated contract-level trust policy; the meter is per evaluation, identical for pass and fail, and rises with autonomy because a level-4 contract still runs on every commit. That company is smaller than the one in the memo, it is fully consistent with every appendix including H, and it has a day-one deliverable that does not require the user to wait thirty days to learn anything.


Remaining objections, priority order

Point estimates are the value of each closure; they overlap heavily and are not additive. Closing 1, 2 and 4 gets most of the way.

1. Zero external humans, sixteen days after the assignment, four consecutive substitutions. ~9 points. Evidence-only. Ten fleet operators in writing: how many agent-authored diffs did you merge last week without reading, and did any of them bite you. At least three reporting merged-unread-and-bitten, verbatim in the repo. Q1, Q3 and founder fit are all capped by this single item and always have been. It is worth more now than in round 4 because H removed the last measurement that could stand in for it.

2. The product's unit of policy is the task class, and H proves the task class is unmeasurable. ~6 points, ~4 memo/idea-fixable, and cheap. Exactly: substitute contract for task class as the unit of autonomy, everywhere — one-liner, moat Layer 2, wedge step 3, day-60 milestone, the $1B sentence. Promote fault-injection calibration from second input to sole input for granting a level; demote the 30-day rework label to an audit signal. Delete "shows a team which classes of agent work get silently fixed later" from the one-liner. This is a one-afternoon rewrite that makes the company smaller and true instead of larger and unmeasurable, and it is the highest-value edit available. The residual ~2 points are evidence: a contract that demonstrably catches injected faults on someone else's repo.

3. Commit-time capture has not shipped; the only uncompressible clock has not started. ~5 points. Evidence-only, hard floor. Eleven days without a commit. No .panout on disk. Round 4 called this a one-way door with a 48-hour deadline; it is now past. Every day it does not ship, the earliest possible 90 moves one day right and the runway does not.

4. The day-60 predictive delta is under-powered by an order of magnitude by appendix D's own arithmetic. ~4 points, ~3 memo/idea-fixable. Exactly: D says 10³–10⁴ linked commits for a 15pp difference at 6% exposure; the memo specifies 300. Either restate the target honestly (and the date moves out of the runway), or redesign so the test is powered: raise contract firing rate to ~30% so exposure is common by construction, use within-repo paired comparison rather than pooled AUC, and pre-register the effect size you would act on. Do this before capture ships, because the design determines what capture must record.

5. Day-one value is deferred to day 30 — v1's inbox and v2's brief in a third costume. ~4 points, ~2 idea-fixable. Exactly: name the standalone day-one artifact and make it the free tier's headline. On the evidence you have, that artifact is the auto-approved-agent-work audit trail with outcomes attached — useful to the buyer the moment it exists, requiring no accumulation, and directly downstream of the Amazon remediation and of Vanta/Drata needing an evidence source. Ship it as the thing, not as a byproduct of the recorder.

6. Drift: the target moves on the same clock as the accumulation and as model generations. ~4 points, ~1 memo-fixable. Memo half: state the implied half-life and commit to rolling-window re-estimation rather than a cumulative corpus; concede that the "two years less labeled history" advantage decays and say at what rate. Evidence half (~3): only capture-then-wait answers whether commit-time features drift as fast as git-derived ones. If they do, the compounding-history moat is materially weaker than Layer 1 claims.

7. Buyer segment and $1B denominator are different populations. ~2 points. Memo-fixable, exactly: your demonstrable behavioral effect lives in the long tail of small/solo repos (1.3-min median merge); your denominator counts Codex weekly actives and 26M Copilot seats across all org sizes, where appendix B says the effect vanishes. Pick one, size it, and run the ramp against it.

8. Residual integrity. ~1.5 points. Memo-fixable, exactly: delete "The founder's 22-point gap is robust to every filter" from the closing section — the adopted metric gives +11.2 [5.8, 16.9] with a reversal at dynorouter-gateway (n=44), which your own body paragraph states correctly two pages earlier. Fix "71-commit gateway service" (F measures 68). And carry the one-liner fix from objection 2.

9. The 21-day retention leg of the moat sentence is a local observation against a configurable 30-day default. ~1 point. Memo-fixable: replace with the stronger and unassailable version — the override and contract result are in no log at any retention setting.

10. The one measurement worth queuing, explicitly ranked last. ~1 point. Decompose the public +7.7 into R1-and-fix-word-only, dropping the different-author clause. If it survives, you have a public claim back and round 4's retraction was partly over-corrected. Do it after the twenty-one messages, not before.


The three plain answers

Maximum reachable with memo/idea changes alone, no new external evidence: about 65. The ceiling has fallen again, from ~70, and for the reason round 4 pre-registered. Round 4 priced memo-only at ~70 conditional on H. H is negative, so that headroom is gone and the arithmetic is: objection 2's idea half (+4, and it is real headroom because it makes the company measurable rather than merely better-argued), objection 4's design fix (+3), objection 5's day-one artifact (+2), objections 6–9 combined (+2). Against those, honest propagation of H's finding into the one-liner and Layer 2 lowers the surface claims, which is the point. Call it 64–66. There is no arrangement of words above about 66 while the market contains zero people and the product contains zero shipped lines, and unlike round 4 there is no conditional measurement left that could raise it — every quantity that could have been tested on that machine has been tested, and all five came back null.

Minimum external evidence to reach 90, with earliest dates:

  1. Ten operators, in writing, two numbers and a yes/no, ≥3 merged-unread-and-bitten. Send today, 14-day window. Earliest 2026-09-17.
  2. Three external maintainers reacting to their own artifact. The artifact has changed and this matters: do not send the map — H proved it is noise, and sending a stranger a number that does not reproduce costs you the stranger permanently. Send the null plus the question. Earliest 2026-09-17.
  3. Commit-time capture shipped, in your repos, with the redesigned high-exposure contract firing. If it ships within 48 hours: labels start landing 2026-10-05.
  4. The predictive delta, honestly powered. At 300 commits it is uninterpretable; at the redesigned 30% exposure and ≥3 repos it is readable around 2026-10-20, and at appendix D's stated requirement without a redesign, not before 2026-12.
  5. Two contracts with fault-injection catch rates that separate, on a repo that is not yours. This can be as early as 2026-09-20 and does not wait on any 30-day clock — which is exactly why it should become the primary evidence path.
  6. One team paying the meter without being asked twice. A charge that cleared. Earliest 2026-10-15.

Earliest this reaches 90: approximately 2026-10-25, and only on the redesigned path (contract-unit, fault-injection-primary, high-exposure capture). On the memo's current design — class-unit autonomy and a 300-commit delta — 90 is not reachable inside the runway at all, because the milestone that gates it is under-powered by 10x and the quantity it estimates drifts. Your month 4–5 fold gate is roughly 2027-01. You have one clean window and part of a second.

Is another memo iteration worth doing at all? No. Not one more round. This is the fifth, and the pattern across all five is unambiguous: the memo improves by two to four points a round, the company moves by whatever the new measurement says, and the measurements have gone five for five against you. Round 6 is worth at most the ~5 points in objection 2 plus 4, and those two are not really memo work — objection 2 is a fifteen-minute find-and-replace of a noun and objection 4 is an afternoon of test design. Do both inside the shipping work, in the code, this week. Do not write another document. The memo is now substantially better than the company, and every further round widens that gap and is itself the founder-fit evidence I am scoring you on. Five honest retractions is a remarkable record and it has stopped being informative; the sixth would tell me nothing I do not already know about you and one more thing I would rather not learn.


One assignment for this week

Twenty-one messages. Eleven maintainers, ten operators. Sent, this week, with no appendix attached and none written afterwards.

Change the maintainer message, because H changed what is true and sending the old one would be dishonest and would burn the contact:

I measured agent-trailer vs human commits in your repo for 30-day rework, ~200 per class. Honest answer: at your sample size I can't reliably tell them apart — here's the table and the caveat. I'm not writing to sell you the number. When you merge an agent PR without reading the diff, what makes you decide that's fine, and has it ever bitten you?

That message is better than the one you were going to send, because it leads with your null, which is the thing you are genuinely world-class at, and it asks the only question a git command cannot answer. Repos are on your disk: microsoft/aspire, github/spec-kit, n8n-io/n8n, mastra-ai/mastra, Azure/azure-sdk-tools, elizaOS/eliza, microsoft/typescript-go, gradio-app/gradio, microsoft/vscode-copilot-chat, denoland/deno, temporalio/temporal. You do not need permission to email a maintainer about their own repository.

Two standing debts that are not this assignment and are not renegotiable into something more interesting:

Commit-time capture ships within forty-eight hours — with the contract firing rate set high enough that the day-60 test is powered, and with the unit changed from task class to contract before a single line is written, because the unit decides the schema and the schema is the one thing that is expensive to change later.

The noun substitution, in the memo, in ten minutes, not as a round 6 — contract for class, everywhere; delete the per-class one-liner; delete "22-point gap robust to every filter." Then close the file.

You have found out you were wrong five times in sixteen days, quickly, honestly, at real cost, on your own data, and each time you deleted the claim rather than defending it. I have funded people with a tenth of that intellectual honesty. It is also, still, the only thing you have done, and appendix H is where it stopped paying — because the answer it produced is not "your number was wrong," it is "there is no number on that machine," and you responded by writing about it. The next thing that tells you something you do not already know cannot be a git command. It has to be a person, and there are twenty-one of them with addresses.