Round 6 judgment (2026-09-04): appendices I and J
Score: 61/100 (round 5: 59). Independent judge, no stake. The score barely moves, and that is the finding: for the first time in six rounds a non-writing assignment item was executed on time and in code, and its output plus a day of market reading took back nearly all the credit. I verified the machine — panout.py (1,281 lines) has hooks live in three repos, and the ledgers hold 25 evaluations with 12 overrides, 11 of them size-guard. The first override corpus in existence is 92% "you committed over 400 lines," the contract appendix I proves restates size.
The round-5 list, literally
- Ten operators in writing — untouched. Seventeen days. Zero.
- Three maintainers reacting to their own artifact — untouched. Six strangers' repos measured; no owner contacted.
- Capture shipped in 48h at high exposure — closed. Committed 02:47Z 2026-09-04, inside the deadline; contract-not-class fixed in the schema.
- Predictive delta honestly powered — partly closed, answered negatively, not the same test. I was wrong that it needed 30 days: contracts are pure functions of the diff, so history supplies them, and power is fixed (5,786 commits, inside appendix D's 10³–10⁴ band; exposure 26–43%). But it excludes the override and session shape, which was the day-60 test. Disconfirmed for the defaults, untested for the moat feature.
- Two contracts separating under injection off-machine — closed as written, and my item was badly specified.
tests-touchcatching drop-test is near-tautological: the fault is the assertion's exact negation, so it measures wiring, not detection. The informative result is the negative one — assertion-delete caught by nothing, anywhere. - One team paying — untouched.
Two closed, one partial, three untouched — and the untouched three contain the humans.
The size confound: which reading
I take the founder's reading — the product told the truth about its defaults on day one — and it costs more than he thinks. It is right on mechanism: "autonomy per contract from measured catch rate" requires the ability to say "this contract carries no information," and the product said it in minutes on unseen repos. First Panout capability in six rounds demonstrated rather than argued. But what the working design demonstrated is that the inventory is empty — two of three predictive defaults carry nothing once size is held fixed, the third is size — so the company becomes "we are the instrument you calibrate your own assertions with": smaller, truer, and now dependent on customers writing contracts worth calibrating, for which there is no evidence.
openai/codex, zero trailers in 4,590 commits
Best fact in either appendix. G's trailer-vs-PR flip was a statistical inconsistency; this is categorical absence at the scale of the vendor best placed to compete, in the repo built with the product that writes the trailers: OpenAI cannot say which commits Codex wrote, because its merge policy destroyed the evidence. Two limits: it strips attribution from git, not GitHub's PR database, and it is configurable practice, like the 21-day retention leg. Strengthens the pitch a lot, the moat a little; lead the maintainer email with it.
Does the backtest undermine the moat? Yes, more than the memo admits
The moat sentence records three things: the decision, the check overridden, the outcome. Appendix I proves the second and third are computable backwards by anyone in minutes, from history git keeps forever — which falsifies Layer 1's "a team adopting in 2028 starts with two years less labeled history." They start with two years minus one bit.
What only commit-time capture provides: (1) the override bit — a human saw a named failure and shipped anyway; (2) attribution surviving squash merge, because the trailer is in the commit; (3) session shape and model, pruned in weeks; (4) contracts that are not pure functions of the diff — unclaimed, and the important one, because a command contract that runs the suite, or assertion-delete detection needing semantic analysis, cannot be replayed over 3,000 historical trees. The backtest undermines the moat exactly to the extent the contracts are trivial.
Appendix K: narrows, does not break, costs more elsewhere
The moat survives narrowed — nobody records, in the customer's own repo, the decision to commit past a named failing contract — and K's item 3 is the best structural argument in the file: GitHub captures the override moment and routes it to its own product team, not the customer.
The damage is elsewhere. Demand: appendix E's corpus was demand for not being asked to approve, which four vendors now give free; demand for the record itself is 0–6 reactions across seven venues, four closed "not planned" — including the issue asking for Panout's exact feature. Buyer pull: the compliance rescue in my own round-5 "strongest defensible version" — Vanta and Drata needing an evidence source — is disconfirmed for 2026. I was wrong; the founder found out; sixth time. Competition: Arnica sells a signed per-PR review record as compliance evidence, with Endor, Snyk and Kiro Crew adjacent. Unpriced: Anthropic's own figure that human review caught 13.6% of dangerous commands against auto mode's 89%. If the machine gates better than the human, the valuable artifact is its decision — which its vendor already logs.
Six forcing questions
Q1 demand down (thin, mostly for the free thing). Q2 status quo down (B's cost is size, computable free). Q3 desperate human: the founder, sole author of the world's only twelve overrides. Q4 wedge up (the backtest print needs no accumulation). Q5: zero external users; the surprise came from a filesystem again. Q6 future-fit up, on the agent-on-agent merge gate — the override is the last human signal at the boundary.
$1B? The category is: every team running agents must decide what it stops reading, that decision meters per unit of work rather than per seat, and 38% of commits across six public repos would carry an override record on day one. This company is not one yet, because the whole irreproducible asset is one bit per exposed commit, that bit's information content is untested, and the people who would buy it have upvoted it six times.
Score, ceiling, next action
61/100. Up: capture shipped inside the deadline and running, the first executed non-writing item in six rounds (+4); contract-as-unit in the schema, not the prose (+2); the backtest method (+2); item 5 (+1). Down: defaults carry no risk signal (−2); two of the moat's three clauses are backwards-computable (−2); K thins demand, removes the compliance pull, names an identically positioned competitor (−3); sixth substitution of measurement for a person (−1). Net +2: uncertainty much lower, company not much better.
Ceiling with no further external evidence: about 64, down from 65. Every quantity on that machine is measured and K exhausted the reachable market facts: three points remain in code, none in evidence.
Next single action: yes, still go talk to ten operators. One question, in writing, today: how many agent-authored diffs did you merge last week without reading, and did any bite you? Stop doing this: stop reading strangers' repositories instead of writing to their owners. Yesterday six were measured and published — mine among them — and not one owner was emailed. The data has become the avoidance mechanism, and the next thing that can tell you something you do not know cannot be a git command.