Panout review packOverviewFinal memoMVPLive outputsJudge roundsAppendicesDecisions

Source: docs/research/2026-09-03-defensibility/rounds/r4-memo.md

Panout defensibility, round 4 (2026-09-03): The rework layer for agent labor

One idea, fourth iteration. Scores so far: 62 and 72 (round 1), 66 (round 2), 67 (round 3). The round 3 judge stated the ceiling: memo and idea changes alone reach 74 to 76; 90 requires ten external operators in writing, three external repos with a stranger's reaction to the rework map, one paying team, and a predictive measurement that cannot exist before 2026-10-05 because it needs 30 days of commit-time capture. This round applies every reachable fix and adds two measurements (appendices F and G). It does not claim to have closed what only the founder can close.

What the founder owes before anyone reads further

The 2026-09-05 gate result, reported today rather than on the day. STRATEGY.md set a behavioral gate: by 2026-09-05 the v2 ledger must have recorded at least five real sessions and the brief must have recovered work, prevented overlap, or changed a next action. Measured 2026-09-03: no .panout directory exists on this machine, no ledger file exists anywhere on disk, and the repo has no commits since 2026-08-23. The dogfood never started. The gate fails as written, two days early, and this memo does not reset it. The reason it never started is the reason v1 died: the product required the founder to run a command and read a brief, and the founder did neither. The wedge below is designed so that the failure mode cannot recur, because it instruments commits the founder is already making (5,294 founder-authored non-merge commits in 180 days across the nine repos the founder controls, by git log) and asks for nothing else.

The ten operators, one question, in writing. Assigned 2026-08-22, still outstanding. No memo replaces it. Everything below is priced by the answer.

The two sentences (rewritten after round 2)

Moat. In this founder's own repos, agent commits get silently rewritten within thirty days about twenty points more often than the commits he writes by hand; across eleven public repos the same gap runs from minus twenty to plus forty-three and averages near zero, so which agent work a team can trust is a property of the team, not of the models, and nothing today connects the outcome back to the session that produced it or the judgment that let it through: git cannot even tell who wrote the commit, the session logs holding the cause are deleted in twenty-one days, and the consequence takes thirty to appear. Panout stands at the commit boundary writing the cause and the override into the commit before they expire, so a team's map of which task classes it can safely stop reading is uncopyable by anyone who was not recording when the commit happened.

$1B. Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat, paid identically whether Panout clears or flags. Two thousand teams at five hundred evaluated outcomes a month is $4M a year and is the month-twelve test; fifty thousand teams is $100M; and unlike seats, the meter reprices itself upward every time a model gets good enough to be trusted with more.

One-liner

Panout closes the rework loop on agent work. It shows a team which classes of agent work get silently fixed later and which do not, derives contracts and earned autonomy from that record, and is paid per outcome evaluated, the same whether it says yes or no.

The constraint set (unchanged, all respected)

Does not control agents. Never in the live agent loop (evaluates at commit, PR, and session-end boundaries). No login, repo and terminal only. No duplicate runs. Raw content stays local; the ledger holds metadata, hashes, contract results, and labels.

The desperate human, named as precisely as the data allows

The founder, measured (git-only backtest, 2026-09-03, appendix A; session census, appendix C). Across nine repos the founder builds alone, the agent share of commits over 180 days runs from 0% to 63%, concentrated in the newest and smallest repos; in one team repo where the founder is a minority committer (checked out twice on this machine, so appendix A counts it twice) it is 85%. Agent-majority is a property of greenfield, and greenfield is where a 48-hour wedge installs. Restricted to commits the founder personally authored, agent commits are reworked within 30 days at 60% versus 38% for human commits. The pooled figure (65% versus 49%) is a mix effect, double-counts one team repo checked out twice, and includes bot commits on the human side; it is reported in the appendix, not relied on here (appendix F has the corrected pooled gap, 22 points with bots excluded). One in six reworked agent commits is reworked by the very next commit. Explicit reverts are 0.16% of commits: correction shows up as rework, not revert, so it is invisible to anyone counting reverts. 99% of agent commits land on the default branch without a merge commit (a weak proxy, since squash merges look identical).

The census that changed this memo (appendix C). 262 Claude Code sessions and 95 Codex sessions in 21 days (the machine's logs begin 2026-08-13; 30-day and 180-day windows are identical for logs). Explicit human rejections of an agent action: 22 in 15,642 tool calls (0.14%). Interrupts: 25. Codex aborts: 14. Stashes machine-wide: zero. Sessions containing any hard negative: 8.4%. Of the four rejections with a recoverable file path, all four were Markdown edits (18 of 22 had no recoverable path). Most sessions run with permission prompts disabled, which removes the moment at which rejection would occur (appendix D). Appendix D, using looser matching, counts 53 rejections and 31 interrupts; appendix C's anchored matching (22 and 25) is the count relied on here because the looser patterns matched marker strings echoed into tool output by earlier measurement runs. Meanwhile 60% of the founder's agent commits are reworked within 30 days; the two rates have different denominators (per tool call, per commit) and are not divided into a ratio. Round 2 called rejections the asset; the data says humans have already stopped rejecting in session and correct afterward, silently, in rework nobody measures. That is the status quo in one number, and it is why the moat, the wedge, and the day-one value are all restated below around rework rather than rejection.

The cohort, observed in public (appendix B, measured 2026-09-03 against the GitHub API). 1.5 million merged PRs from Copilot's coding agent, 5.8 million merged PRs on Codex branches, 694 thousand on Cursor branches, 346 thousand carrying a Claude co-author trailer. In a 150-PR sample of merged Copilot agent PRs, the median time from open to merge is 1.3 minutes and 70% merge within 10 minutes, at a mean of 397 added lines; the control sample of non-agent PRs in popular repos merges at a 25-minute median. Zero-review merge rates are not distinctive (49% for the control), so the claim is not "agents get merged without review", it is "agent output is merged faster than it can be read, into a review culture that was already thin." Sonar's 1,100-developer survey: 96% believe AI code is not reliably correct and 48% say they always check it. Peter Steinberger, publicly: "I don't read much code anymore."

The named consequence. Amazon, March 2026: a "trend of incidents" with "high blast radius" tied to "Gen-AI assisted changes", a roughly six-hour retail outage, and a remediation that is the whole problem in one sentence: every AI-assisted change must now be approved by a senior engineer before deployment, which Amazon called "controlled friction." Amazon disputes the framing and has said the one AI-involved root cause was an engineer acting on stale internal advice; the dispute is itself the point, because the remediation stands either way. That is the status quo's only answer to agent-caused incidents: more human reading, applied to everything, forever. The person who needs Panout is the engineering leader who has just been told to read everything and knows it cannot scale, and the fleet operator one incident away from the same order. DORA 2025 (about 5,000 respondents) finds AI adoption at 90% and continuing to increase change failure and rework; GitClear measures code churn rising from a 3.3% pre-AI baseline to 7.1% in 2025.

Demand in writing, unprompted (appendix E, counts read from the GitHub API on 2026-09-03). The request to stop approving agent actions one at a time has 249 upvotes on a single VS Code issue and 59 on Codex; a Claude Code meta-issue titled "Permissions matching is fundamentally broken, 30+ open issues, community building workarounds" has 78. GitHub code search finds on the order of 15,000 shell scripts invoking the skip-permissions flag and 2,800 settings files enabling bypass mode, plus auto-approve hooks shipped to customers by Railway and Render. Anthropic's own February 2026 measurement: full auto-approve use rises from about 20% of sessions for new users to over 40% by 750 sessions, and Anthropic concludes that approving every action "will create friction without necessarily producing safety benefits." Retention is a live grievance: issues about Claude Code silently deleting transcripts after 30 days carry 31 and 23 upvotes. Per-unit payment for review of agent output is established: Cursor Bugbot at $1 to $1.50 per run, Greptile at $1 per review, CodeRabbit at $0.25 per file beyond quota and an estimated $40M ARR growing seven-fold (secondary source). What nobody asks for in words is "tell me what not to read"; people ask to read better (180 upvotes for a diff-review UI). The behavioral record says they gave up reading instead. Panout does not compete for the auto-approve ask; Claude Code Auto mode, VS Code terminal auto-approve, and Codex execpolicy own it, free. The 15,000 scripts and the 40% of sessions in full auto-approve are the size of the population that has already made the trade, and none of them can see what it cost. That is why the wedge leads with the rework bill, not with a reading tool: the ask that exists is "stop making me approve", the cost that exists is the rework gap, and Panout connects the two. The person who buys is the one who already took auto-approve and now owns the consequence.

What gets them promoted or fired. Shipping volume without incidents. The lever they control is how much agent output they read, and today that lever has two positions: everything, or nothing.

Status quo and its cost

Position A, read everything: caps agent throughput at human reading speed. The team repo in the backtest merges about 1,200 agent commits a month, 85% of all its commits; nobody is reading that. Position B, read nothing: git diff --stat, skim, merge. Cost is the measured rework gap: for commits the founder personally authored, agent work is reworked within 30 days 22 points more often than hand-written work (60% versus 38%); pooled across all repos the gap is 16 points; per repo it runs from minus 20 to plus 35 points and reverses sign in the large team repo and in three repos with one to ten agent commits (appendix A). The correction lands a median 17 hours later. Because the line-identity metric has a high floor (human work is reworked 38 to 51%) and could be counting reformatting, appendix F recomputed the same 2,859 commits under stricter definitions. The gap survives every one of them. Whitespace and move insensitivity changes it by under two points, so reformatting was not what the metric was counting. Requiring the rework to look corrective (a follow-up with a fix-word subject or touching a test) cuts the human floor to 34% on the founder's commits and 29% pooled while the agent rate barely moves, so the founder-only gap is 20 points with a bootstrap interval of 14 to 26, and 11 points with an interval of 6 to 17 under the strictest variant that drops the test-path clause. The team repo's reversal was an artifact: 36 bot release commits classed as human that overwrite the same version strings on every release and score 100% reworked; under the corrective definition that repo flips to agent work reworked 32 points more, and with bots excluded the line-identity gap there is under two points, meaning gone. Two corrections to appendix A surfaced in the process: the two team repos are one git history checked out twice, so pooled figures there double-count, and 28% of "human" commits pooled are bots (dependabot, release republishing, a research pipeline); with bots excluded the pooled gap is 22 points. After the filters, no sign reversal remains that is backed by more than ten agent commits. Industry-wide, churn is doubling (GitClear) and instability rises with adoption (DORA 2025), and Veracode finds models pick the insecure implementation in 45% of tasks with no improvement across model generations; but appendix G shows the agent-versus-human rework gap itself is team-specific and near zero pooled across public repos, so the cost of Position B is not a universal tax, it is an unknown that differs per team and per task class, and that unknown is the product. Position C, second-agent review (gstack /review, Cursor Bugbot at $1 to $1.50 per run, Copilot review, Codex review): a reviewer with no memory of the team's bar and no record of whether it has ever caught anything. It produces more text to read, which is the problem restated. Bugbot's shift to per-review pricing in June 2026 is also the market telling us that per-unit pricing at this boundary is accepted.

Verified against vendor documentation on 2026-09-03 (appendices B and E): across Copilot code review, GitHub rulesets and merge queue, Linear Coding Sessions and Guided Reviews, Cursor Bugbot, Claude Code Auto mode and hooks, VS Code auto-approve, Codex execpolicy and review, none learns a per-team bar from that team's own history, none enforces autonomy per task class, and none joins what was approved to what was later reworked. The closest approximations are hand-written allowlists, a stateless per-call classifier, and effort dials. The narrow gap claim, and the only one the evidence supports, is learned from your own history and recorded durably.

None of these learn. Every day of agent work today produces zero durable information about what this team accepts.

The wedge: value in the first minute, zero new behavior, 48 hours to build

Panout's v1 died because it required the founder to read an inbox; v2 never started because it required the founder to run a collector and read a brief. This wedge produces its first result from data that already exists and then instruments behavior that already happens.

  1. Minute one: the rework map. panout init runs the appendix A measurement on the repo's own history and prints one table: for each task class (derived from paths and commit shape), the share of agent commits reworked within 30 days, the median hours to rework, and the same numbers for human commits. The map uses the corrective definition from appendix F (rework that is whitespace and move insensitive and whose follow-up has a fix-word subject or touches a test), with bot authors excluded, because that is the version that does not print backwards on any repo with a meaningful sample. On this founder's repos it says: agent commits reworked 54% versus 34%, one in six by the very next commit. No streak, no contract, no login.

What the map prints on strangers' repos (appendix G), stated before anyone else finds it. The same measurement was run on 11 public repositories with heavy agent traffic (aspire, spec-kit, n8n, mastra, azure-sdk-tools, eliza, typescript-go, gradio, vscode-copilot-chat, deno, temporal; 2,042 commits). Pooled, agent-trailer commits are reworked 2.5 points more than human commits, with a 95% interval from minus 1.8 to plus 6.9. Per repo the gap runs from minus 20 to plus 43: positive in five, negative in three, noise in three. Matching samples by calendar month moves the pooled gap to minus 1.3. Switching from trailer attribution to PR-level attribution flips the sign in all three repos where both could be computed, and in the repo with the largest gap the whole effect is one contributor. The founder's 22-point gap is real for the founder and is not a property of agent code in general. Three consequences follow and the memo accepts all of them. First, "agents get reworked more" is retracted as a general claim; the hook is not that agents are worse, it is that nobody knows which of their own task classes are, and the answer differs by team, which is the argument for a per-team calibrated map and the final argument against the global index idea from round 1. Second, the day-one map from git alone is weaker than round 3 claimed: it may print "no difference" for a stranger, and its per-class ranking is readable only where samples are large; the one pooled per-class pattern with any support is agent work reworked more in source and config and less in tests and docs, on repos whose individual gaps disagree. Third, git cannot attribute authorship: trailer and PR-level rules give opposite answers on identical commits, because a trailer means an agent was in the room, not that it shipped the change. That is the strongest evidence yet that the features have to be captured at commit time from the session that produced the commit; it is also a limit on any product that promises a map from history alone, Panout included. Whether a stranger finds their own map useful is now the open question the founder must answer with three external repos and their maintainers' reactions, and no measurement on this machine can substitute. This is the demo, the hook, and the reason to keep the tool installed.

  1. Commit time: capture the features. A commit-time step records, as metadata only, the task class, harness and model (from trailers and supported session surfaces), session shape, and the result of any contracts. Nothing is read from transcripts; the census in appendix C was produced with exactly this discipline.
  2. Contracts for the two worst classes. For the two task classes with the highest rework rate, panout init proposes acceptance contracts from the existing test setup and AGENTS.md; evaluation prints only failures. Committing after a printed failure is an override and is recorded as one. Acceptance is the commit.
  3. Thirty days later: the label attaches. Each captured commit is labeled reworked or not, and the map updates. Autonomy levels per task class are derived from it, with the catch-rate calibration below as the second input.
  4. One optional key turns a rework or override into a decision record in decisions/, summarized into AGENTS.md so every harness reads it at session start.

The fourteen-day test needs no new habit: does the founder keep committing with the capture on, and does the rework map change what gets a contract? The measurement in appendix D tried to test, on the same commits, whether session evidence predicts rework better than git alone. It could not be answered, and the reason is the product argument in miniature: harness logs on this machine are retained for 21 days, the rework label needs 30, and the two windows never overlap, so retrospective joining is impossible for anyone. Squash merges compound it: of 49 commit hashes agents printed in sessions, 69% never landed under that hash, and in the largest repo 0 of 40 did. Retroactive linkage recovers about 13% and the recoverable subset is the reviewed one. The features must be captured at commit time under a stable identifier written into the commit itself, or they are gone within three weeks. The first honest predictive delta therefore has a date: 30 days after capture starts, 2026-10-05 at the earliest if the wedge ships this week. Until then the claim that commit-time features add predictive power over git is a hypothesis with a measurement scheduled, not a result.

The moat, in three layers that GitHub cannot copy

Layer 1: the override record and the commit-time join. Concession first: task class (from paths and commit shape) and harness and model (from commit trailers) are recoverable from git by anyone, and appendix A classified 15,440 agent commits over 180 days using nothing else. GitHub can compute those. What is not in git, and not in any vendor's logs, is three things. First, the override: a human committed after a printed contract failure, on a specific diff, for a specific reason. That is a human judgment about agent work, it is generated by the act the human already performs, it is the negative class that in-session rejection turned out not to be, and no incumbent records it because no incumbent runs contracts at the local commit boundary. Second, the contract result itself, per commit, per class. Third, session shape from supported harness surfaces, which appendix D shows is pruned after about 21 days on this machine, nine days before the rework label can arrive. The join of those three to the 30-day label cannot be built retrospectively by anyone, GitHub included; a team that adopts in 2028 starts with two years less of its own labeled history than a team that adopted in 2026.

Layer 2: earned autonomy as the team's calibrated policy. Each task class carries a level from 0 (human reads everything) to 4 (auto-merge). Levels are computed from acceptance streaks, rejection history, and contract catch rate (see calibration below), and enforced at the merge boundary. This is the switching cost: leave Panout and you go back to reading everything, because the policy is derived from a history no other party holds. It also answers "what happens when GitHub ships AI review confidence": GitHub's confidence becomes one more evidence input to the contract, the way a CI check is. GitHub's signal is global and mechanical; the policy is the team's.

Layer 3: cross-team priors (the Radar effect). Each team's bar is private. What is shared, with consent and as aggregates, is the pattern library: which fault classes get rejected, which contract shapes catch them, and which task signatures earn autonomy fastest. A new team's contracts start warm. This is why one merchant's fraud history is worthless and Stripe Radar's is a business: the per-customer data is thin, the cross-customer prior is thick, and no customer can get the prior without joining. Neither a harness vendor (sees one agent) nor GitHub (sees no rejections) can build this prior.

The tension the moat has to survive, named. The market argument says models improve and teams let agents do more. A moat built on refusals thins on the same curve, and the census shows refusals are already near zero on this machine (appendix C, table 3: 21 days of history, half the negatives on one day, model comparison confounded with time, no trend claim possible). So the asset is not refusals. It is the join between commit-time features and the 30-day rework label, which does not thin as models improve: rework fell from 38% for human commits to a floor, not to zero, and the human rate itself is 38 to 51%. Better models shift which task classes are safe; they do not remove the need to know which ones.

Calibration, so the moat is measurable and the product is not a linter. Off the hot path, Panout injects a small library of agent-characteristic faults into a sample of accepted diffs and re-runs the contract, producing a catch rate per contract. Autonomy requires both a clean streak and a catch rate above threshold. Weak contracts cannot earn autonomy. This converts the round 1 day-45 kill criterion ("contracts are theater") from a risk into a metric that is measured weekly.

Incumbent response, and why the idea survives each

IncumbentWhat they shipWhy it does not close the seam
GitHubCopilot code review can now submit approving reviews that satisfy branch protection (public preview, appendix B), plus rulesets, merge queue, auto-mergeGitHub has stepped into the approval decision. It still sees survivors only, has no local session or rejection evidence, applies rules globally and mechanically, and is conflicted via Copilot. The defense is switching cost, not neutrality: a team that leaves Panout goes back to reading everything, because its autonomy policy is derived from a rejection history nobody else holds. Copilot's approval becomes one more evidence input to the contract.
LinearCoding Sessions, Guided Reviews, per-issue model choiceCloud-resident graph, sees only its own hosted sessions, bills AI credits for execution, so cannot judge its own agent or see local work. Panout syncs cleared outcomes back to the issue.
Harness auto-approval (Claude Code Auto mode, VS Code terminal auto-approve, Codex execpolicy, Gemini policy files)Per-call classified approval (Claude Code, a remote classifier) or static allowlists (the others)This is the closest shipped surface and the memo does not claim it is empty. All of it decides per action, in session, statelessly: none learns from the team's own acceptance or rework history, none is per task class, none leaves an acceptability record, and every high-vote issue against them is about being unavailable, over-permissive, or over-restrictive. Panout consumes their decisions as evidence and answers the question they cannot: which of the things you auto-approved got fixed later.
Anthropic, OpenAI, CursorIn-session review, auto modes, thumbs up/down, memory from correctionsEach sees one agent, in session, with no post-hoc outcome and no cross-vendor comparison. A rating of your own work is not a rating. Panout consumes their supported event surfaces.
Entire (entireio/cli; the public sweep in appendix B could not independently verify its feature set, so this row rests on the repo's 2026-08-23 research)Cross-agent transcript archive, checkpoints, judge, expertsPreserves everything; no acceptance labels, no enforcement; its value grows with raw retention, which conflicts with metadata-only. Optional evidence input.
Harness, Jellyfish, SwarmiaLeader dashboards on sessions, tokens, attributionRead-only reporting for managers; no enforcement, no rejections, no per-team bar.
Vanta, DrataAI-code control modules pulling GitHub evidenceTop-down compliance; they will need an evidence source for local agent work and Panout's receipts are that source. Channel, not competitor.

Business model: metered per outcome evaluated, paid the same for yes and no

Why this fixes the round 1 objection. Revenue tracks agent volume. A team of two humans and forty agents pays Linear $32 a month; the same team produces thousands of evaluable outcomes a month. Cursor's Bugbot already charges $1 to $1.50 per review run at this boundary and moved there from per-seat pricing in June 2026, so per-unit pricing is accepted by exactly this buyer. Prosumer conversion rates are irrelevant when the meter is on the agents.

$1B math, stated as a ramp with a measured denominator. The denominator is built from usage figures rather than one vendor's team count: Codex reports over 3 million weekly active developers, Copilot over 26 million users, Claude Code an estimated 4% of public commits (appendix B; the last two are secondary sources). Taking only the 3 million Codex weekly actives and a team size of six gives roughly 500,000 agent-using teams today, before Claude Code, Cursor, and Copilot are added. The ramp: 2,000 teams evaluating 500 outcomes a month at a third of a dollar is $4M a year and is the month-12 test of whether the meter works, at 0.4% penetration of that denominator; 10,000 teams is $20M at 2%; 50,000 teams is $100M at 10% of today's population, and the 2028 population is larger by whatever multiple agent adoption grows. Merged agent PRs visible on public GitHub, an upper bound of about eight million cumulative, are a volume reference, not a market share. Every model improvement raises the evaluated count because teams let agents do more. Trust infrastructure trades at multiples that make $100M of metered revenue a $1B company: Chainguard raised at $3.5B on $40M of revenue growing seven-fold, LMArena at $1.7B four months after its first product. Stripe Radar is cited for its network prior (89% of cards it screens have been seen before), not for its pricing, which has moved to tiers.

Why now, and why not in 2028

Supported event surfaces (Claude Code hooks and OpenTelemetry, Codex OpenTelemetry) exist as of this year, so rejections are observable without private parsing. Agent volume per human is crossing the point where reading everything is impossible. Every incumbent has just committed to a position that disqualifies it as referee: GitHub to Copilot, Linear to hosted execution, vendors to their own agents. In 2028 the rejection corpora will exist somewhere; the question is whether they sit in one neutral place or are lost inside six vendors' logs.

Founder fit

Repo and terminal only. Zero new habits: the behavior the product instruments is committing, rejecting, and writing decisions, all of which the founder does daily and prolifically. Solo-buildable in the runway: the ledger and commit-time evaluation exist as a spike; the wedge is two contracts and a rejection adapter.

Plan against five months, with kill criteria

Evidence appendices

What the evidence does not show, stated so nobody has to discover it. No external human has used Panout. Agent work is not reworked more than human work in general: across 11 public repos the pooled gap is 2.5 points with an interval including zero, and attribution from git is unreliable enough that two defensible rules give opposite signs (appendix G). The founder's 22-point gap is robust to every filter (appendix F) and is the founder's. In-session rejection is too rare on this machine (0.14% of tool calls) to certify an autonomy level for any task class; at that base rate the required sample is about 100 times what is on disk, which is why autonomy is derived from the rework label and contract catch rate instead. Logs cover 21 days, so no trend by model version can be claimed. Zero-review merge rates do not distinguish agent PRs from human PRs in popular repos; the distinguishing signal is merge latency. Nobody is asking for AI-code provenance yet. The rework measurement is line-identity, not semantics, and the founder's largest repos are team repos where the founder is a minority committer.