Panout review packOverviewFinal memoMVPLive outputsJudge roundsAppendicesDecisions

Source: docs/research/2026-09-03-defensibility/rounds/r3-memo.md

Panout defensibility, round 3 (2026-09-03): The rework layer for agent labor

One idea, third iteration. Round 1 scored 62 (best single) and 72 (best combination). Round 2 scored 66 with twelve objections; nine were memo or idea fixes and are applied here, three required measurements that are attached as appendices C to E. Two remain that only the founder can close, and they are stated as such in the first section rather than buried.

What the founder owes before anyone reads further

The 2026-09-05 gate result, reported today rather than on the day. STRATEGY.md set a behavioral gate: by 2026-09-05 the v2 ledger must have recorded at least five real sessions and the brief must have recovered work, prevented overlap, or changed a next action. Measured 2026-09-03: no .panout directory exists on this machine, no ledger file exists anywhere on disk, and the repo has no commits since 2026-08-23. The dogfood never started. The gate fails as written, two days early, and this memo does not reset it. The reason it never started is the reason v1 died: the product required the founder to run a command and read a brief, and the founder did neither. The wedge below is designed so that the failure mode cannot recur, because it instruments commits the founder is already making (2,800 in 180 days across founder-controlled repos) and asks for nothing else.

The ten operators, one question, in writing. Assigned 2026-08-22, still outstanding. No memo replaces it. Everything below is priced by the answer.

The two sentences (rewritten after round 2)

Moat. Human review of agent work has not been replaced, it has moved: on this machine humans reject 0.14% of agent actions in session and then rework 60% of agent commits within 30 days, and nobody joins the two. Panout is installed at the only point where the features that explain rework (task class, harness, model, session shape, contract result) can be captured, at commit time, across every harness; the label arrives from git 30 days later, and a team's calibrated map of which task classes it can stop reading cannot be rebuilt by anyone who was not capturing the features when the commit happened.

$1B. Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat. The meter is paid the same whether Panout clears or flags, and it compounds every time models get good enough to be trusted with more, which makes it the one revenue line in developer tooling that strengthens as models improve.

One-liner

Panout closes the rework loop on agent work. It shows a team which classes of agent work get silently fixed later and which do not, derives contracts and earned autonomy from that record, and is paid per outcome evaluated, the same whether it says yes or no.

The constraint set (unchanged, all respected)

Does not control agents. Never in the live agent loop (evaluates at commit, PR, and session-end boundaries). No login, repo and terminal only. No duplicate runs. Raw content stays local; the ledger holds metadata, hashes, contract results, and labels.

The desperate human, named as precisely as the data allows

The founder, measured (git-only backtest, 2026-09-03, appendix A; session census, appendix C). Across nine repos the founder builds alone, the agent share of commits over 180 days runs from 0% to 63%, concentrated in the newest and smallest repos; in two team repos where the founder is a minority committer it is 85%. Agent-majority is a property of greenfield, and greenfield is where a 48-hour wedge installs. Restricted to commits the founder personally authored, agent commits are reworked within 30 days at 60% versus 38% for human commits. The pooled figure (65% versus 49%) is a mix effect and reverses in the two team repos, where agent work is reworked less than human work; it is reported in the appendix, not relied on here. One in six reworked agent commits is reworked by the very next commit. Explicit reverts are 0.16% of commits: correction shows up as rework, not revert, so it is invisible to anyone counting reverts. 99% of agent commits land on the default branch without a merge commit (a weak proxy, since squash merges look identical).

The census that changed this memo (appendix C). 262 Claude Code sessions and 95 Codex sessions in 21 days (the machine's logs begin 2026-08-13; 30-day and 180-day windows are identical for logs). Explicit human rejections of an agent action: 22 in 15,642 tool calls (0.14%). Interrupts: 25. Codex aborts: 14. Stashes machine-wide: zero. Sessions containing any hard negative: 8.4%. Every rejection with a recoverable path was a Markdown edit; none touched code, tests, or config. Most sessions run with permission prompts disabled, which removes the moment at which rejection would occur (appendix D). Meanwhile 60% of the founder's agent commits are reworked within 30 days. The two definitions of "negative" disagree by roughly 460 times. Round 2 called rejections the asset; the data says humans have already stopped rejecting in session and correct afterward, silently, in rework nobody measures. That is the status quo in one number, and it is why the moat, the wedge, and the day-one value are all restated below around rework rather than rejection.

The cohort, observed in public (appendix B, measured 2026-09-03 against the GitHub API). 1.5 million merged PRs from Copilot's coding agent, 5.8 million merged PRs on Codex branches, 694 thousand on Cursor branches, 346 thousand carrying a Claude co-author trailer. In a 150-PR sample of merged Copilot agent PRs, the median time from open to merge is 1.3 minutes and 70% merge within 10 minutes, at a mean of 397 added lines; the control sample of non-agent PRs in popular repos merges at a 25-minute median. Zero-review merge rates are not distinctive (49% for the control), so the claim is not "agents get merged without review", it is "agent output is merged faster than it can be read, into a review culture that was already thin." Sonar's 1,100-developer survey: 96% believe AI code is not reliably correct and 48% say they always check it. Peter Steinberger, publicly: "I don't read much code anymore."

The named consequence. Amazon, March 2026: a "trend of incidents" with "high blast radius" tied to "Gen-AI assisted changes", a roughly six-hour retail outage, and a remediation that is the whole problem in one sentence: every AI-assisted change must now be approved by a senior engineer before deployment, which Amazon called "controlled friction." Amazon disputes the framing and has said the one AI-involved root cause was an engineer acting on stale internal advice; the dispute is itself the point, because the remediation stands either way. That is the status quo's only answer to agent-caused incidents: more human reading, applied to everything, forever. The person who needs Panout is the engineering leader who has just been told to read everything and knows it cannot scale, and the fleet operator one incident away from the same order. DORA 2025 (about 5,000 respondents) finds AI adoption at 90% and continuing to increase change failure and rework; GitClear measures code churn rising from a 3.3% pre-AI baseline to 7.1% in 2025.

Demand in writing, unprompted (appendix E, counts read from the GitHub API on 2026-09-03). The request to stop approving agent actions one at a time has 249 upvotes on a single VS Code issue and 59 on Codex; a Claude Code meta-issue titled "Permissions matching is fundamentally broken, 30+ open issues, community building workarounds" has 78. GitHub code search finds on the order of 15,000 shell scripts invoking the skip-permissions flag and 2,800 settings files enabling bypass mode, plus auto-approve hooks shipped to customers by Railway and Render. Anthropic's own February 2026 measurement: full auto-approve use rises from about 20% of sessions for new users to over 40% by 750 sessions, and Anthropic concludes that approving every action "will create friction without necessarily producing safety benefits." Retention is a live grievance: issues about Claude Code silently deleting transcripts after 30 days carry 31 and 23 upvotes. Per-unit payment for review of agent output is established: Cursor Bugbot at $1 to $1.50 per run, Greptile at $1 per review, CodeRabbit at $0.25 per file beyond quota and an estimated $40M ARR growing seven-fold (secondary source). What nobody asks for in words is "tell me what not to read"; people ask to read better (180 upvotes for a diff-review UI). The behavioral record says they gave up reading instead. That is why the wedge leads with the rework bill, not with a reading tool: the ask that exists is "stop making me approve", and the cost that exists is 60% rework, and Panout connects the two.

What gets them promoted or fired. Shipping volume without incidents. The lever they control is how much agent output they read, and today that lever has two positions: everything, or nothing.

Status quo and its cost

Position A, read everything: caps agent throughput at human reading speed. The two team repos in the backtest merge about 1,200 to 1,300 agent commits a month each, 85% of all commits; nobody is reading that. Position B, read nothing: git diff --stat, skim, merge. Cost is the measured rework gap: agent work reworked within 30 days 15 to 22 points more often than human work in the same repos, with the correction landing a median 17 hours later. Industry-wide the same signal shows as churn doubling (GitClear) and instability rising with adoption (DORA 2025), and Veracode finds models pick the insecure implementation in 45% of tasks with no improvement across model generations. Position C, second-agent review (gstack /review, Cursor Bugbot at $1 to $1.50 per run, Copilot review, Codex review): a reviewer with no memory of the team's bar and no record of whether it has ever caught anything. It produces more text to read, which is the problem restated. Bugbot's shift to per-review pricing in June 2026 is also the market telling us that per-unit pricing at this boundary is accepted.

Verified against vendor documentation on 2026-09-03 (appendices B and E): across Copilot code review, GitHub rulesets and merge queue, Linear Coding Sessions and Guided Reviews, Cursor Bugbot, Claude Code Auto mode and hooks, VS Code auto-approve, Codex execpolicy and review, none learns a per-team bar from that team's own history, none enforces autonomy per task class, and none joins what was approved to what was later reworked. The closest approximations are hand-written allowlists, a stateless per-call classifier, and effort dials. The narrow gap claim, and the only one the evidence supports, is learned from your own history and recorded durably.

None of these learn. Every day of agent work today produces zero durable information about what this team accepts.

The wedge: value in the first minute, zero new behavior, 48 hours to build

Panout's v1 died because it required the founder to read an inbox; v2 never started because it required the founder to run a collector and read a brief. This wedge produces its first result from data that already exists and then instruments behavior that already happens.

  1. Minute one: the rework map. panout init runs the appendix A measurement on the repo's own history and prints one table: for each task class (derived from paths and commit shape), the share of agent commits reworked within 30 days, the median hours to rework, and the same numbers for human commits. On this founder's repos that table already says: agent commits reworked 60% versus 38%, one in six by the very next commit. No streak, no contract, no login. This is the demo, the hook, and the reason to keep the tool installed.
  2. Commit time: capture the features. A commit-time step records, as metadata only, the task class, harness and model (from trailers and supported session surfaces), session shape, and the result of any contracts. Nothing is read from transcripts; the census in appendix C was produced with exactly this discipline.
  3. Contracts for the two worst classes. For the two task classes with the highest rework rate, panout init proposes acceptance contracts from the existing test setup and AGENTS.md; evaluation prints only failures. Committing after a printed failure is an override and is recorded as one. Acceptance is the commit.
  4. Thirty days later: the label attaches. Each captured commit is labeled reworked or not, and the map updates. Autonomy levels per task class are derived from it, with the catch-rate calibration below as the second input.
  5. One optional key turns a rework or override into a decision record in decisions/, summarized into AGENTS.md so every harness reads it at session start.

The fourteen-day test needs no new habit: does the founder keep committing with the capture on, and does the rework map change what gets a contract? The measurement in appendix D tried to test, on the same commits, whether session evidence predicts rework better than git alone. It could not be answered, and the reason is the product argument in miniature: harness logs on this machine are retained for 21 days, the rework label needs 30, and the two windows never overlap, so retrospective joining is impossible for anyone. Squash merges compound it: of 49 commit hashes agents printed in sessions, 69% never landed under that hash, and in the largest repo 0 of 40 did. Retroactive linkage recovers about 13% and the recoverable subset is the reviewed one. The features must be captured at commit time under a stable identifier written into the commit itself, or they are gone within three weeks. The first honest predictive delta therefore has a date: 30 days after capture starts, 2026-10-05 at the earliest if the wedge ships this week. Until then the claim that commit-time features add predictive power over git is a hypothesis with a measurement scheduled, not a result.

The moat, in three layers that GitHub cannot copy

Layer 1: the rework join (features at commit time, label 30 days later). The label, "was this agent commit reworked within 30 days", is computable from git by anyone with the repository, GitHub included. The features that make the label useful are not in git: task class, harness and model, session shape (tool calls, errors, duration, subagent share), which contracts ran and what they returned, and whether a human read the diff. Those exist only at the local commit boundary, only at the moment of the commit, and only across harnesses if a neutral party is standing there. Harness vendors have the features for their own agent and lose the label when the session ends; GitHub has the label and none of the features; Linear has neither for local work. Panout is the join. A team that adopts it in 2028 starts with two years less of its own labeled history than a team that adopted in 2026, and cannot backfill the features: on this machine the harness logs that hold them are pruned after about 21 days (appendix D), so the raw material for the join is destroyed faster than the label can arrive.

Layer 2: earned autonomy as the team's calibrated policy. Each task class carries a level from 0 (human reads everything) to 4 (auto-merge). Levels are computed from acceptance streaks, rejection history, and contract catch rate (see calibration below), and enforced at the merge boundary. This is the switching cost: leave Panout and you go back to reading everything, because the policy is derived from a history no other party holds. It also answers "what happens when GitHub ships AI review confidence": GitHub's confidence becomes one more evidence input to the contract, the way a CI check is. GitHub's signal is global and mechanical; the policy is the team's.

Layer 3: cross-team priors (the Radar effect). Each team's bar is private. What is shared, with consent and as aggregates, is the pattern library: which fault classes get rejected, which contract shapes catch them, and which task signatures earn autonomy fastest. A new team's contracts start warm. This is why one merchant's fraud history is worthless and Stripe Radar's is a business: the per-customer data is thin, the cross-customer prior is thick, and no customer can get the prior without joining. Neither a harness vendor (sees one agent) nor GitHub (sees no rejections) can build this prior.

The tension the moat has to survive, named. The market argument says models improve and teams let agents do more. A moat built on refusals thins on the same curve, and the census shows refusals are already near zero on this machine (appendix C, table 3: 21 days of history, half the negatives on one day, model comparison confounded with time, no trend claim possible). So the asset is not refusals. It is the join between commit-time features and the 30-day rework label, which does not thin as models improve: rework fell from 38% for human commits to a floor, not to zero, and the human rate itself is 38 to 51%. Better models shift which task classes are safe; they do not remove the need to know which ones.

Calibration, so the moat is measurable and the product is not a linter. Off the hot path, Panout injects a small library of agent-characteristic faults into a sample of accepted diffs and re-runs the contract, producing a catch rate per contract. Autonomy requires both a clean streak and a catch rate above threshold. Weak contracts cannot earn autonomy. This converts the round 1 day-45 kill criterion ("contracts are theater") from a risk into a metric that is measured weekly.

Incumbent response, and why the idea survives each

IncumbentWhat they shipWhy it does not close the seam
GitHubCopilot code review can now submit approving reviews that satisfy branch protection (public preview, appendix B), plus rulesets, merge queue, auto-mergeGitHub has stepped into the approval decision. It still sees survivors only, has no local session or rejection evidence, applies rules globally and mechanically, and is conflicted via Copilot. The defense is switching cost, not neutrality: a team that leaves Panout goes back to reading everything, because its autonomy policy is derived from a rejection history nobody else holds. Copilot's approval becomes one more evidence input to the contract.
LinearCoding Sessions, Guided Reviews, per-issue model choiceCloud-resident graph, sees only its own hosted sessions, bills AI credits for execution, so cannot judge its own agent or see local work. Panout syncs cleared outcomes back to the issue.
Harness auto-approval (Claude Code Auto mode, VS Code terminal auto-approve, Codex execpolicy, Gemini policy files)Per-call classified approval (Claude Code, a remote classifier) or static allowlists (the others)This is the closest shipped surface and the memo does not claim it is empty. All of it decides per action, in session, statelessly: none learns from the team's own acceptance or rework history, none is per task class, none leaves an acceptability record, and every high-vote issue against them is about being unavailable, over-permissive, or over-restrictive. Panout consumes their decisions as evidence and answers the question they cannot: which of the things you auto-approved got fixed later.
Anthropic, OpenAI, CursorIn-session review, auto modes, thumbs up/down, memory from correctionsEach sees one agent, in session, with no post-hoc outcome and no cross-vendor comparison. A rating of your own work is not a rating. Panout consumes their supported event surfaces.
Entire (entireio/cli; the public sweep in appendix B could not independently verify its feature set, so this row rests on the repo's 2026-08-23 research)Cross-agent transcript archive, checkpoints, judge, expertsPreserves everything; no acceptance labels, no enforcement; its value grows with raw retention, which conflicts with metadata-only. Optional evidence input.
Harness, Jellyfish, SwarmiaLeader dashboards on sessions, tokens, attributionRead-only reporting for managers; no enforcement, no rejections, no per-team bar.
Vanta, DrataAI-code control modules pulling GitHub evidenceTop-down compliance; they will need an evidence source for local agent work and Panout's receipts are that source. Channel, not competitor.

Business model: metered per outcome evaluated, paid the same for yes and no

Why this fixes the round 1 objection. Revenue tracks agent volume. A team of two humans and forty agents pays Linear $32 a month; the same team produces thousands of evaluable outcomes a month. Cursor's Bugbot already charges $1 to $1.50 per review run at this boundary and moved there from per-seat pricing in June 2026, so per-unit pricing is accepted by exactly this buyer. Prosumer conversion rates are irrelevant when the meter is on the agents.

$1B math, stated as a ramp with its denominator. Merged agent PRs visible on public GitHub are an upper bound of about eight million cumulative across Codex, Copilot, Cursor, Claude, Devin, and Jules (rows overlap and branch-prefix queries include human commits; appendix B). Cursor reports 50,000 paying teams (secondary source). Those are reference populations, not Panout's market share. The ramp: 2,000 teams evaluating 500 outcomes a month at a third of a dollar is $4M a year and is the month-12 test of whether the meter works; 10,000 teams is $20M; 50,000 teams, roughly today's count of teams paying for one agent product, is $100M, and by 2028 the population of teams running agent fleets is a multiple of that. Every model improvement raises the evaluated count because teams let agents do more. Trust infrastructure trades at multiples that make $100M of metered revenue a $1B company: Chainguard raised at $3.5B on $40M of revenue growing seven-fold, LMArena at $1.7B four months after its first product. Stripe Radar is cited for its network prior (89% of cards it screens have been seen before), not for its pricing, which has moved to tiers.

Why now, and why not in 2028

Supported event surfaces (Claude Code hooks and OpenTelemetry, Codex OpenTelemetry) exist as of this year, so rejections are observable without private parsing. Agent volume per human is crossing the point where reading everything is impossible. Every incumbent has just committed to a position that disqualifies it as referee: GitHub to Copilot, Linear to hosted execution, vendors to their own agents. In 2028 the rejection corpora will exist somewhere; the question is whether they sit in one neutral place or are lost inside six vendors' logs.

Founder fit

Repo and terminal only. Zero new habits: the behavior the product instruments is committing, rejecting, and writing decisions, all of which the founder does daily and prolifically. Solo-buildable in the runway: the ledger and commit-time evaluation exist as a spike; the wedge is two contracts and a rejection adapter.

Plan against five months, with kill criteria

Evidence appendices

What the evidence does not show, stated so nobody has to discover it. No external human has used Panout. In-session rejection is too rare on this machine (0.14% of tool calls) to certify an autonomy level for any task class; at that base rate the required sample is about 100 times what is on disk, which is why autonomy is derived from the rework label and contract catch rate instead. Logs cover 21 days, so no trend by model version can be claimed. Zero-review merge rates do not distinguish agent PRs from human PRs in popular repos; the distinguishing signal is merge latency. Nobody is asking for AI-code provenance yet. The rework measurement is line-identity, not semantics, and the founder's largest repos are team repos where the founder is a minority committer.