Panout defensibility, final (2026-09-03): The commit-boundary record for agent labor
One idea, fifth iteration. Scores so far: 62 and 72 (round 1), 66 (round 2), 67 (round 3), 64 (round 4, after appendix G retracted the general "agents get reworked more" claim). The round 4 judge's ceiling without new external evidence is about 70, conditional on appendix H showing that a repo's rework gap is a stable property rather than a draw. 90 requires: ten operators in writing, three external maintainers reacting to their own map, one paying team, and a predictive measurement that cannot exist before 2026-10-05 because it needs 30 days of commit-time capture that has not started. This round applies the remaining reachable fixes and adds appendix H. It does not claim to have closed what only the founder can close.
The strongest defensible version, after eight measurements
Panout is the commit-boundary recorder and enforcement point for teams that have already stopped reading agent output. It ships as a local commit hook that runs a small set of contracts, assertions the team writes or that Panout proposes from the existing test setup, prints only failures, and records immutably in the commit every time a human shipped past one. Each contract earns a trust level from fault injection, not from an observational rework rate: Panout injects known agent-characteristic faults into accepted diffs off the hot path and measures the contract's catch rate, a designed experiment with a same-day answer and a sample size the founder chooses. Autonomy is granted per contract, never per task class, because appendix H shows the task class is below the resolution of any sample a real team produces. The 30-day rework label is kept for exactly one job, a slow audit of whether trusted contracts are silently failing, and is never sold as a map. The day-one artifact the buyer receives is an audit trail of auto-approved agent work with outcomes attached: the thing Amazon's remediation implies every engineering organization will be asked for, the thing 15,000 skip-permissions scripts and 40%-auto-approve sessions have made unrecordable, and the evidence source Vanta and Drata will need. The moat is that this record exists nowhere else and cannot be reconstructed backwards by anyone, GitHub included. The switching cost is the accumulated contract-level trust policy. The meter is per evaluation, identical for pass and fail, and rises with autonomy because a level-4 contract still runs on every commit. This company is smaller than the one in earlier rounds, is consistent with every appendix including the disconfirming ones, and does not ask the user to wait thirty days to learn anything.
Moat, two sentences. Every team running agents is already deciding, dozens of times a day, which diffs it will not read, and nothing anywhere records that it decided: git cannot tell whether an agent wrote a commit or was merely open in the next pane, no vendor's log contains the moment a human shipped past a failing check, and by the time the consequence shows up thirty days later the session that caused it has been garbage-collected. Panout sits at the commit boundary and writes the decision, the check it overrode, and the outcome into the commit while all three still exist, so whatever a team eventually learns about what it can stop reading is derived from a record no one who was not there can reconstruct.
$1B, two sentences. Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat, paid identically whether Panout clears or flags. Two thousand teams at five hundred evaluated outcomes a month is $4M a year and is the month-twelve test; fifty thousand teams is $100M; and unlike seats, the meter reprices itself upward every time a model gets good enough to be trusted with more, because a level-4 contract still runs on every commit and what stops is the human reading, not the evaluation.
Score and ceiling, stated by the independent judge. Five rounds scored 62/72, 66, 67, 64, 59. Memo and idea changes alone cap at about 65. Reaching 90 requires: ten fleet operators in writing (earliest 2026-09-17), three external maintainers reacting to their own artifact (2026-09-17), commit-time capture shipped in the founder's repos within 48 hours with contracts firing on roughly 30% of commits so the day-60 test is powered (labels from 2026-10-05), two contracts with fault-injection catch rates that separate on a repo that is not the founder's (2026-09-20), the predictive delta honestly powered (about 2026-10-20), and one team paying the meter without being asked twice (about 2026-10-15). Earliest 90: about 2026-10-25, on this design only. The judge's last instruction was that no sixth memo is worth writing.
What the founder owes before anyone reads further
The 2026-09-05 gate result, reported today rather than on the day. STRATEGY.md set a behavioral gate: by 2026-09-05 the v2 ledger must have recorded at least five real sessions and the brief must have recovered work, prevented overlap, or changed a next action. Measured 2026-09-03: no .panout directory exists on this machine, no ledger file exists anywhere on disk, and the repo has no commits since 2026-08-23. The dogfood never started. The gate fails as written, two days early, and this memo does not reset it. The reason it never started is the reason v1 died: the product required the founder to run a command and read a brief, and the founder did neither. The wedge below is designed so that the failure mode cannot recur, because it instruments commits the founder is already making (5,294 founder-authored non-merge commits in 180 days across the nine repos the founder controls, by git log) and asks for nothing else.
The ten operators, one question, in writing. Assigned 2026-08-22, still outstanding. No memo replaces it. Everything below is priced by the answer.
The two sentences (rewritten after round 2)
Moat. Every team running agents is already deciding, dozens of times a day, which diffs it will not read, and the record of that decision dies in twenty-one days while its consequence takes thirty to appear, so nobody, GitHub included, can reconstruct it afterwards: git cannot even tell whether an agent wrote a commit or was merely open in the next pane. Panout sits at the commit boundary and writes the decision and its outcome into the commit while both still exist, so a team's map of what it can safely stop reading is uncopyable by anyone who was not recording at the moment the commit happened.
$1B. Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat, paid identically whether Panout clears or flags. Two thousand teams at five hundred evaluated outcomes a month is $4M a year and is the month-twelve test; fifty thousand teams is $100M; and unlike seats, the meter reprices itself upward every time a model gets good enough to be trusted with more, because a level-4 class is still evaluated on every commit and what stops is the human reading, not the evaluation.
One-liner
Panout records and enforces the decision to stop reading agent work. It runs the team's contracts at the commit boundary, records every override with its outcome, grants autonomy per contract from measured catch rates, and is paid per outcome evaluated, the same whether it says yes or no.
The constraint set (unchanged, all respected)
Does not control agents. Never in the live agent loop (evaluates at commit, PR, and session-end boundaries). No login, repo and terminal only. No duplicate runs. Raw content stays local; the ledger holds metadata, hashes, contract results, and labels.
The desperate human, named as precisely as the data allows
The founder, measured (git-only backtest, 2026-09-03, appendix A; session census, appendix C). Across nine repos the founder builds alone, the agent share of commits over 180 days runs from 0% to 63%, concentrated in the newest and smallest repos; in one team repo where the founder is a minority committer (checked out twice on this machine, so appendix A counts it twice) it is 85%. Agent-majority is a property of greenfield, and greenfield is where a 48-hour wedge installs. Restricted to commits the founder personally authored, agent commits are reworked within 30 days at 60% versus 38% for human commits. The pooled figure (65% versus 49%) is a mix effect, double-counts one team repo checked out twice, and includes bot commits on the human side; it is reported in the appendix, not relied on here (appendix F has the corrected pooled gap, 22 points with bots excluded). One in six reworked agent commits is reworked by the very next commit. Explicit reverts are 0.16% of commits: correction shows up as rework, not revert, so it is invisible to anyone counting reverts. 99% of agent commits land on the default branch without a merge commit (a weak proxy, since squash merges look identical).
The census that changed this memo (appendix C). 262 Claude Code sessions and 95 Codex sessions in 21 days (the machine's logs begin 2026-08-13; 30-day and 180-day windows are identical for logs). Explicit human rejections of an agent action: 22 in 15,642 tool calls (0.14%). Interrupts: 25. Codex aborts: 14. Stashes machine-wide: zero. Sessions containing any hard negative: 8.4%. Of the four rejections with a recoverable file path, all four were Markdown edits (18 of 22 had no recoverable path). Most sessions run with permission prompts disabled, which removes the moment at which rejection would occur (appendix D). Appendix D, using looser matching, counts 53 rejections and 31 interrupts; appendix C's anchored matching (22 and 25) is the count relied on here because the looser patterns matched marker strings echoed into tool output by earlier measurement runs. Meanwhile 60% of the founder's agent commits are reworked within 30 days; the two rates have different denominators (per tool call, per commit) and are not divided into a ratio. Round 2 called rejections the asset; the data says humans have already stopped rejecting in session and correct afterward, silently, in rework nobody measures. That is the status quo in one number, and it is why the moat, the wedge, and the day-one value are all restated below around rework rather than rejection.
The cohort, observed in public (appendix B, measured 2026-09-03 against the GitHub API). 1.5 million merged PRs from Copilot's coding agent, 5.8 million merged PRs on Codex branches, 694 thousand on Cursor branches, 346 thousand carrying a Claude co-author trailer. In the long tail of small repos, a 150-PR sample of merged Copilot agent PRs shows a median time from open to merge of 1.3 minutes and 70% merging within 10 minutes, at a mean of 397 added lines. In popular repos (over 500 stars) the difference disappears: Copilot agent PRs merge at a 19.6-minute median against a 25-minute control, and Codex-branch PRs at 28 minutes, slower than the control. Zero-review merge rates are not distinctive either (49% for the control). So the populations differ: in small and solo repos agent output is merged faster than it can be read; in popular repos review culture was already thin for everyone. Panout's first buyer is the first population, the small team running fleets, and the memo does not claim the second. Sonar's 1,100-developer survey: 96% believe AI code is not reliably correct and 48% say they always check it. Peter Steinberger, publicly: "I don't read much code anymore."
The named consequence. Amazon, March 2026: a "trend of incidents" with "high blast radius" tied to "Gen-AI assisted changes", a roughly six-hour retail outage, and a remediation that is the whole problem in one sentence: every AI-assisted change must now be approved by a senior engineer before deployment, which Amazon called "controlled friction." Amazon disputes the framing and has said the one AI-involved root cause was an engineer acting on stale internal advice; the dispute is itself the point, because the remediation stands either way. That is the status quo's only answer to agent-caused incidents: more human reading, applied to everything, forever. The person who needs Panout is the engineering leader who has just been told to read everything and knows it cannot scale, and the fleet operator one incident away from the same order. DORA 2025 (about 5,000 respondents) finds AI adoption at 90% and continuing to increase change failure and rework; GitClear measures code churn rising from a 3.3% pre-AI baseline to 7.1% in 2025.
Demand in writing, unprompted (appendix E, counts read from the GitHub API on 2026-09-03). The request to stop approving agent actions one at a time has 249 upvotes on a single VS Code issue and 59 on Codex; a Claude Code meta-issue titled "Permissions matching is fundamentally broken, 30+ open issues, community building workarounds" has 78. GitHub code search finds on the order of 15,000 shell scripts invoking the skip-permissions flag and 2,800 settings files enabling bypass mode, plus auto-approve hooks shipped to customers by Railway and Render. Anthropic's own February 2026 measurement: full auto-approve use rises from about 20% of sessions for new users to over 40% by 750 sessions, and Anthropic concludes that approving every action "will create friction without necessarily producing safety benefits." Retention is a live grievance: issues about Claude Code silently deleting transcripts after 30 days carry 31 and 23 upvotes. Per-unit payment for review of agent output is established: Cursor Bugbot at $1 to $1.50 per run, Greptile at $1 per review, CodeRabbit at $0.25 per file beyond quota and an estimated $40M ARR growing seven-fold (secondary source). What nobody asks for in words is "tell me what not to read"; people ask to read better (180 upvotes for a diff-review UI). The behavioral record says they gave up reading instead. Panout does not compete for the auto-approve ask; Claude Code Auto mode, VS Code terminal auto-approve, and Codex execpolicy own it, free. The 15,000 scripts and the 40% of sessions in full auto-approve are the size of the population that has already made the trade, and none of them can see what it cost. That is why the wedge leads with the rework bill, not with a reading tool: the ask that exists is "stop making me approve", the cost that exists is the rework gap, and Panout connects the two. The person who buys is the one who already took auto-approve and now owns the consequence.
What gets them promoted or fired. Shipping volume without incidents. The lever they control is how much agent output they read, and today that lever has two positions: everything, or nothing.
Status quo and its cost
Position A, read everything: caps agent throughput at human reading speed. The team repo in the backtest merges about 1,200 agent commits a month, 85% of all its commits; nobody is reading that. Position B, read nothing: git diff --stat, skim, merge. Cost is the measured rework gap: for commits the founder personally authored, agent work is reworked within 30 days 22 points more often than hand-written work (60% versus 38%); pooled across the founder's repos excluding upstream clones the gap is 16 points (13 with clones included); per repo it runs from minus 20 to plus 35 points and reverses sign in the large team repo and in four repos with one to ten agent commits (appendix A). The correction lands a median 17 hours later. Because the line-identity metric has a high floor (human work is reworked 38 to 51%) and could be counting reformatting, appendix F recomputed the same 2,859 commits under stricter definitions. The gap survives every one of them except the different-author test, which is degenerate on a single-author slice by construction. Whitespace and move insensitivity changes it by under two points, so reformatting was not what the metric was counting. Requiring the rework to look corrective (a follow-up with a fix-word subject or touching a test) cuts the human floor to 34% on the founder's commits and 29% pooled while the agent rate barely moves, so the founder-only gap is 20 points with a bootstrap interval of 14 to 26, and 11 points with an interval of 6 to 17 under the strictest variant that drops the test-path clause. The team repo's reversal was an artifact: 36 bot release commits classed as human that overwrite the same version strings on every release and score 100% reworked; under the corrective definition that repo flips to agent work reworked 32 points more, and with bots excluded the line-identity gap there is under two points, meaning gone. Two corrections to appendix A surfaced in the process: the two team repos are one git history checked out twice, so pooled figures there double-count, and 28% of "human" commits pooled are bots (dependabot, release republishing, a research pipeline); with bots excluded the pooled gap is 22 points. Under the corrective definition no sign reversal remains that is backed by more than ten agent commits; under the strictest variant, which the product adopts, one appears (a 68-commit gateway service, 44 agent commits, minus 12 points), so the honest statement is that the founder's gap is robust in direction across definitions and its size is between 11 and 23 points depending on how strictly "rework" is defined. Industry-wide, churn is doubling (GitClear) and instability rises with adoption (DORA 2025), and Veracode finds models pick the insecure implementation in 45% of tasks with no improvement across model generations; but appendix G shows the agent-versus-human rework gap itself is team-specific and near zero pooled across public repos, so the cost of Position B is not a universal tax, it is an unknown that differs per team and per task class, and that unknown is the product. Position C, second-agent review (gstack /review, Cursor Bugbot at $1 to $1.50 per run, Copilot review, Codex review): a reviewer with no memory of the team's bar and no record of whether it has ever caught anything. It produces more text to read, which is the problem restated. Bugbot's shift to per-review pricing in June 2026 is also the market telling us that per-unit pricing at this boundary is accepted.
Verified against vendor documentation on 2026-09-03 (appendices B and E): across Copilot code review, GitHub rulesets and merge queue, Linear Coding Sessions and Guided Reviews, Cursor Bugbot, Claude Code Auto mode and hooks, VS Code auto-approve, Codex execpolicy and review, none learns a per-team bar from that team's own history, none enforces autonomy per task class, and none joins what was approved to what was later reworked. The closest approximations are hand-written allowlists, a stateless per-call classifier, and effort dials. The narrow gap claim, and the only one the evidence supports, is learned from your own history and recorded durably.
None of these learn. Every day of agent work today produces zero durable information about what this team accepts.
The wedge: value in the first minute, zero new behavior, 48 hours to build
Panout's v1 died because it required the founder to read an inbox; v2 never started because it required the founder to run a collector and read a brief. This wedge produces its first result from data that already exists and then instruments behavior that already happens.
- Minute one: the rework map, now demoted to a diagnostic (see appendix H below).
panout initruns the appendix A measurement on the repo's own history and prints one table: for each task class (derived from paths and commit shape), the share of agent commits reworked within 30 days, the median hours to rework, and the same numbers for human commits. The map uses the strictest corrective definition from appendix F (rework that is whitespace and move insensitive and whose follow-up either comes from a different author or carries a fix-word subject; the test-path clause is dropped because appendix F shows it is mostly measuring "the follow-up touched a test"), with bot authors excluded. That definition is adopted because it is the most defensible, not because it flatters: on this founder's repos it says agent commits reworked 34% versus 23%, a gap of 11 points with a bootstrap interval of 6 to 17, and under it one founder repo with 44 agent commits reverses sign. No streak, no contract, no login.
What the map prints on strangers' repos (appendix G), stated before anyone else finds it. The same measurement was run on 11 public repositories with heavy agent traffic (aspire, spec-kit, n8n, mastra, azure-sdk-tools, eliza, typescript-go, gradio, vscode-copilot-chat, deno, temporal; 2,042 commits). Pooled, agent-trailer commits are reworked 2.5 points more than human commits, with a 95% interval from minus 1.8 to plus 6.9. Per repo the gap runs from minus 20 to plus 43: positive in five, negative in three, noise in three. Matching samples by calendar month moves the pooled gap to minus 1.3. Switching from trailer attribution to PR-level attribution flips the sign in all three repos where both could be computed, and in the repo with the largest gap the whole effect is one contributor. The founder's 22-point gap is real for the founder and is not a property of agent code in general. Three consequences follow and the memo accepts all of them. First, "agents get reworked more" is retracted as a general claim; the hook is not that agents are worse, it is that nobody knows which of their own task classes are, and the answer differs by team, which is the argument for a per-team calibrated map and the final argument against the global index idea from round 1. Second, the day-one map from git alone is weaker than round 3 claimed: it may print "no difference" for a stranger, and its per-class ranking is readable only where samples are large; the one pooled per-class pattern with any support is agent work reworked more in source and config and less in tests and docs, on repos whose individual gaps disagree. Third, git cannot attribute authorship: trailer and PR-level rules give opposite answers on identical commits, because a trailer means an agent was in the room, not that it shipped the change. That is the strongest evidence yet that the features have to be captured at commit time from the session that produced the commit; it is also a limit on any product that promises a map from history alone, Panout included. Appendix H then tested whether the per-team gap is a property or a draw, and it is a draw. Splitting each of the 11 public repos into two halves by time, a repo's gap in the first half does not predict its gap in the second: rank correlation 0.16 with a repo-clustered interval from minus 0.62 to plus 0.90, signs agreeing in 7 of 11 where a coin gives 5.5, and two repos flipping sign by more than sampling error allows. A random re-split of the same commits produces a median correlation of 0.48, so the observed value sits at the fifth percentile of pure noise, which is the signature of drift over calendar time, not of a stable trait measured imprecisely. Per-class orderings inside a repo are not measurable at 100 agent commits per half, and the two worst classes are readable in 2 of 11 repos. Under the strictest corrective definition the pooled gap across public repos is 7.7 points with a clustered interval of 1.3 to 13.5, but that separation is carried by the "reworked by a different author" clause, which measures file ownership as much as quality. The consequence for the wedge is stated without softening: a rework map computed from git history is not a day-one hook for a stranger. It is noise-dominated at the sample sizes real repos have, it drifts, and it cannot attribute authorship. What the rework label is good for is the thing round 3 said it was for before the hook was bolted on: a 30-day outcome attached to features that were captured at commit time, on the team's own future commits, accumulating into a per-class record over months. Minute-one value therefore cannot come from history. It has to come from the commit-time step itself: contracts that print only failures, and the override record they generate. That is a linter with a memory on day one, and the memo says so rather than promising a map it cannot print. This is the demo, the hook, and the reason to keep the tool installed.
- Commit time: capture the features. A commit-time step records, as metadata only, the task class, harness and model (from trailers and supported session surfaces), session shape, and the result of any contracts. Nothing is read from transcripts; the census in appendix C was produced with exactly this discipline.
- Contracts for the two worst classes. For the two task classes with the highest rework rate,
panout initproposes acceptance contracts from the existing test setup andAGENTS.md; evaluation prints only failures. Committing after a printed failure is an override and is recorded as one. Acceptance is the commit. - Thirty days later: the label attaches. Each captured commit is labeled reworked or not, and the map updates. Autonomy levels per task class are derived from it, with the catch-rate calibration below as the second input.
- One optional key turns a rework or override into a decision record in
decisions/, summarized intoAGENTS.mdso every harness reads it at session start.
The fourteen-day test needs no new habit: does the founder keep committing with the capture on, do contract failures get printed and either fixed or overridden, and does the override count per class fall? The measurement in appendix D tried to test, on the same commits, whether session evidence predicts rework better than git alone. It could not be answered, and the reason is the product argument in miniature: harness logs on this machine are retained for 21 days, the rework label needs 30, and the two windows never overlap, so retrospective joining is impossible for anyone. Squash merges compound it: of 49 commit hashes agents printed in sessions, 69% never landed under that hash, and in the largest repo 0 of 40 did. Retroactive linkage recovers about 13% and the recoverable subset is the reviewed one. The features must be captured at commit time under a stable identifier written into the commit itself, or they are gone within three weeks. The first honest predictive delta therefore has a date: 30 days after capture starts, 2026-10-05 at the earliest if the wedge ships this week. Until then the claim that commit-time features add predictive power over git is a hypothesis with a measurement scheduled, not a result.
The moat, in three layers that GitHub cannot copy
Layer 1: the override record and the commit-time join. Concession first: task class (from paths and commit shape) and harness and model (from commit trailers) are recoverable from git by anyone, and appendix A classified 15,440 agent commits over 180 days using nothing else. GitHub can compute those. What is not in git, and not in any vendor's logs, is three things. First, the override: a human committed after a printed contract failure, on a specific diff, for a specific reason. That is a human judgment about agent work, it is generated by the act the human already performs, it is the negative class that in-session rejection turned out not to be, and no incumbent records it because no incumbent runs contracts at the local commit boundary. Second, the contract result itself, per commit, per class. Third, session shape from supported harness surfaces, which appendix D shows is pruned after about 21 days on this machine, nine days before the rework label can arrive. The join of those three to the 30-day label cannot be built retrospectively by anyone, GitHub included; a team that adopts in 2028 starts with two years less of its own labeled history than a team that adopted in 2026.
Layer 2: earned autonomy as the team's calibrated policy, per contract. Each contract (not each task class; appendix H shows the class is unmeasurable at real sample sizes) carries a level from 0 (human reads everything) to 4 (auto-merge). Levels are computed from acceptance streaks, rejection history, and contract catch rate (see calibration below), and enforced at the merge boundary. This is the switching cost: leave Panout and you go back to reading everything, because the policy is derived from a history no other party holds. It also answers "what happens when GitHub ships AI review confidence": GitHub's confidence becomes one more evidence input to the contract, the way a CI check is. GitHub's signal is global and mechanical; the policy is the team's.
Layer 3: cross-team priors (the Radar effect). Each team's bar is private. What is shared, with consent and as aggregates, is the pattern library: which fault classes get rejected, which contract shapes catch them, and which task signatures earn autonomy fastest. A new team's contracts start warm. This is why one merchant's fraud history is worthless and Stripe Radar's is a business: the per-customer data is thin, the cross-customer prior is thick, and no customer can get the prior without joining. Neither a harness vendor (sees one agent) nor GitHub (sees no rejections) can build this prior.
The tension the moat has to survive, named. The market argument says models improve and teams let agents do more. A moat built on refusals thins on the same curve, and the census shows refusals are already near zero on this machine (appendix C, table 3: 21 days of history, half the negatives on one day, model comparison confounded with time, no trend claim possible). So the asset is not refusals. It is the join between commit-time features and the 30-day rework label, which does not thin as models improve: rework fell from 38% for human commits to a floor, not to zero, and the human rate itself is 38 to 51%. Better models shift which task classes are safe; they do not remove the need to know which ones.
Calibration, so the moat is measurable and the product is not a linter. Off the hot path, Panout injects a small library of agent-characteristic faults into a sample of accepted diffs and re-runs the contract, producing a catch rate per contract. Autonomy requires both a clean streak and a catch rate above threshold. Weak contracts cannot earn autonomy. This converts the round 1 day-45 kill criterion ("contracts are theater") from a risk into a metric that is measured weekly.
Incumbent response, and why the idea survives each
| Incumbent | What they ship | Why it does not close the seam |
|---|---|---|
| GitHub | Copilot code review can now submit approving reviews that satisfy branch protection (public preview, appendix B), plus rulesets, merge queue, auto-merge | GitHub has stepped into the approval decision. It still sees survivors only, has no local session or rejection evidence, applies rules globally and mechanically, and is conflicted via Copilot. The defense is switching cost, not neutrality: a team that leaves Panout goes back to reading everything, because its autonomy policy is derived from a rejection history nobody else holds. Copilot's approval becomes one more evidence input to the contract. |
| Linear | Coding Sessions, Guided Reviews, per-issue model choice | Cloud-resident graph, sees only its own hosted sessions, bills AI credits for execution, so cannot judge its own agent or see local work. Panout syncs cleared outcomes back to the issue. |
| Harness auto-approval (Claude Code Auto mode, VS Code terminal auto-approve, Codex execpolicy, Gemini policy files) | Per-call classified approval (Claude Code, a remote classifier) or static allowlists (the others) | This is the closest shipped surface and the memo does not claim it is empty. All of it decides per action, in session, statelessly: none learns from the team's own acceptance or rework history, none is per task class, none leaves an acceptability record, and every high-vote issue against them is about being unavailable, over-permissive, or over-restrictive. Panout consumes their decisions as evidence and answers the question they cannot: which of the things you auto-approved got fixed later. |
| Anthropic, OpenAI, Cursor | In-session review, auto modes, thumbs up/down, memory from corrections | Each sees one agent, in session, with no post-hoc outcome and no cross-vendor comparison. A rating of your own work is not a rating. Panout consumes their supported event surfaces. |
| Entire (entireio/cli; the public sweep in appendix B could not independently verify its feature set, so this row rests on the repo's 2026-08-23 research) | Cross-agent transcript archive, checkpoints, judge, experts | Preserves everything; no acceptance labels, no enforcement; its value grows with raw retention, which conflicts with metadata-only. Optional evidence input. |
| Harness, Jellyfish, Swarmia | Leader dashboards on sessions, tokens, attribution | Read-only reporting for managers; no enforcement, no rejections, no per-team bar. |
| Vanta, Drata | AI-code control modules pulling GitHub evidence | Top-down compliance; they will need an evidence source for local agent work and Panout's receipts are that source. Channel, not competitor. |
Business model: metered per outcome evaluated, paid the same for yes and no
- Free: one repo, unlimited agents, forever. Recording, contracts, and the exception queue cost nothing. Distribution comes from the free tier being the best way to see what your agents did.
- Metered: a fee per outcome evaluated against a contract, from day one, identical whether the result is cleared or flagged. Evaluation is the billing event because it is countable, not because it is the deliverable: the customer's CI already runs the checks, and what the fee buys is the policy, meaning the rework map, the autonomy level per task class, the override record, and the calibration behind them. This is the round 2 fix: a referee paid per clearance granted is issuer-pays, which is the conflict of interest the whole thesis is built on avoiding. Paid per evaluation, Panout has no incentive to raise autonomy, and revenue does not wait on the day-45 autonomy milestone. Skin in the game comes from a published miss rate per team and a credit when a cleared class regresses. No per-seat pricing, no per-token pricing, as a charter commitment.
- Teams: flat per-repo tier for shared policy, sync, and permissions, for buyers who want predictable bills.
- Named, not built: attestation exports for audit, and diagnostic reports sold to harness vendors about their own agent's clearance rate on consenting teams. Timing note from the evidence: as of September 2026 no auditor, framework, or standard questionnaire asks for the AI-authored share of code, and EU AI Act Article 50 duties fall on model providers, not teams. Attestation is a 2027-28 budget. It is not in this plan's revenue.
Why this fixes the round 1 objection. Revenue tracks agent volume. A team of two humans and forty agents pays Linear $32 a month; the same team produces thousands of evaluable outcomes a month. Cursor's Bugbot already charges $1 to $1.50 per review run at this boundary and moved there from per-seat pricing in June 2026, so per-unit pricing is accepted by exactly this buyer. Prosumer conversion rates are irrelevant when the meter is on the agents.
$1B math, stated as a ramp with a measured denominator. The denominator is built from usage figures rather than one vendor's team count: Codex reports over 3 million weekly active developers, Copilot over 26 million users, Claude Code an estimated 4% of public commits (appendix B; the last two are secondary sources). Taking only the 3 million Codex weekly actives and a team size of six gives roughly 500,000 agent-using teams today, before Claude Code, Cursor, and Copilot are added. A second, independent denominator lands in the same band: Cursor's reported 7 million monthly actives at the same team size is about 1.2 million teams, and Copilot's 26 million users at a team size of ten is 2.6 million, so 50,000 paying teams is between 2% and 10% of any of them. The ramp: 2,000 teams evaluating 500 outcomes a month at a third of a dollar is $4M a year and is the month-12 test of whether the meter works, at 0.4% penetration of that denominator; 10,000 teams is $20M at 2%; 50,000 teams is $100M at 10% of today's population, and the 2028 population is larger by whatever multiple agent adoption grows. Merged agent PRs visible on public GitHub, an upper bound of about eight million cumulative, are a volume reference, not a market share. Every model improvement raises the evaluated count because teams let agents do more. Trust infrastructure trades at multiples that make $100M of metered revenue a $1B company: Chainguard raised at $3.5B on $40M of revenue growing seven-fold, LMArena at $1.7B four months after its first product. Stripe Radar is cited for its network prior (89% of cards it screens have been seen before), not for its pricing, which has moved to tiers.
Why now, and why not in 2028
Supported event surfaces (Claude Code hooks and OpenTelemetry, Codex OpenTelemetry) exist as of this year, so rejections are observable without private parsing. Agent volume per human is crossing the point where reading everything is impossible. Every incumbent has just committed to a position that disqualifies it as referee: GitHub to Copilot, Linear to hosted execution, vendors to their own agents. In 2028 the rejection corpora will exist somewhere; the question is whether they sit in one neutral place or are lost inside six vendors' logs.
Founder fit
Repo and terminal only. Zero new habits: the behavior the product instruments is committing, rejecting, and writing decisions, all of which the founder does daily and prolifically. Solo-buildable in the runway: the ledger and commit-time evaluation exist as a spike; the wedge is two contracts and a rejection adapter.
Plan against five months, with kill criteria
- Days 1 to 14. Wedge above, in this repo and one other. The git-only rework baseline is done (appendix A); the next measurement is whether local evidence (rejections, interrupts, contract results) predicts rework better than git alone, on the same commits. Kill if the founder stops committing with the check on, or the surfaced count does not fall.
- Days 15 to 45. Codex as second harness; ten-fault calibration spike; contracts on the two worst classes in three repos. Kill if catch rates do not separate contracts. The autonomy milestone is deliberately not here: a level 1 requires 30-day-old commit-time-captured commits, which cannot exist before 2026-10-05 if capture ships this week, so it moves to day 60 and is stated as moved rather than left to drift.
- Day 60. First autonomy level earned by at least one contract with a fault-injection catch rate above threshold and zero regressions, and the first predictive delta (git-only versus git plus commit-time features, with a bootstrap interval, both outcome classes present). Appendix D's power note says a 15-point difference at 6% exposure needs thousands of commits, so contracts must fire on roughly 30% of commits by design and the comparison is paired within repo; 300 commits at 6% exposure would be uninterpretable. Kill if no contract reaches level 1 or the delta is indistinguishable from zero.
- Days 45 to 90. Three external fleet operators from the watering hole running it, with rejection capture on. Publish the first number nobody has: the rejection rate of agent work by task class, across harnesses. Kill if fewer than two operators keep the check on for 14 days.
- Month 4 to 5. Continue only with external teams paying the meter or a fundable signal.
Evidence appendices
- Appendix A, local git backtest:
appendices/A-local-git-backtest.md(method, per-repo tables, caveats; git-only, rework is a proxy). - Appendix B, public evidence with URLs:
appendices/B-public-evidence.md(GitHub API measurements reproducible viagh, incident sources, DORA, GitClear, Veracode, vendor documentation for each incumbent, comparables; items the sweep could not verify are marked UNVERIFIED and are not relied on above except where labeled). - Appendix C, negative-class census:
appendices/C-negative-class-census.md(262 Claude Code and 95 Codex sessions, 21 days, anchored matching, per-session and per-project tables, statistical adequacy). - Appendix D, session-to-commit linkage and rework prediction:
appendices/D-session-commit-linkage.md. - Appendix E, written demand signals:
appendices/E-written-demand-signals.md. - Appendix F, restricted rework metric on the same commits:
appendices/F-restricted-rework-metric.md. - Appendix G, the rework map run on public agent-heavy repositories:
appendices/G-public-repo-rework-map.md. - Appendix H, split-half stability of the per-repo gap and the corrective metric on public repos:
appendices/H-split-half-stability.md.
What the evidence does not show, stated so nobody has to discover it. No external human has used Panout. A per-repo rework gap computed from history is not stable across time within the same repo (appendix H), so the day-one map is a diagnostic, not a product. Agent work is not reworked more than human work in general: across 11 public repos the pooled gap is 2.5 points with an interval including zero, and attribution from git is unreliable enough that two defensible rules give opposite signs (appendix G). The founder's gap is robust in direction across definitions and is between 11 and 23 points depending on how strictly rework is defined (appendix F); under the adopted strictest definition it is 11 points with one reversal at 44 agent commits. It is the founder's. In-session rejection is too rare on this machine (0.14% of tool calls) to certify an autonomy level for any task class; at that base rate the required sample is about 100 times what is on disk, which is why autonomy is derived from the rework label and contract catch rate instead. Logs cover 21 days, so no trend by model version can be claimed. Zero-review merge rates do not distinguish agent PRs from human PRs in popular repos; the distinguishing signal is merge latency. Nobody is asking for AI-code provenance yet. The rework measurement is line-identity, not semantics, and the founder's largest repos are team repos where the founder is a minority committer.