Panout review packOverviewFinal memoMVPLive outputsJudge roundsAppendicesDecisions

Source: docs/research/2026-09-03-defensibility/rounds/r2-judgment.md

Panout round 2 — investment judgment: "The clearance layer for agent labor"

Judge: Garry Tan lens. Methodology: gstack office-hours startup mode (six forcing questions, anti-sycophancy rules, end with an assignment) plus plan-ceo-review cognitive patterns (inversion reflex, proxy skepticism, focus as subtraction). Date: 2026-09-03. Dogfood gate: 2026-09-05 — two days away, unreported. Question scored: would I write the check today for a defensible, $1B-capable company, given this founder's constraints and this evidence? Calibration: a prior judge scored the best single round-1 idea 62 and the best round-1 combination 72. I read that report for scale only. My conclusions are my own and are checked against the new memo and both appendices.


Score: 66/100

Verdict: the first Panout document with a moat mechanism I would actually repeat out loud — rejected agent work is a signal that is destroyed rather than stored, and nobody else is standing where it dies. It is 66 and not 85 because the moat's raw material has never been counted, the number the memo quotes for it appears in neither appendix and is contradicted by the one you wrote, and the dated kill criterion you set for yourself lands in two days and this memo silently starts a new fourteen-day clock instead of reporting it.


Dimension table

#DimensionScoreReasoning
1Q1 Demand reality5Real progress: you now have behavior at population scale — 1.3-minute median open-to-merge on a 397-line mean diff across 150 merged Copilot agent PRs, against a 25-minute control. That is the first evidence in any Panout document that the target behavior exists outside your own head. But it is evidence that a behavior exists, not that demand exists. Zero external humans have used, asked for, or paid for anything. And the 14-day plan is the same shape as the 14-day plan set on 2026-08-22, which produced three research documents and no daily habit. A credible 14-day plan loses credibility the second time it is written.
2Q2 Status quo7Best-improved dimension and the reason the score moved. Positions A/B/C are named with costs: read-everything caps throughput at reading speed; read-nothing carries a measured 15.6-point (pooled) to 21.7-point (founder-only) rework gap with correction landing a 17-hour median later; second-agent review is text that produces more text. Bugbot's June 2026 move to $1.00–$1.50 per run is a genuine market signal that per-unit pricing at this boundary clears. Held at 7, not 9: rework is a proxy your own appendix disclaims ("a proxy for quality, not a measurement of it"), the human control is 49% — a metric where half the control fires is nearly saturated — and nobody has priced one incident in dollars.
3Q3 Desperate specificity5You finally have a consequence with a name attached to it: Amazon, March 2026, "high blast radius," ~6-hour retail outage, and a remediation that is your pitch inverted — every AI-assisted change now needs senior sign-off, which Amazon itself called "controlled friction." That is a real person being handed an unscalable order. But it is not your person: no Amazon engineering leader has spoken to you, Steinberger is a quote not a customer, and "the fleet operator one incident away from the same order" is a hypothesis about someone else's future. Also, your best incident is contested — Amazon disputes the framing and Fortune reported the root cause as an engineer acting on bad advice from an outdated wiki. Your evidence file says so. Your memo does not.
4Q4 Narrowest wedge8The sharpest wedge in any Panout document. panout init, two contracts, prints only failures at the commit boundary; acceptance is the commit, and committing after a printed failure is an override recorded as one. That is the real unlock — you found the zero-new-behavior form of consent. No login, no integration, works with any harness. Held at 8: day-one value is a linter (you cannot stop reading anything until a streak exists), the check sits in the human's commit loop which brushes your own observation-only invariant, and "48 hours" is an estimate from a founder whose last twelve days produced 15,000 words and no shipped check.
5Q6 Future fit8Strong and for a reason competitors structurally cannot claim: agent output per engineer is rising, human reading capacity is fixed, and every incumbent has just committed to a disqualifying position — GitHub to Copilot, Linear to metered hosted execution, each vendor to rating its own agent. One point held back for an inversion nobody in the memo names: your asset is rejections, and rejections are the thing that declines as models improve. The market grows and your raw material thins at the same time. See objection 5.
6Moat mechanism7The survivorship argument is the best new thinking in twelve days and it directly answers the prior round's strongest objection. Everyone computing quality from merged code is training on survivors and never sees the negative class; git can be re-mined in 2031, a diff that was reset before commit cannot. That is structurally non-reconstructible, which is the correct shape of a data moat. Layer 3 (cross-team priors, the Radar analogy) is the right second-order argument. Held at 7 because: the volume of the negative class is unmeasured; your own evidence file records that Claude Code's OTel already logs tool-permission denials locally, so the harness sees the rejection before you do and Anthropic could hold its own; and your pricing rebuilds the conflict of interest the moat is built on avoiding.
7Incumbent response6The table is honest and mostly right, and "GitHub's confidence becomes one more evidence input" is the correct posture rather than a claim of immunity. But you omitted the single most threatening line in your own appendix B: Copilot code review can now submit approving reviews that satisfy repository approval requirements (public preview). GitHub just stepped from commenting into the approval decision — the exact boundary you are claiming. Leaving that out of the incumbent table while citing the same page for rulesets is the kind of omission a partner finds in ten minutes.
8$1B credibility6Materially better than the prosumer-conversion story it replaces: metering on agent volume rather than human seats is the right structural answer, and it makes the sentence "revenue grows every time the models improve" true rather than decorative. But the arithmetic is 50,000 teams × ~$2,000/yr = $100M, and 50,000 is the entire measured population of Cursor's paying teams. Your $100M case requires 100% penetration of your own reference market. Also: your flagship comparable moved off per-transaction pricing — Radar is now $10–$70/month tiers — which weakens "per-transaction trust pricing is an accepted shape." And the eight-million-PR figure is an upper bound (your appendix says branch-prefix queries mix human commits and "counts across rows overlap") presented as a floor.
9Founder fit6On paper this is the best fit yet: repo, terminal, commits, decision records, multi-harness, no login, no new habits. Down from where the prior judge had it, because the revealed behavior of the last twelve days is now in evidence and it cuts the other way. Assigned two things on 2026-08-22 (ship the crappiest gate; run 10–15 non-pitch conversations) and one thing on 2026-09-03 (ask ten operators one question). Delivered: a Linear teardown, a competitive memo, five defensibility ideas, one sharpened idea, two appendices. Fifteen conversations is nine hours. You have had three hundred and thirty-six. The three invariants are respected; the fourth invariant — that this founder ships and talks to humans — is the one under test and currently failing.
10Evidence integrity5Genuinely split, so a middling score rather than a bad one. Credit, and real credit: you ran Sample E and let the control kill your own headline ("zero-review merge rates are not distinctive"); you reported that nobody asks for AI-code provenance and that EU AI Act duties fall on providers, killing your own 2027 revenue line; you tagged UNVERIFIED items and mostly kept them out; you wrote a "what the evidence does not show" section. Very few founder memos do that. Against it, and it is serious: the load-bearing rejection numbers are unsourced and internally contradicted (see objection 2); "11 founder-controlled repos" describes two repos your own appendix labels multi-author team repos where you are a minority committer; "27% (dotfiles) to 63% (a gateway service)" cherry-picks the top of a range that actually starts at 0.6%; "roughly 2,500 agent commits a month each" is 2× wrong (the real figures are ~1,264 and ~1,195/month; 2,458 is the combined total); and the pooled rework gap reverses sign in the two largest repos, undisclosed.

Raw: 63/100.

Gut adjustment: +3. Two things are worth paying for and neither is in the rubric. First, the survivorship-of-the-negative-class argument is the first genuinely non-reconstructible asset any Panout document has proposed — everything before it was a time series, and a calendar is not a moat. Second, acceptance is the commit; committing after a printed failure is an override is a real product insight: it converts consent into an act the user already performs, which is the only thing that has ever made this category work. I held the adjustment to +3 rather than +8 because of the 09-05 silence. A founder who lets a self-imposed dated kill criterion pass without reporting it has told me something about how the next four gates will go.

Final: 66/100.


The two sentences, verified

Moat, as written

Panout sits at the only boundary where agent work that a human rejected can be seen: the local commit boundary, before GitHub, Linear, or the harness vendor ever learns it existed. GitHub sees only survivors; vendors see only their own agent; Panout is the sole holder of each team's rejections, and rejections are what calibrate trust.

Would I repeat this to a partner? Not as written — two words are false and a partner will find both. "The only boundary" is wrong: the harness sees the rejection first. Your own appendix B documents that Claude Code's OpenTelemetry stream records tool-permission decisions and who made them. Anthropic is standing upstream of you with the data already flowing. "Sole holder" is aspiration wearing the clothes of a fact. Your real claim is narrower and still good: cross-vendor, and durable because nobody is storing it.

Rewrite:

Rejected agent work is the only quality signal in software that gets destroyed instead of stored: GitHub and every analytics vendor compute from merged survivors, and each harness sees only its own agent's refusals. Panout records the negative class at the local boundary across every harness, and a trust model trained on rejections cannot be rebuilt later by anyone who was not capturing them at the time.

$1B, as written

Every unit of agent labor needs a clearance decision, humans cannot make them by reading, and the party that clears them is paid per decision, so revenue scales with agent volume rather than human seats. Stripe Radar for software labor: per-transaction pricing on a cross-customer trust model that gets better with every customer.

Would I repeat this? The first half yes, the second half no. "Paid per decision" is the sentence a partner turns into a weapon: you are the referee, and you are paid for each thing you tell the human they no longer need to read. That is issuer-pays. And "Stripe Radar: per-transaction pricing" is now factually stale — Radar is subscription-tiered at $10/$14/$20. Cite Radar for the network prior, which is the part that is actually your argument, and drop the pricing analogy.

Rewrite:

Agent output per engineer is now growing faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than a fee per human seat. Ten thousand teams already paying for agent fleets at a third of a dollar per evaluated outcome is a $20M business, fifty thousand is $100M, and the meter compounds automatically every time the models get good enough to be trusted with more — which is the only revenue line in developer tooling that gets stronger as the models improve.

Note what I changed and why: evaluated, not cleared. It removes the conflict of interest, it starts the meter on day one instead of day forty-five, and it does not require you to grant autonomy in order to get paid.


The single hardest push

Your entire company is one dataset: the negative class. Everything else in this memo — the earned autonomy levels, the cross-team priors, the Radar analogy, the reason GitHub cannot follow you — is downstream of one claim, which is that rejected agent work exists in recordable volume at the local boundary and that you are the one recording it. Good. That is the right thing to build a company on, and it is the first time in three weeks of documents that you have proposed an asset a well-funded copier cannot buy.

So tell me how big it is.

You wrote: "26 rejections and 27 interrupts in 30 days on this machine across 245 Claude Code sessions." That number appears in neither appendix. Not in A, not in B. And appendix A, which you wrote today, section 6, says the opposite: eight session JSONL files total, all dated 2026-08-30, all under /tmp/jones-agent-guidance/, zero in any directory corresponding to a /code repo, ~/.claude/sessions/ empty, and — in your own words — "the local session store has been pruned or rotated and cannot corroborate the git-trailer signal." You measured your session store, found it empty, wrote that down honestly, and then quoted a session count from somewhere else in the memo the appendix is supposed to support. A partner reads appendix A section 6 and the memo's fifth paragraph and the diligence call ends there. Not because you lied — I do not think you did — but because the one number that decides whether the company exists was not measured, and you did not notice that you had not measured it.

Now take the number at face value and it is worse. Fifty-three negative events in thirty days. You are, plausibly, in the top thousandth of a percent of agent users on earth — 245 sessions a month, sixty-four percent of your commits agent-authored — and your negative class is fifty-three events. Split that across task classes and you have single digits per class per month. Stripe Radar's prior is thick because 89% of cards presented have been seen before across seventy trillion data points. Yours would be a dozen refusals a month per class from the heaviest user alive. And it gets thinner: your Q6 argument is that models keep improving and teams let agents do more, which is precisely the condition under which humans reject less. Your market expands and your raw material depletes on the same curve. Radar's fraudsters fight back; yours retire.

And then the pricing. You built the whole thesis on conflict of interest — GitHub cannot referee because it owns Copilot, Linear cannot referee because it bills the credits, no vendor can rate its own work. Correct, and it is your best structural argument. Then you priced yourself per read avoided. You are paid, per unit, for granting autonomy. Moody's got paid by the issuer, per rating, for issuing the rating, and that arrangement survived exactly as long as nothing went wrong. Yours is worse, because the customer buying the clearance is also the party that gets hurt by a bad one, so the day your level-3 class ships an incident you are simultaneously the vendor, the referee and the cause. You cannot sell neutrality on the front page and issuer-pays on the pricing page. Somebody will read both.

You did not need users to find any of this out. You needed one afternoon with jq over the event surfaces you already have on disk, and you spent it writing your third strategy document in twelve days. Two weeks ago you were told to talk to fifteen people. Twelve days ago the whole gate rested on a habit you were supposed to be forming. In two days a kill criterion you wrote yourself and dated yourself comes due, and this memo does not mention it — it opens a fresh "Days 1 to 14." That is the exact behavior the criterion existed to prevent. You wrote in STRATEGY.md, unprompted, that you were writing the gate down "now so month-5 us can't move the goalposts." Month-5 you did not move them. Day-12 you did.

Count the rejections. If the answer is fifty-three, say so out loud and tell me why a company survives on it.


Every remaining objection keeping this below 90, in priority order

1. The 2026-09-05 gate comes due in two days and this memo does not report it; day one has been silently reset. STRATEGY.md: "2026-09-05 (day ~14 of dogfood): the brief has not recovered work, prevented overlap, or changed the next action → a readable log is not a product; narrow or stop. The test is behavior, not intention." The memo's plan reads "Days 1 to 14" with a new day one. Fix: Not a memo fix — evidence. By 2026-09-05, publish the gate result as a pass or a fail with the specific behavior that did or did not happen. A failed gate honestly reported costs you nothing with me; a gate quietly re-dated costs you the whole meeting. If it failed, the next document is one page: "the gate failed, here is what I am narrowing to."

2. The rejection volume — the moat's raw material — is unsourced in the memo and contradicted by your own appendix. "245 Claude Code sessions… 26 rejections and 27 interrupts" appears in neither appendix; appendix A §6 measured eight session files, none mapping to a /code repo, and states the store cannot corroborate anything. Fix: Evidence, one day of work, no permission required. Count, from the surfaces actually on disk (Claude Code OTel/hook records, Codex denial events, git reflog/stash/reset-discarded diffs), over 30 days: total negative events, split by type, split by task class, with n per class. Publish the count even if it is embarrassing. Then state the minimum n per class at which an autonomy level is statistically meaningful, and compare. This single number decides whether layers 1–3 exist.

3. Zero external humans, thirteen days after the assignment. Population evidence is not demand evidence. The memo defers external operators to days 45–90, which puts the first non-founder human at roughly month four of a five-month runway. Fix: Evidence only. Ten fleet operators, in writing: how many agent-authored diffs did you merge last week without reading them, and did any of them bite you. Two numbers and a yes/no. Nine hours. It has now been outstanding for twelve days and it prices this entire memo.

4. The pricing recreates the conflict of interest the moat is built on. Paid per read avoided = paid per clearance granted = issuer-pays. It also delays all revenue past the day-45 kill criterion, so the paid product depends on the riskiest unproven step in the plan. Fix: Idea fix, one line. Meter per outcome evaluated, not per outcome cleared. Revenue starts at level 0 on day one, the referee is paid identically whether it says yes or no, and if you want to keep skin in the game add a published miss rate plus a credit when a cleared class regresses. Say the words "we are paid the same whether we clear it or flag it" on the pricing page.

5. The negative class depletes as the market grows. Better models → fewer rejections → thinner corpus, on exactly the curve that makes the market big. The memo's Q6 argument and its Layer 1 argument are in tension and neither notices. Fix: Memo fix plus a measurement. Name the tension, then state the hedge: the durable asset is not raw rejection count but contract catch-rate per fault class and task-signature-to-outcome mappings, which survive model improvement because the fault distribution shifts rather than vanishes. Then measure it: rejection rate per session, by model version, over the last six months. If it is flat or rising, you have a much stronger memo. If it is falling steeply, you have a different company.

6. The headline rework gap reverses sign in your two largest repos, undisclosed. The two team repos supply 64.5% of the agent commits in the rework sample but only 7.9% of the human commits, and in both of them agent work is reworked less than human work (68.7% vs 86.3%; 66.6% vs 84.9%). The pooled 65-vs-49 is a mix effect. The clean number is the within-founder slice, 59.5% vs 37.8%. Fix: Memo fix. Lead with the founder-only 60-vs-38, disclose the reversal in the team repos in the same breath, and drop the pooled figure to a footnote. Also fix "roughly 2,500 agent commits a month each" — it is ~1,264 and ~1,195; 2,458 is the combined total. A 2× error in the one sentence that says "nobody is reading that" is the sentence a partner recomputes.

7. Selective framing of the local measurement. "11 founder-controlled repos" includes two repos your appendix labels multi-author team repos where you are a minority committer. "27% (dotfiles) to 63% (a gateway service)" omits work-brain at 0.6%, app-platform-feedback at 0%, jones at 7.9%, crossword at 6.5% and poe-browser at 16.1%. Fix: Memo fix. "Across nine repos I build alone the agent share runs 0% to 63%, concentrated in the newest and smallest; in two team repos it is 85% and I am a minority committer." That sentence is more interesting than the one you wrote, because it says agent-majority is a property of greenfield, and greenfield is where your wedge installs in 48 hours.

8. GitHub's move into the approval decision is omitted from the incumbent table. Appendix B: Copilot code review can now submit approving reviews that satisfy repository approval requirements (public preview). That is GitHub taking a position on the clearance decision itself, and you cite the same doc page for rulesets. Fix: Memo fix. Name it in the table and answer it with a switching cost, not with neutrality. Neutrality is a claim; the switching cost is "leaving means reading everything again, because the policy is derived from a rejection history nobody else holds." That sentence is already in your Layer 2. Move it to the GitHub row.

9. The $1B arithmetic needs full penetration of its own reference population. 50,000 teams × $2,000 = $100M, and 50,000 is the entire measured count of Cursor's paying teams. The 8M-PR figure is an upper bound (overlapping rows, branch-prefix queries include human commits) presented as a floor. Radar has moved off per-transaction pricing. Fix: Memo fix. State a penetration ramp with a de-duplicated population and a defensible 2028 denominator, drop the "measured, not imagined" claim about Cursor's secondary-sourced revenue, cite Radar for the network prior only, and label the PR count an upper bound the way your appendix does.

10. The measurement that decides whether the moat has predictive value is still not done. Does local evidence (rejections, interrupts, contract results) predict rework better than git history alone, on the same commits? The memo lists this as "the next measurement." It was the previous round's central objection and it is still open. Without it, "clearance" is a dashboard with a good story. Fix: Evidence. Same commit set as appendix A, git-only baseline versus baseline-plus-local-evidence, one stated delta. Gated on objection 2 — you cannot run it until the negative class is counted.

11. The strongest incident is contested and the memo omits the dispute. Amazon disputes the framing; Fortune reported the root cause as an engineer acting on inaccurate advice inferred from an outdated wiki; the sourcing is secondary reporting attributed to FT and CNBC. Fix: Memo fix, one clause. "Amazon disputes the framing, which is itself the point: the remediation stands either way." Volunteering the dispute makes the citation stronger, not weaker.

12. Entire is treated as a known product though your evidence file could not find it. Appendix B: "UNVERIFIED — I could not find any product named 'Entire' with experts or judge features in public search." Fix: Memo fix. Drop the row, or relabel it as an unverified private tool. You wrote the discipline into the appendix; apply it in the memo.


One assignment for this week

Count the negative class. By 2026-09-05, publish one table: negative events over the last 30 days, from surfaces actually on disk, split by type — human rejection of a proposed action, interrupt, permission denial, diff discarded by reset/checkout/stash, commit reverted, commit reworked within 30 days — and split by the two task classes you repeat weekly. Report n per class. Then state, in one line, the minimum n per class at which an autonomy level would mean anything, and whether you are above or below it.

Nothing else. No new memo. No fourth strategy document. This is a day with jq and git reflog on data you already have, it requires no user and no permission, and it is the cheapest available thing that can kill the company — which is exactly why it goes first. If the negative class is thick and separates by class, your moat sentence writes itself and I will take a second meeting on the strength of that one table. If it is fifty-three events, you have learned in one day that the asset is not there, and you will have five months left instead of four to do something about it.

Two things I am not letting you replace with this. The 2026-09-05 gate result is still due on 2026-09-05 — publish it as a pass or a fail. And the ten operators, one question, in writing, remain outstanding from twelve days ago; that assignment does not expire and it does not get renegotiated into a better one. A rejection corpus with only your own repo in it is not a corpus. It is a diary.