Panout defensibility ideas, round 1 — investment judgment
Judge: Garry Tan lens, gstack office-hours methodology (startup mode, six forcing questions, anti-sycophancy rules) plus plan-ceo-review cognitive patterns. Date: 2026-09-03. Dogfood gate: 2026-09-05 (two days out, unresolved). Question scored: would I write the check today for a defensible, $1B-capable company, given this founder's constraints and evidence?
Opening position, before the individual scores
Two weeks ago you were assigned two things: ship the crappiest possible gate, and have 10 to 15 non-pitch conversations. Since then you have produced a competitive teardown of Linear, a product-overlap map, a decision log with nine entries, and now five defensibility memos. The 2026-09-02 doc says the conversations are "already planned." Planned. Fifteen calls is nine hours of work and you have had three hundred and thirty-six.
So the honest read on all five ideas: the variance between them is smaller than the variance between "any of them, with one external human attached" and "any of them, with none." You are optimizing the wrong term. Every score below is capped by that, and I am not going to pretend otherwise to make the memo feel productive.
Second position. Four of these five are the 2026-09-02 recommended business — decision, contract, acceptance, earned autonomy, routing — decomposed and re-labeled. That is not worthless; decomposition is how you find the load-bearing member. But you should know which memo you are reading. Ideas 4 and 1's wedges are the 09-02 doc restated. Idea 3 is the 09-02 doc's explicitly deferred item ("name the enterprise budget without building for it") promoted to the lead. Idea 5 is STRATEGY.md's "canonical index — unproven prize" with a business model bolted on. Only Idea 2 and Idea 1's moat mechanism are new thinking.
Third. Nobody in this set names a human. "The solo founder or 2-5 person team running 5-40 sessions a day" is a filter with a number in it, not a person. You cannot email a range. The one real user in these documents is you, and office hours #1 recorded you saying you were not sold. Two weeks of memos did not change that; it deferred it.
Idea 1: The underwriter of agent labor
Score: 62/100. Verdict: the only one with a structural moat argument I would repeat to a partner, and its data asset is half-owned by GitHub already.
| Dim | Score | Reasoning |
|---|---|---|
| Q1 Demand reality | 4 | Zero external behavior. Founder is the user and has been burned — that counts for something, and the 14-day exception-queue test is genuinely falsifiable. But N=1, and the same founder failed the same shape of test in v1. |
| Q2 Status quo | 5 | git diff --stat, skim, merge — specific. The cost is not: "incidents," "silent drift," "growing dread." Nobody has priced a single incident. Dread is not a line item. |
| Q3 Desperate specificity | 4 | A range with a session count, not a name. No career consequence for anyone but the founder. |
| Q4 Narrowest wedge | 7 | Days, one repo, terminal, no login, no integration. Real. But the claimed day-one value ("you stop reading the classes that pass") is false on day one — you cannot stop reading until a streak exists. Day one you get a log. |
| Q6 Future fit | 8 | More agent volume makes reading-everything more impossible, and provenance demand rises on a regulatory clock. The reason a competitor cannot claim it — "you may not rate your own work" — is real. |
| Moat mechanism | 6 | Three structural claims stacked: conflict-of-interest neutrality, data that only accrues over time, potential format standard. Not "we move fast." But see the push. |
| Incumbent response | 5 | Survives Anthropic/OpenAI/Cursor/Linear on neutrality. Does not obviously survive GitHub, which owns the merge and the git history the loss labels are computed from. |
| $1B credibility | 5 | Moody's / Verisk / Sigstore+Chainguard are legible comparables. The chain has four legs with four different buyers, and your own strategy doc prices the middle leg at 70-100k free users for 2,000 paid. |
| Founder fit | 8 | Terminal, repo, decision-writing, multi-harness, no login. Respects invariants 1, 3, 4, 5. Commit-time evaluation brushes invariant 2 — it is not in the agent's loop, but it is in the human's. Say so explicitly before someone else does. |
| Novelty vs 09-02 | 7 | Wedge is restatement. Late-binding loss labels and the rating-agency structure are new and are the best new idea in the document. |
Raw 59. Gut +3: this is the only idea where the wedge, the moat, and the founder's actual daily behavior are the same object. That alignment is rare and I pay for it.
Hardest push. Your moat is "late-binding loss labels — only Panout can join pre-merge evidence to the accepted outcome to what happened next." Take that apart. The second half of that join — reverted, fixed forward, reworked within 30 days touching the same lines — is computable from public git history alone, and GitHub has it for every repository on earth without asking anyone's permission. So your exclusive contribution is the pre-merge half: contract results and session metadata. Which means your entire moat rests on an unproven empirical claim: that session metadata materially improves the prediction of rework over what git history alone already predicts. You have not measured that. You could measure it this week on your own repos without a single user, and you wrote five memos instead. If session metadata adds two points of AUC over a git-only baseline, you do not have a rating agency, you have a dashboard with a good story. And the actuarial framing needs teams times task classes times 30-day windows to reach significance — with five months of runway and one repo, you are not years from a calibrated probability, you are one order of magnitude of users away from the first honest confidence interval. Show me the backtest, then say "underwriter."
To 90+: (a) A git-only baseline versus baseline-plus-session-evidence rework-prediction backtest on 90 days of your own repos, with a stated AUC delta. If the delta is large, the moat sentence writes itself and this jumps 15 points on its own. (b) Three external repos running the acceptance layer with the exception queue answered daily for 14 consecutive days, logged. (c) One named human, title and company, who merged agent work unread, got burned, and said out loud they would pay to not repeat it. (d) A stated answer to "what happens when GitHub ships AI review confidence" that is not "we are neutral" — neutrality is a claim, not a barrier; name the switching cost.
Idea 3: Work receipts, the SBOM for agent labor
Score: 48/100. Verdict: the best buyer and the best regulatory clock in the set, attached to a sales motion this founder will not run.
| Dim | Score | Reasoning |
|---|---|---|
| Q1 Demand reality | 4 | Zero evidence, but uniquely cheap to get: 20 emails to heads of engineering asking "has a customer asked you what fraction of your code is AI-written?" is a two-day test with a yes/no answer. |
| Q2 Status quo | 6 | Best in the set. A policy document nobody follows, plus a spreadsheet. Cost is denominated in audit findings and delayed deals, which is actual money, not dread. |
| Q3 Desperate specificity | 6 | Best in the set. Title named, consequence named, and the consequence lands on a specific person's quarter. Still no actual name. |
| Q4 Narrowest wedge | 6 | panout attest is days of work, git trailer, no login. But it is parasitic: a receipt attesting to checks that do not exist yet is a receipt for nothing. It needs Idea 1's ledger underneath it. |
| Q6 Future fit | 8 | EU AI Act phases through 2026-27; the questionnaire line item only hardens. Genuinely more essential in three years. |
| Moat mechanism | 5 | Format adoption is a coin flip, and SLSA, Sigstore, in-toto, SPDX and CycloneDX already exist as the obvious hosts for an authorship extension. Note that the companies that monetized SBOM monetized curation, not the format. |
| Incumbent response | 4 | Weakest in the set. "GitHub could issue receipts for Copilot only" is wrong — GitHub Actions already signs build provenance and can attest to anything crossing its API. Vanta and Drata will ship an AI-code control module and pull evidence via the GitHub API, top-down, into a budget they already own. |
| $1B credibility | 6 | Best in the set. Vanta, Drata, Chainguard. Compliance sells top-down at high prices. |
| Founder fit | 2 | The memo admits it. Enterprise sales-led motion, no daily use for the founder, contradicts every piece of revealed behavior. Solo, five months. |
| Novelty vs 09-02 | 5 | Reverses a deliberate 09-02 decision rather than adding a mechanism. That is a legitimate move, but it is a reversal, not novelty. |
Raw 52. Gut -4: founder fit is not a dimension you average away. A solo technical founder with five months of cash running a compliance sales cycle whose median length exceeds his runway is a near-certain zero regardless of how good the market is.
Hardest push. You wrote the kill criterion yourself and then ranked this third anyway: "Founder fit is poor: this is a sales-led enterprise motion with no daily use for the founder." Read that back. You are a solo founder with roughly twenty weeks of money who does not log into trackers, and you are proposing to sell audit artifacts to Series B heads of engineering in fintech, where the procurement cycle alone is twelve to sixteen weeks and the first three calls are with someone who cannot sign. You will spend your remaining runway learning a motion you have never run, for a buyer you have never met, against Vanta and Drata, who already have the security questionnaire in front of that exact person every quarter. The correct use of this idea is not a company, it is a sentence in the deck: "the enterprise budget for this is compliance attestation, and here is the artifact we already produce." Name the budget, do not chase it. That is what your own 09-02 doc said, and it was right.
To 90+: Not reachable by this founder in this configuration, and I would rather you not try. What would move it: a co-founder who has closed six-figure compliance deals, plus ten heads of engineering confirming they were asked the AI-provenance question by a customer (not an auditor) in the last 90 days, plus a signed design-partner LOI at a real price. That is a different company with a different cap table.
Idea 4: The re-litigation firewall for decisions
Score: 47/100. Verdict: the best wedge in the set and the thinnest moat; ship it, do not pitch it.
| Dim | Score | Reasoning |
|---|---|---|
| Q1 Demand reality | 5 | Highest in the set on plausible near-term behavior, because the pain is near-universal among agent users this week and your own repo already generates the artifact. Still zero external evidence. |
| Q2 Status quo | 5 | Workaround specific and real: AGENTS.md prose that agents read and forget. But the per-incident cost is annoyance, not incident. People will not switch tools over annoyance; they will grumble and re-fix the dependency. |
| Q3 Desperate specificity | 4 | "Any developer who has watched an agent re-introduce a dependency" is deliberately broad. Broad and mild. No career consequence. |
| Q4 Narrowest wedge | 9 | Best in the set by a distance. Ships in a day or two. Works with every harness with zero integration because they all read repo instruction files at session start. No login. Prints one line. |
| Q6 Future fit | 4 | This is where it dies. Claude auto memory and Copilot Memory already learn from corrections — your own overlap doc says so. In three years the harness does this natively and for free. Less essential, not more. |
| Moat mechanism | 2 | The memo says "thin alone" and undersells how thin. Per-repo, portable, plain-text corpus means zero switching cost and zero network effect. It is a markdown convention plus a lint check. |
| Incumbent response | 3 | Absorbed as a default, not fought. Every harness ships rules files and memory. |
| $1B credibility | 2 | Memo admits $100-300M ceiling. Standards make big companies only when paired with a hosted service, and you know it. |
| Founder fit | 9 | Highest in the set. This is literally what you already do by hand. |
| Novelty vs 09-02 | 2 | Objects 1 and 2 of the 09-02 recommended business, restated with a directory name. |
Raw 45. Gut +2: because it is the shortest path from here to the behavioral evidence that every other idea in this document is waiting on, and shipping it costs 48 hours.
Hardest push. ADRs have failed for twenty years and your explanation for why they will work now is "an agent can draft one in seconds and every harness reads repo files." Half of that is a claim about writing, which was never the binding constraint — the binding constraint was that nobody read them and nothing enforced them. Your enforcement is a commit-time check that prints a violation. What happens the fourth time it prints a violation you disagree with at 11pm on a Friday? You add the exception, and then the file is prose again. That is the exact failure mode of every lint rule anyone has ever bypassed, and you have no answer for it except that it will be you doing the bypassing. Meanwhile your own kill criterion — "agents comply from AGENTS.md prose alone, making enforcement redundant" — is the likely outcome, because model instruction-following got materially better across 2026 and will keep going. You are betting a company on models staying bad at something they are visibly getting good at. That is the wrong side of a rising tide.
To 90+: Not as a company. As a wedge, what would move the combination it feeds: instrument the check for 14 days and publish the count — how many diffs violated a recorded decision, out of how many. If that number is under 3%, your kill criterion fired and you should say so in public. If it is over 15%, you have the first real number anyone has published about agent decision drift, and that number is a distribution asset in the watering hole you already identified.
Idea 2: Mutation-calibrated acceptance contracts
Score: 46/100. Verdict: the most original thinking in the document and the most honest self-assessment; it is a component, and it knows it.
| Dim | Score | Reasoning |
|---|---|---|
| Q1 Demand reality | 3 | Nobody has ever asked for a catch rate. This is a founder's pain about his own product's credibility, wearing a user's clothes. |
| Q2 Status quo | 4 | Tests, CI, second-agent review. The cost of not knowing your catch rate is invisible to the person paying it, which is the definition of a problem people do not buy solutions to. |
| Q3 Desperate specificity | 3 | "The same fleet operator, at the moment they ask whether they can trust the check." That moment is a hypothesis about an internal monologue. |
| Q4 Narrowest wedge | 6 | Ships in days, background, no login. But the deliverable is a number, and numbers do not retain. Also it cannot ship first: it needs a corpus of accepted diffs, which requires Idea 1 or 4 running already. |
| Q6 Future fit | 6 | More agent volume means more need to calibrate. Against that: your fault library depreciates every time a model ships, so this is a treadmill, not an accruing asset. |
| Moat mechanism | 5 | "Vendors will not catalog their own failure modes" is true and clever. But academics and independent researchers will, for free and for citations, and mutation-testing literature is already public. |
| Incumbent response | 5 | Survives, mostly because nobody wants to build it. Surviving by being unwanted is not the same as being defensible. |
| $1B credibility | 2 | Memo says $50-200M standalone. Correct, and disqualifying as a lead. |
| Founder fit | 7 | Local, background, terminal, engineering-heavy in a way that suits you. But it is weeks of build in a twenty-week window. |
| Novelty vs 09-02 | 8 | Highest in the set. It directly defuses the 09-02 day-45 kill criterion ("contracts are too weak and the product is a linter"), which is the single most likely way the recommended business dies. |
Raw 49. Gut -3: dependency ordering is wrong (it needs another idea shipped first), the fault library depreciates, and as a standalone it consumes runway to produce a metric rather than a customer.
Hardest push. The reason this is interesting is that it defeats the "you are a linter" objection, and the reason it is dangerous is that it is the most intellectually satisfying thing in the document. You are a founder with no users who has spent two weeks writing analysis, and you have just invented a project that consists entirely of measuring your own instrument. That is the most seductive possible way to spend a month not talking to anyone. Also test your own premise: if the mutation results come back showing every contract catches 80% of injected faults, you have learned nothing and shipped a vanity number; if they come back at 15% across the board, you have proven your product does not work and you will have spent three weeks proving it. Only the middle case — where catch rates separate contracts and the separation predicts real rework — is worth anything, and you can find out which world you are in with a two-day spike on ten faults, not a fault library. Do the two-day version or do not do it.
To 90+: Not as a standalone; the ceiling is stated and I believe it. What raises its contribution: a two-day, ten-fault spike showing catch rates that both separate across contracts and correlate with observed rework in your own git history. That correlation is the whole ballgame — without it this is mutation testing with better marketing.
Idea 5: The neutral rework index across harnesses
Score: 38/100. Verdict: step three of a plan whose steps one and two do not exist, and naming it now is an active distraction.
| Dim | Score | Reasoning |
|---|---|---|
| Q1 Demand reality | 2 | No evidence, and the paying party (vendors) cannot be sold until you have data you do not have. |
| Q2 Status quo | 4 | Benchmarks are gamed and everyone says so. But nobody's job depends on fixing it — leads pick harnesses on vibes and their quarter still closes. |
| Q3 Desperate specificity | 3 | Two personas, both weak. A lead justifying a budget is mildly inconvenienced. A vendor wanting a number is a customer, not a desperate user, and only after you are already credible. |
| Q4 Narrowest wedge | 2 | Fails outright. Cannot ship in days with value. Requires hundreds of consenting repos before the first useful output exists. |
| Q6 Future fit | 7 | Benchmark saturation is real and the need grows. So does everyone else's ability to serve it. |
| Moat mechanism | 7 | Strongest in the set if reached. Corpus plus methodology credibility plus a neutrality charter is genuinely winner-take-most. Trust concentrates. |
| Incumbent response | 6 | Survives agent vendors on conflict of interest. Does not survive an academic consortium, an LMArena-style nonprofit, or Harness and Jellyfish, who already have the enterprise telemetry contracts and the sales channel. |
| $1B credibility | 4 | Gartner is legible but took decades. Chatbot Arena is the closer analogue on mechanism and has approximately no revenue. Neutral index economics are famously hard to convert. |
| Founder fit | 4 | Requires distribution, consent, methodology review and press before any revenue. Solo, five months. Wrong ordering at every level. |
| Novelty vs 09-02 | 4 | STRATEGY.md already lists "canonical index — unproven prize." This is an elaboration with pricing attached. |
Raw 43. Gut -5: this is the highest-status idea in the document and the one most likely to eat a month. It requires the thing you do not have (users) to produce the thing that makes it defensible (corpus), and there is no version that works at N=1.
Hardest push. Every founder with no distribution eventually invents the index, because the index is the version of the business where you are important and nobody can compete with you, and it always sits exactly one impossible step past where you actually are. You need "a few hundred repos contributing" — your own kill criterion — and you currently have one, which is yours, and you have not yet passed its 14-day test. Worse, your neutrality is not a barrier, it is a temporary vacancy: the moment this matters, a foundation or a university lab publishes the same index for free with better methodology credibility than a venture-backed startup can ever claim, and your paying customers are the vendors you are rating, which is the conflict you built the whole company to avoid. Moody's got away with issuer-pays because regulation mandated ratings. Nothing mandates yours. Delete this from the deck until you have three hundred repos, and if you ever get three hundred repos you will have a better idea by then.
To 90+: Structurally unreachable pre-distribution. It becomes gradeable at a few hundred consenting repos with a published methodology that survives outside review, and at that point it is a follow-on, not an idea to score.
Ranking
- Idea 1 — Underwriter of agent labor, 62. Only one with a moat sentence I would repeat out loud.
- Idea 3 — Work receipts, 48. Best market, wrong founder, wrong runway.
- Idea 4 — Re-litigation firewall, 47. Best wedge, no moat. Ship it as a wedge, never pitch it as the company.
- Idea 2 — Mutation-calibrated contracts, 46. Best new mechanism, smallest ceiling, wrong to do first.
- Idea 5 — Neutral rework index, 38. Correct destination, fantasy sequencing.
Note what this ranking says: the two ideas with the best markets (3, 5) are the two the founder cannot execute, and the two the founder can execute best (4, 2) have the weakest moats. Idea 1 is ranked first because it is the only one that resolves that tension rather than picking a side.
The combination that beats any single idea
Idea 4 as the wedge, Idea 1 as the moat, Idea 2 as the proof, Ideas 3 and 5 named in the deck and not built. Estimated 72/100.
The mechanism, in the order the pieces have to arrive:
- Idea 4 ships in 48 hours and produces daily human behavior with no login and no integration, which is the only thing that raises Q1 and Q4 for everything downstream. It is not the company; it is the reason you have data at all.
- Idea 2 as a two-day, ten-fault spike defeats the single most likely cause of death — the day-45 "your contracts are theater, you are a linter" finding — and it does it with a measurement rather than an argument.
- Idea 1 supplies the compounding asset and the only credible $1B story, because loss labels are the one thing in this document that a well-funded copier cannot get by copying the interface.
- Idea 3 is one sentence in the deck ("the enterprise budget for this is attestation of agent-written code, and the receipt is a byproduct of the ledger we already keep"). Zero build.
- Idea 5 is one sentence in the vision slide, gated on three hundred repos.
Why 72 and not higher: the combination fixes wedge, novelty, kill-criterion resilience and $1B legibility, and fixes nothing about Q1, Q3 or the GitHub problem. Those are not fixed by better arrangement of memos. They are fixed by two humans who are not you.
The three objections keeping the best idea below 90
- Q1 is N=1, and the same founder already failed the same shape of test once. Zero external humans have done anything. Office hours #1 recorded "not sold myself" on 2026-08-22 and the resolution mechanism was a 14-day behavioral gate that lands in two days and has produced three strategy documents instead of a daily habit. Until an external repo answers an exception queue for fourteen consecutive days, all five of these are theses.
- The loss-label moat is half-owned by GitHub and unmeasured. The post-merge half of the join — revert, fix-forward, rework within 30 days on the same lines — is public git data at planetary scale, and GitHub also owns the merge boundary. Your differentiator is the pre-merge session evidence, and you have never tested whether it adds predictive power over a git-only baseline. This is a one-day backtest on data you already have, and it is the difference between a rating agency and a dashboard.
- The middle of the $1B chain is a business your own strategy doc calls brutal. Free repo to per-repo team sync requires 70-100k free users for 2,000 paid, by your own arithmetic. The leg after it (attestation) requires an enterprise sales motion you do not run. So the path from wedge to a billion currently routes through two segments the founder is on record as unable to serve. Either name the person who serves them, or find the version where the compounding asset itself is the product sold to the party that already has budget.
The assignment
One action, this week, not a strategy.
By 2026-09-10, get ten fleet operators to answer one question in writing: "How many agent-authored diffs did you merge last week without reading them, and did any of them bite you?" Two numbers and a yes/no, per person. DM them in the watering hole you identified two weeks ago — gstack forks, gstack issues, OpenClaw users. Do not pitch. Do not describe Panout. Do not send a link. Record the answers verbatim in the repo.
This is the number that prices all five ideas at once, and nobody on earth has published it:
- If the honest answer is mostly zero — people read everything — then there is no market for reading less, and Ideas 1, 2 and 4 are all dead in one afternoon.
- If people merge unread and nothing has ever bitten them, there is also no market, because the status quo is working and dread is not a purchase order.
- The company exists only inside the band where operators merge unread and get burned, and every score above is really a bet on the width of that band.
You have written roughly fifteen thousand words about defensibility against a market whose central quantity you have not measured and could measure in nine hours of conversation. Do the nine hours. Then come back and we will talk about which of these five it is — and if the band is wide, Idea 1 is the answer and you will be able to say so in two sentences without hedging.