Panout review packOverviewFinal memoMVPLive outputsJudge roundsAppendicesDecisions

Source: docs/research/2026-09-03-defensibility/rounds/r4-judgment.md

Panout round 4 — investment judgment: "The rework layer for agent labor"

Judge: Garry Tan lens. Methodology: gstack office-hours startup mode (six forcing questions, anti-sycophancy rules, end with an assignment) plus plan-ceo-review cognitive patterns (inversion reflex, proxy skepticism, focus as subtraction, classification instinct). Date: 2026-09-03. Self-imposed dogfood gate: 2026-09-05 — reported by the founder as a failure, two days early, and independently re-verified below. Question scored: would I write the check today for a defensible, $1B-capable company, given this founder's constraints and this evidence? Calibration: prior rounds scored 62/72, 66, 67. I read those for scale only. Every number in the memo was re-checked against the appendix it cites; four were re-derived from the repos themselves.


Score: 64/100

Verdict: you applied every fix the last round priced, you ran the measurement that could kill the hook, it killed it, and you led with the retraction — and the company is worth less today than it was yesterday, because what you retracted was the only empirical claim that made anyone want this.


Dimension table

#DimensionScoreReasoning
1Q1 Demand reality6Appendix E is unchanged and still the strongest thing in the file: 249 👍 on one VS Code issue, 78 on a Claude Code meta-issue whose title says the community is building workarounds, ~14.8k shell files invoking --dangerously-skip-permissions, auto-approve-*.sh hooks shipped to customers by Railway and Render, Anthropic's own 20%→40% auto-approve curve. The memo's framing of it is materially better this round — you now concede the auto-approve ask is owned free by three vendors and name your buyer as "the person who already took auto-approve and now owns the consequence." That is the correct sentence. Held at 6, and it should have gone to 7, because of something I found in your own appendix: the memo's lead behavioral claim, "agent output is merged faster than it can be read," compares a 150-PR unfiltered Copilot sample (median 1.3 min) against a stars>500 filtered control (25.1 min). The matched comparison is in the same appendix — Copilot at stars>500 is 19.6 min, and codex-branch at stars>500 is 28.4 min, slower than the control. Matched on repo popularity the effect disappears and in one sample reverses. Appendix B's own confidence summary tells the reader latency is the defensible claim; on its own numbers, it isn't. And thirteen days in, still zero humans.
2Q2 Status quo6Down one, and appendix G is why. Positions A/B/C are still the cleanest competitive framing anyone has written for this space, and appendix F genuinely hardens the founder-repo cost: the gap survives whitespace/move insensitivity, survives a corrective filter, survives bot exclusion, and the bootstrap lower bound stays above +5 under every definition. But the cost of Position B for anyone who is not you is now measured and it is approximately nothing: pooled +2.5 points across 11 stranger repos with a CI spanning zero, and −1.3 after calendar-month matching. The memo's honest restatement — "not a universal tax, an unknown that differs per team, and that unknown is the product" — is the right retreat, but a status-quo cost of "unknown, possibly zero, possibly negative" is a weaker thing to sell than a 22-point tax.
3Q3 Desperate specificity4Down one. Amazon is still the best named consequence and volunteering the dispute is still right. What changed: the one human whose pain you actually measured is now provably unrepresentative. Your +22 is yours; eleven strangers pool to +2.5, three run negative, and in deno the entire +42.8 is one contributor whose agent-trailer commits are reworked 63.8% against 10.3% for their own hand-written ones — which appendix G correctly says cannot be separated from "this person uses the agent on the churny parts." You have an N of one, that one is at the 99.9th percentile of agent volume, runs everything in bypassPermissions, has zero stashes machine-wide, and 91.6% of his sessions contain no negative event at all. Thirteen days ago you were told to get two numbers and a yes/no from ten humans. Nine hours of work.
4Q4 Narrowest wedge6Down two, and this is the most expensive line in the table. The wedge's structure is still the best thing you have built: value in minute one from data that already exists, zero new behavior, instruments commits you already make, all four invariants respected. Steps 2 through 5 are untouched by appendix G. But step 1 is the acquisition mechanic — you call it "the demo, the hook, and the reason to keep the tool installed" — and G measured what it prints on strangers: positive in 5 of 11, negative in 3, noise in 3, per-class ranking unreadable in 4 and pinned at a 96–100% ceiling in 2. Worse, the metric you chose for the product (the corrective definition) was never run on a single stranger's repo — appendix G is R0 and R1 only. So the sentence "the map uses the corrective definition… because that is the version that does not print backwards on any repo with a meaningful sample" is supported by your repos and nobody else's, and is presented as though it generalizes. Your hook is a coin flip whose coin you have not yet tested.
5Q6 Future fit7Down one. The core argument still holds and is argued rather than asserted: rework has a floor, not a zero; better models shift which task classes are safe rather than removing the need to know which ones. Volume growth makes the question more urgent, not less. The deduction is that you have now moved the load-bearing quantity from a level (retracted) to a variance, and variance products are structurally harder — the customer has to believe their own number is different from the average before they will pay to learn it, and the average is zero.
6Moat mechanism7Up one, earned exactly where the last round said it would be. You conceded that task class, harness and model are git-derivable — appendix A proves it with 15,440 commits — and promoted the override record to Layer 1. That is right: a human committing after a printed contract failure, on a specific diff, is a human judgment about agent work, it exists in no vendor's log and no git history, and unlike in-session rejection it is generated by the act the user already performs. The retention asymmetry (21-day logs against a 30-day label) is still the best physical idea in the file, and appendix G added a second, unintended support for it: trailer and PR-level attribution give opposite signs on identical commits, so post-hoc authorship cannot be recovered at all. Held at 7 by two things. The override rate is still zero-measured — nobody has printed a contract failure yet. And a corpus of a noisy label is not an asset: if the 30-day rework label is near-zero-signal for the median team, you are compounding a private history of a coin flip.
7Incumbent response7Unchanged and still good. The table is honest where it hurts — harness auto-approval is "the closest shipped surface and the memo does not claim it is empty," Entire is UNVERIFIED, GitHub's approving-reviews preview is disclosed and answered with switching cost rather than neutrality. One new item that cuts against you: appendix G is a demonstration that GitHub could run your map across every repo on earth tomorrow and would find pooled +2.5 — which is both why they may not bother, and why "we found what they didn't" is no longer available to you as a defense.
8$1B credibility7Up one. All three fixes landed and I checked the arithmetic: 2,000 teams × 500 outcomes × $0.33 = $3.96M; 50,000 = $99M; denominator 3M Codex weekly actives ÷ 6 = ~500k teams, giving 0.4% / 2% / 10% penetration, stated as such. Radar cited for the prior, not the pricing. The meter now explicitly prices the policy rather than the check, which answers "why does running pytest cost thirty-three cents." Two unfixed: $100M is still 10% of a denominator built from one vendor's weekly actives and an assumed team size of six, and nobody has stated what happens to the meter as autonomy rises — see the sentence check below.
9Founder fit5Held, and the hold is doing work in both directions. On paper this remains the best-fitting version of the company: repo, terminal, commits, no new habit, 5,294 founder-authored commits in 180 days as the substrate — I re-ran that count myself and got 5,580 across all refs in the nine repos, so it ties, which discharges last round's residue. Against it: I verified again today that /code/research/panout has zero commits since 2026-08-23 and no .panout directory exists anywhere. Capture has not shipped. And the revealed-preference datum of this round is the one that matters: you were told to point the map at ten public repos and send each maintainer their table with one question. You cloned eleven, measured 2,042 commits, ran a calendar-matched sensitivity, a PR-level attribution check and a within-author control — in a day, which is genuinely fast — and appendix G's own first paragraph records that no maintainer was contacted. You did the half of the assignment a machine could do and skipped the half that required a stranger. Third time.
10Evidence integrity8Up one, and it is the best dimension in the file by a distance. All five of last round's residues are discharged and I checked each: the commit count now ties, the per-repo range is stated with its actual values instead of a pooled band, appendix C's anchored 22/25 is reconciled against D's looser 53/31 with a stated reason, the 4-of-22 recoverable-path caveat is carried into the memo, and the meaningless 460× ratio is deleted with an explicit note that the two rates have different denominators. Then you commissioned appendix G, which disconfirmed your central claim, and you put the retraction in the first sentence of your moat paragraph. That is rare and I pay for it. Four new residues cost the tenth point, and three are load-bearing: (i) the product's default metric was selected because it does not print backwards. Appendix F caveat 4 says the test-path clause "is not really a corrective signal" and calls R4-strict "the more defensible corrective definition"; under R4-strict your founder gap halves to +11.2 and a new reversal appears in dynorouter-gateway at n=44 agent commits — so the memo's "after the filters, no sign reversal remains that is backed by more than ten agent commits" is false under the strictest variant it cites two sentences earlier. (ii) The corrective metric has never been run on a stranger's repo, as above. (iii) The merge-latency comparison is sample-mismatched, as above. (iv) Minor: "pooled across all repos the gap is 16 points" is the excl-OSS-clones figure (all-repos pooled is 13.2); "the gap survives every one of them" omits that R2 inverts on the founder slice (F discloses it as degenerate; the memo does not mention it); "three repos with one to ten agent commits" is four.

Raw: 63/100.

Gut adjustment: +1. Stated plainly, both halves. Up +3: you ran the experiment that removed your own hook, on your own initiative, within hours of being told to, and you led the memo with the result instead of burying it in an appendix. Round 3 relocated the company when the data said no; round 4 deleted a claim when the data said no, which is a harder and rarer act, and it is the single most predictive founder behavior in this entire file. Down −2: in the same document you chose the definition of rework that does not embarrass you, made it the product default, and wrote a sentence about sign reversals that your own appendix contradicts — and you executed the machine half of an assignment whose entire point, stated in one line last round, was "face validity from someone who is not you is the whole test."

Final: 64/100.

I am aware this is below the previous round. That is the honest arithmetic. The memo got roughly four points better and the company got roughly six points worse, and you are the person who found out. The score is not a grade on the document; it is a probability on the outcome.


The two sentences, verified

Moat, as written

In this founder's own repos, agent commits get silently rewritten within thirty days about twenty points more often than the commits he writes by hand; across eleven public repos the same gap runs from minus twenty to plus forty-three and averages near zero, so which agent work a team can trust is a property of the team, not of the models, and nothing today connects the outcome back to the session that produced it or the judgment that let it through: git cannot even tell who wrote the commit, the session logs holding the cause are deleted in twenty-one days, and the consequence takes thirty to appear. Panout stands at the commit boundary writing the cause and the override into the commit before they expire, so a team's map of which task classes it can safely stop reading is uncopyable by anyone who was not recording when the commit happened.

Would I repeat this to a partner? No — and this time it is not because it is wrong, it is because it argues against itself out loud.

  1. It is 110 words with its own refutation in the middle. A partner hearing "averages near zero" at second nine stops listening; nothing after that clause gets processed. Honesty in the appendix is a virtue. Honesty inside the moat sentence is a category error — the retraction belongs in the demand section, where it is a fact about the market, not in the sentence whose only job is to say why you cannot be copied.
  2. "a property of the team, not of the models" is the load-bearing claim and it is unmeasured. G measured 11 repos once, on one day, with no repeat run (G caveat 9). Its own §6.1 shows the per-repo gap moving 10 to 23 points and flipping sign under a trivial calendar-month control — aspire +5.2 → −2.1, deno +42.8 → +19.7, vscode-copilot-chat +25.2 → +11.2. You have not shown the per-team gap is a property of the team rather than a property of the sampling window. Until you do, "the team's map" may be a map of noise, and everything downstream — the contracts, the autonomy levels, the switching cost, the cross-team prior — is built on it.
  3. You are still making the rework gap do two jobs. It was the problem (why anyone cares) and the moat (why nobody can copy you). G destroyed its ability to do the first. The correct response is to stop asking it to do the second as well. Your moat does not need the gap at all — it needs the override, the contract result, and the retention asymmetry, none of which depend on agent code being worse than human code.
  4. What is genuinely excellent and must survive any rewrite: "git cannot even tell who wrote the commit, the session logs holding the cause are deleted in twenty-one days, and the consequence takes thirty to appear." Three independent, verifiable, physical facts that together make retrospective reconstruction impossible for anyone including GitHub. That is a moat sentence. It is buried at position four in a sentence that opens by conceding your effect size is zero.

Rewrite — what I would actually repeat:

Every team running agents is already deciding, dozens of times a day, which diffs it will not read — and the record of that decision dies in twenty-one days while its consequence takes thirty to appear, so nobody, including GitHub, can reconstruct it afterwards: git cannot even tell whether an agent wrote a commit or was merely open in the next pane. Panout sits at the commit boundary and writes the decision and its outcome into the commit while both still exist, so a team's map of what it can safely stop reading is uncopyable by anyone who was not recording at the moment the commit happened.

No number a partner can look up and discount. No retracted claim. No "only," no "explains." Defensible against every line of every appendix you have, including G.

$1B, as written

Agent output per engineer now grows faster than human reading capacity, so every team is being forced to decide what it will stop reading, and that decision is worth a metered fee per unit of work evaluated rather than per human seat, paid identically whether Panout clears or flags. Two thousand teams at five hundred evaluated outcomes a month is $4M a year and is the month-twelve test; fifty thousand teams is $100M; and unlike seats, the meter reprices itself upward every time a model gets good enough to be trusted with more.

Yes, with one clause added. I checked the arithmetic and it holds. It has a denominator, a ramp, a dated test, and no absolutes. The issuer-pays fix — paid identically for yes and no — is the correct structural answer to the referee problem and it starts revenue on day one instead of day forty-five.

The missing clause is the one a CFO asks in the first meeting and neither this memo nor any previous round has answered: your product's stated goal is that teams stop reading, and your revenue unit is a thing you charge for evaluating. If autonomy level 4 means auto-merge without evaluation, your meter shrinks precisely as you succeed, and every improvement in your product is a cut to your revenue. I believe your architecture already answers this — the contract is what permits the auto-merge, so evaluation continues at every level and what falls is human reading, not machine evaluation — but the memo never says it, and unstated it reads as a business model that eats itself.

Add: "…and autonomy raises the meter rather than lowering it, because a level-4 class is still evaluated on every commit — what stops is the human reading, not the evaluation."


The single hardest push

Round two your company was rejection. You measured it. Twenty-two. Round three your company was the rework gap. You measured it. Noisy, high-floored, reversing in your two largest repos. Round four your company was that agent work gets reworked more. You measured it on eleven strangers. Plus two point five, confidence interval through zero, minus one point three once you match the calendar.

And within the same document you have already relocated again — to "the gap is a property of the team." That is the fourth quantity, and here is what you need to sit with: it is the first one you have not measured.

You could have. The eleven repos are on your disk right now. Split each one's window in half and ask whether a repo's gap in the first half predicts its gap in the second, and whether its per-class ranking in the first half agrees with the second. Report the rank correlation. That is four hours on data you already cloned, and it is the single question on which the entire post-G company rests, because "per-team calibrated map" is only a product if a team's number is a property rather than a draw. Your own appendix already tells you which way to bet: a calendar-month control — the most trivial possible perturbation — moves five of your eleven repos by more than ten points and flips two signs. That is not what a stable property looks like. That is what a noisy estimator looks like.

Now the harder half, and it is not about statistics.

Last round you were given one assignment. It had two clauses joined by the word and: point the map at ten public repos and send each maintainer their own table with one question attached. The reasoning was stated in one line — face validity from someone who is not you is the whole test. You executed the first clause superbly and to a standard most seed companies never reach: eleven repos, 2,042 commits, a calendar-matched sensitivity, a PR-level attribution check, a within-author control, ten caveats. Then appendix G's third line records, in your own words: no maintainer was contacted.

Thirteen days ago: ten operators, two numbers and a yes/no, nine hours. Not done. Today: eleven maintainers, one table, one question. Not done. In between: four memos, seven appendices, three fourteen-day plans, one failed gate, zero commits, zero humans.

You are not avoiding the humans because you are lazy — nobody lazy produces appendix G in a day. You are avoiding them because a measurement can only tell you that you were wrong, and a person can tell you that the whole thing is wrong, and you have arranged thirteen days of extremely productive work so that only the first kind of answer can reach you. Every one of those measurements was a way of asking the market a question without letting it answer.

Here is the inversion. What would make this company fail? Not a bad metric — you have proven four times over that you will find and report a bad metric faster than anyone I have funded. It fails because you spend a six-month runway becoming the world's leading expert on the rework characteristics of your own laptop, and in month five somebody tells you in one sentence that they would never install this, and that sentence was available on day one for the price of an email you did not send.

You are extremely good at finding out you are wrong. You have now exhausted every way of finding it out that does not involve another person. There is no fifth measurement on that machine. There are eleven maintainers whose repos you have already cloned, and ten operators, and an email client.


Every remaining objection keeping this below 90, in priority order

These are the value each closure unlocks; they overlap and are not additive. Closing 1, 2 and 6 gets most of the way.

1. Zero external humans, thirteen days after the assignment. ~9 points. Evidence-only. Ten fleet operators from the watering hole, in writing: how many agent-authored diffs did you merge last week without reading, and did any of them bite you. Two numbers and a yes/no per person, verbatim in the repo, with at least three reporting merged-unread-and-bitten. This was worth nine points last round and it is worth more now, because appendix G removed the last thing that could substitute for it. Q1, Q3 and founder fit are all capped by this single item.

2. The rescue reframe — "the gap is a property of the team" — is the first load-bearing quantity you have not measured, and your own sensitivity analysis argues against it. ~7 points, ~3 idea-fixable. Idea fix, exactly: split-half stability on the eleven repos already on disk. Compute each repo's gap and per-class ranking independently in two non-overlapping time windows; report the across-window rank correlation of the gap and of the class ordering, with a repo-clustered interval. Four hours. If a repo's own gap does not predict its own gap, there is no per-team map, the moat's raw material is noise, and the company needs a different asset — better to know on Friday than in January. The remaining ~4 points are evidence: three external maintainers confirming that the map's two worst classes are the two they would have named unprompted.

3. The minute-one hook fails or prints backwards on roughly half of strangers' repos, and the metric the product actually ships has never been run on one. ~6 points, ~3 idea-fixable. Idea fix, exactly: re-run appendix G's eleven repos under R4 and R4-strict with bot authors excluded — the corrective definition your wedge actually uses — and report how many of the eleven print a positive, readable, class-level gap with n≥25 per class. The repos are cloned; the script is measure_restricted.py; this is one day. If fewer than half print something usable, the hook is not a hook and step 1 of the wedge must be replaced before you show it to anyone. Remaining ~3 points: the stranger's reaction to the corrected map.

4. Metric selection on outcome. ~4 points. Memo/idea-fixable, exactly: Adopt R4-strict as the product default — appendix F caveat 4 calls it "the more defensible corrective definition" and says the test-path clause dominating R3 "is not really a corrective signal." State the founder gap as +11.2 points, bootstrap interval 5.8 to 16.9, not 54% versus 34%. Delete "after the filters, no sign reversal remains that is backed by more than ten agent commits" — under R4-strict dynorouter-gateway reverses at −11.7 with 44 agent commits, which appendix F §5 states in terms. And fix the mismatch between the definition you describe in the wedge (fix-word-or-test) and the numbers you print (which are R4, including the different-author clause). Choosing your metric because it does not print backwards is the one place in four rounds where the document shades toward its conclusion, and it is in the paragraph that defines the product.

5. Founder fit: revealed behavior is now a thirteen-day pattern with a same-day instance. ~5 points. Evidence-only. Zero commits since 2026-08-23, no .panout on disk, capture not shipped, and an assignment executed in its machine half and abandoned in its human half. One shipped thing plus one contacted human changes this entirely and nothing else does.

6. The predictive delta is unmeasurable until 2026-10-05 at the earliest, and the clock has not started. ~5 points. Evidence-only, hard floor. Appendix D's reasoning is correct and its null is honest: 21-day log retention against a 30-day label means retrospective joining is impossible for anyone, and squash-merge severs 100% of SHA linkage in the largest repo. Requirement when it lands: ≥300 commit-time-captured commits across ≥3 repos, git-only AUC versus git-plus-commit-time-features, bootstrap CI on the delta, both outcome classes present. Every day capture does not ship, the earliest possible 90 moves one day right and the runway does not.

7. The public behavioral claim is sample-mismatched. ~3 points. Memo-fixable, exactly: Compare like with like or drop it. Copilot at stars>500 is 19.6 minutes against a 25.1-minute control; codex-branch at stars>500 is 28.4 minutes, slower than the control; the 1.3-minute median comes from the unfiltered long tail of solo repos. Restate as: "in the long tail of small repos agent PRs merge at a 1.3-minute median against 397 added lines; in popular repos the difference disappears" — and then say which of those two populations you are selling to, because they are different companies.

8. The meter's relationship to autonomy is unstated and reads as self-cannibalising. ~2 points. Memo-fixable, exactly: one clause saying that a level-4 class is still evaluated on every commit, so autonomy raises the evaluated count; what stops is human reading, not evaluation.

9. $1B's top rung is 10% of a single-vendor denominator. ~2 points. Partly memo-fixable: run the same ramp against a second, independently derived denominator (Copilot's 26M users, or Claude Code's ~4%-of-public-commits estimate, each with a stated team size) and show the rungs land in the same band. Right now one secondary-sourced number carries the whole market.

10. Residual arithmetic labels. ~1 point. Memo-fixable: "pooled across all repos the gap is 16 points" is the excl-OSS-clones figure (all-repos is 13.2); "the gap survives every one of them" should say "every one except R2, which is degenerate on a single-author slice by construction"; "three repos with one to ten agent commits" is four.


The two questions, answered plainly

Maximum reachable with memo/idea changes alone, no new external evidence: about 70 — and the ceiling has fallen.

Last round's judge put memo-only at 74 to 76. That headroom has been spent: you applied essentially every fix he priced, and they landed roughly where he said they would (evidence integrity +1, moat +1, $1B +1, and the restricted metric was worth its point in Q2 before G took two back). What remains reachable without a stranger is: objection 4 (+1, and it lowers your headline number, which is the point), objection 2's idea half (+2, conditional), objection 3's idea half (+1, conditional), and objections 7 through 10 (+1.5 combined). That is 69 to 70.

The two conditional items are conditional in both directions, and you should hear that clearly: if the split-half stability test and the corrective-metric-on-public-repos re-run come back negative, the honest score goes down, not up — because the post-G company rests entirely on the per-team map being a property rather than a draw. There is no arrangement of words above 70 while the moat sentence's central noun is unmeasured and the market contains zero people.

Minimum external evidence to reach 90, and the earliest date it can exist:

  1. Split-half stability, on repos already on your disk. A repo's gap and class ranking reproducing across two independent windows, rank correlation reported. Earliest: 2026-09-05. This is now a gate on items 2 and 3 — if it fails, there is nothing to show a maintainer and no map to sell.
  2. Ten operators, in writing, two numbers and a yes/no, ≥3 merged-unread-and-bitten. Fourteen days. Earliest: 2026-09-17.
  3. Three external repos, corrective map printed, the maintainer's reaction recorded — specifically whether the two worst classes are the two they would have named unprompted. Same fourteen days, same delivery vehicle. Earliest: 2026-09-17, conditional on item 1.
  4. The predictive delta. ≥300 commit-time-captured commits across ≥3 repos, git-only versus git-plus-commit-time AUC, bootstrap CI, both outcome classes. Thirty days after capture ships. Capture has not shipped as of today. Earliest: 2026-10-05, and only if it ships within forty-eight hours.
  5. One of those operators paying the meter without being asked twice. A charge that cleared. Not an LOI, not a waitlist.

Earliest date this reaches 90: approximately 2026-10-10 — unchanged from last round in calendar terms, but now carrying a new precondition (item 1) that can kill it outright, and resting on a capture clock that still has not started. Your month 4–5 fold gate is roughly 2027-01. You have three of those windows left.


One assignment for this week

Send all eleven maintainers their own table and one question. Eleven messages. The deliverable is eleven sent messages, not a twelfth appendix.

Do four hours of work first, and only because sending a stranger a number that does not reproduce costs you the stranger: split each of the eleven repos' windows in half, check that a repo's gap and class ranking survive the split, and re-run the map under the corrective definition you actually ship. Then send. The message is three lines — here is what your repo's agent-versus-human rework looks like by task class, here are the two classes where it is worst, does that match what you'd have guessed and did any of it bite you. Attach the caveats; they make you more credible, not less. microsoft/aspire, github/spec-kit, n8n-io/n8n, mastra-ai/mastra, Azure/azure-sdk-tools, elizaOS/eliza, microsoft/typescript-go, gradio-app/gradio, microsoft/vscode-copilot-chat, denoland/deno, temporalio/temporal. You already have their data. You do not need permission to email a maintainer about their own repository.

Two things this does not replace, and they are not renegotiable into something more interesting.

Commit-time capture ships within forty-eight hours, in your own repos. It is the only clock in this company that cannot be compressed, and it has not started. Thirty days is thirty days whether it starts Thursday or in October, and the number it produces — whether commit-time features beat git alone — is the number that decides whether you have a rating layer or a dashboard with a very good story. Everything else in your plan is a two-way door. This one is a one-way door on the calendar.

The ten operators, in writing, this week. Fourteen days outstanding. Nine hours of work. The eleven maintainer emails are not a substitute for it — they are the warm version of it, and if you write both this week you will have done, in five days, the two things that four memos and seven appendices have been standing in for.

You have found out you were wrong three times in thirteen days, quickly, honestly, at real cost, on your own data, and you deleted a claim rather than defending it. Almost nobody does that. It is also, at this point, the only thing you have done. The next person to tell you something you do not know cannot be a git command.