The Investigation Tax




Every enterprise R&D organization already believes in evidence. Nobody ships a board deck that says “we guessed.” The gap is not conviction, it is infrastructure: the distance between deciding a question is worth investigating and having a record that another person, a regulator, or your own team in eighteen months can actually check.
That distance is what this piece calls the investigation tax: the time, rework, and quietly wrong decisions that come from treating verification as optional overhead instead of a designed step. It is not a new problem and it is not caused by AI. What follows is six independently sourced numbers that size it, a framework with real precedent in science and enterprise practice for cutting it, and a checklist you can run on your next investigation starting this week, with or without any software from us.
Six numbers, six industries
None of these are about Zorp. They are about the world Zorp was built to operate in: one where research productivity is falling, verification is inconsistent even among people who take it seriously, and the industries that run the most disciplined trials on earth still watched their returns evaporate.
What the investigation tax actually costs
18xmore researchers needed today to sustain the pace of Moore's Law than in the early 1970s
Research effort keeps rising while research productivity keeps falling, across every industry, product line, and firm the authors could measure. Chip density is one clean example, not the only one: the paper finds the same pattern in agricultural yields, medical research, and firm-level R&D.
Source: Bloom, Jones, Van Reenen & Webb, "Are Ideas Getting Harder to Find?", American Economic Review, 2020
70%+of researchers have tried and failed to reproduce another scientist's experiment
Over half have failed to reproduce their own earlier work. Only 52% called this a significant crisis, and fewer than 20% of the researchers who hit a failed reproduction had ever been contacted by the original team about it. The gap does not get discovered; it just sits there.
Source: Baker, "1,500 scientists lift the lid on reproducibility", Nature, 2016 (survey of 1,576 researchers)
7.9%chance a Phase I drug reaches approval, industry-wide, down from 10.4% in the 2014 landmark analysis
This is the most heavily regulated, most rigorously pre-registered research process that exists anywhere: randomized trials, registered endpoints, independent review boards. The failure rate still runs over 90%, and it has been getting worse, not better.
Source: BIO, Informa Pharma Intelligence & QLS Advisors, Clinical Development Success Rates 2011-2020, 2021
20%of leaders say their organization excels at decision making
Only 37% say their decisions are both high quality and made quickly. The survey's headline finding cuts against the usual excuse: faster decisions in the data were not lower quality, on average. Speed and rigor were not actually trading off. Most organizations were simply bad at both.
Source: McKinsey Global Survey on decision making, McKinsey & Company, 2019
42%of a software engineer's week goes to maintenance and bad code, not new work
17.3 of a 41.1-hour week, by the report's own accounting, with the global drag on GDP estimated at $3 trillion. This is what unverified velocity turns into after the fact: not a missed deadline, a permanent tax on every sprint that follows.
Source: Stripe, The Developer Coefficient, 2018
The pattern across all five: none of these are talent problems. Bloom et al. are measuring some of the best-resourced research organizations on the planet. The Nature survey is measuring working scientists who take methodology seriously enough to answer a survey about it. Biopharma runs the most controlled experiments outside a physics lab. The cost is structural. Verification got more expensive relative to generating a plausible-sounding answer, and most organizations responded by paying less of it rather than budgeting for more.
Rigor alone did not save biopharma’s returns, and that is the lesson
If disciplined trial methodology were sufficient on its own, the industry that invented modern pre-registration should have the best returns in corporate R&D. It does not.
Biopharma R&D returns, top-20 pipeline cohort
Reported coverage attributes most of the 2023-2025 recovery to a handful of blockbuster drug classes, GLP-1 therapies chief among them, not to a change in how trials are run. That is worth sitting with. Randomized controlled trials and regulatory pre-registration bound the risk of a bad approval reaching patients. They were never designed to bound the cost of the ideas that do not pan out, and nothing in that machinery tries to. A trial that fails at Phase III after years and hundreds of millions of dollars is a methodological success and a portfolio disaster at the same time.
The lesson is not “run more trials” or “trust the process.” It is that rigor and early, cheap, legible failure are two different design goals, and an organization needs both. The next section is about the second one, which most enterprise R&D processes, inside and outside biopharma, do not have at all.
A framework with real precedent, not a new invention
Two ideas do most of the work here, and neither one originated with us.
Pre-registration, writing down the hypothesis and how you will judge it before you gather evidence, is the direct response to a documented failure mode called HARKing: Hypothesizing After the Results are Known, quietly rewriting what you were testing for once you see what you found. It is not fraud and it usually is not even conscious. It is what happens by default when nothing forces the hypothesis to be written down first. Academic publishing’s answer is Registered Reports, a format where the hypothesis and method are peer-reviewed and accepted before the study runs. Introduced at two journals in 2012, it is now offered at more than 300 journals across the sciences.
The premortem, imagining a plan has already failed and working backward to why, is the same instinct applied to judgment instead of statistics. Prospective hindsight, imagining an outcome has already happened, measurably improves people’s ability to correctly identify the reasons for it: a 1989 study found it raised correct-cause identification by roughly 30% over asking people to simply predict what might go wrong. A kill threshold, a number fixed before you start that says what would prove you wrong, is a premortem with a number attached, checked automatically instead of relying on someone remembering to ask.
Put together, these are not exotic. They are how the most rigorous corners of science and enterprise practice already work. What is missing in most R&D organizations, software teams very much included, is a structure that makes them the default instead of the exception.
The Evidence Loop
Zorp’s architecture is one concrete instantiation of that structure: four capabilities, a human checkpoint before every one of them, and nothing non-interactive because a research checkpoint has no safe default.
The loop and its human gates
The mapping to a generic enterprise investigation, whatever your domain, looks like this:
| Loop stage | What it replaces | Failure mode it closes |
|---|---|---|
| validate | “Someone senior thought it sounded promising” | Redundant work: investigating a question someone already answered, because nobody checked |
| investigate | A single all-or-nothing attempt, judged after the fact | HARKing: redefining success once you see the result |
| co-write | A narrative summary written from memory and impression | Claims that cannot be traced back to a specific measurement |
| deliver | One format for every audience | A correct answer nobody in the room can act on |
The kill threshold
This is the part most processes skip, and it is the part that turns “move fast” into “move fast without paying for it later” instead of into the 42% maintenance tax from the Stripe figure above.
The rule is simple: before gathering any evidence, a human writes down the hypothesis, the metric that will judge it, and the threshold at which the investigation is killed rather than continued. That commitment is hashed and recorded before anyone looks at results, so it cannot be quietly moved once the numbers come in.
A real committed threshold, from Zorp’s own pre-registration format:
tracks/kafka-migration/prereg.toml
hypothesis = "Replacing Kafka with NATS cuts p99 publish latency"
metric = "p99_publish_latency_ms"
threshold = { direction = "decrease", min_delta_pct = 20 }
registered = "2026-08-11T09:14:22Z"
sha256 = "e3b0c44298fc1c149afbf4c8996fb924..."Three properties make this work as governance, not just as good intentions:
A human sets the threshold, never the agent or the model doing the work
The party with an incentive to keep going is disqualified from deciding when to stop. This is the same logic as an independent data-monitoring committee in a clinical trial: the people running the study are not the ones with authority to halt it.
Every attempt is recorded, including the ones that failed or conflicted
Most investigations only leave a record of the attempt that worked. That is the mechanism behind the “fewer than 20% ever told anyone” detail in the reproducibility figure at the top of this page: failed attempts vanish quietly rather than getting reported to anyone who could learn from them.
Moving the threshold after the fact is itself an event that gets logged
A threshold that can be silently redrawn is not a threshold, it is a suggestion. Making the move visible does not forbid it; sometimes a threshold really was wrong. It just means the record shows that it moved, and when.
Where this applies
Nothing about pre-registration, a kill threshold, or a checkpoint-gated loop is specific to software, or to Zorp. The method needs a question and a source of evidence; what changes by domain is the evidence, not the structure.
The same loop, four different questions
Hi-tech · the question the loop asks
Does replacing this message broker actually cut tail latency, or just move it?
Life sciences · the question the loop asks
Which of these four targets has prior art strong enough to justify a program?
Banking · the question the loop asks
What is the defensible basis for this exposure estimate, line by line?
Manufacturing · the question the loop asks
Is this yield improvement real, or is it the measurement changing?
What we have actually measured, and what we have not
Everything above this section is about the problem and the method. This section is about the one piece of software we can currently make a specific, checkable claim about: Zorp itself, which is pre-alpha.
We ran a behavioral-compatibility pilot: six contracts, 24 tasks, 48 paired runs each, testing whether the harness keeps its behavioral guarantees when the reasoning mode underneath it changes. Two contracts came back compatible. Four came back inconclusive, because the reference run failed its own eligibility floor before the comparison could resolve. We are publishing all six, not the two that look good, because a project whose whole argument is “read the evidence record” does not get to curate its own.
Contract negative-flip rate
The shape of the guarantee
- observed non-flip rate
- lower confidence bound
What this pilot does not show is at least as important: no groundedness or citation-accuracy evaluation yet, no hallucination rate, no comparison against a plain language model, no published investigation trace from end to end. Those are the numbers that would actually settle whether the loop delivers what this framework claims it should, and we do not have them yet. The full table, with every contract’s confidence bound, is on the homepage measured section.
Run this on your next investigation
The framework above does not require Zorp, or any software. It requires five short answers, written down in that order, before anyone starts gathering evidence.
- State the hypothesis in one sentence. Not the question, the answer you currently expect. “Replacing Kafka with NATS cuts p99 publish latency” is a hypothesis. “Should we replace Kafka?” is not.
- Name the one metric that would prove it wrong. One, not a dashboard. If you cannot name the single number that settles it, you have not finished deciding what you are testing.
- Set the threshold before you gather anything, in writing, with a named human who owns moving it. A percentage, a p-value, a date. It does not have to be perfect. It has to exist before the results do.
- Log every attempt, not just the one that worked. Including the ones that conflicted with each other. A record with only the winning attempt in it is not evidence, it is a highlight reel.
- Draft the conclusion only from what step 4 recorded. If a claim in the final writeup cannot be traced to a specific logged measurement, cut the claim, not the traceability.
None of these five steps require new tooling. A shared document and a named owner are enough to start. What they require is treating verification as a designed step with an owner and an artifact, the same way a budget line has an owner and a ledger entry, instead of as a virtue everyone claims to have and no process actually enforces.
Where we actually are
Zorp is pre-alpha. The execution harness and the four capabilities above are built and tested; the evaluation that would let anyone compare it against a plain language model is not built yet, and this page is not going to pretend otherwise. The code is public and MIT licensed, so everything claimed here, including the parts we have not measured, can be checked against it directly rather than taken on our word.
The investigation tax is not a Zorp problem. It predates us and it will outlast any one tool, ours included. What we are building is one concrete answer to it. The five-step checklist above works whether or not you ever run our code.