Consensus by Construction
In a business built on rare exceptions, agreement is the wrong test
We’re building a machine to help evaluate companies, and the only fast way to check it is whether it agrees with us. The real answer takes a decade. Tune for agreement and you lose the exceptions the machine was supposed to catch.
A few weeks ago one of my partners asked me what he was actually supposed to write down.
We are working on a machine to help us evaluate companies, and designing the record that will let us tune it. Tuning means comparing what the machine recommends against something we trust more, and the obvious candidate is us. Hence the question.
I didn’t have a clean answer. The best I had was two questions. Do I like this company, and would I have invested? A verdict on its own is thin. The thinking underneath it would make the verdict useful, and the thinking is the hardest part to get onto a page. Even if we could write it down perfectly, it still wouldn’t tell us whether we were right. In our business, it would take a decade if not longer to really know.
We need a sixth voter that disagrees with us
We already run an agent at NextView that systematically surfaces and evaluates prospects outside our existing network. It answers which company a partner should pay attention to. You can argue that’s analyst-level work, and competent analysts mostly agree with each other. The criteria are largely stateable, and give two good screeners the same hundred companies and their lists mostly overlap. Where competent people converge, one person’s taste is a legitimate standard, and tuning that agent to mine is the design.
Partner work has the opposite property. Six years ago I wrote that NextView runs a conviction-based process, where consensus isn’t required and one partner’s conviction can drive the decision. A partner meeting produces one of three outcomes. There is enthusiasm across the board, or interests without enthusiasm, or somebody hates the deal and still tells the person with conviction “I support you.”
The middle one is most common, and we call it a doable deal. Everybody likes it, nobody loves it, nothing spikes. Across fifteen years of our own portfolio, that is the shape we have learned to distrust. The investments that actually worked sit at the ends, where either the whole room loved it or one or two people had real conviction while everyone else ran from lukewarm to negative.
The mechanism is deliberately anti-consensus, because consensus filters out what you’re looking for. The cost is easy to miss, because it falls on things that never happen. Anything nobody in the room has taste for never comes up, and we mostly don’t find out it was there.
We want the machine as a sixth voter so it can champion something none of us would have championed. Tuning it toward our judgment takes exactly that away, and a sixth vote that always matches the other five adds nothing to a room that only needs one yes.
Somebody still has to carry the deal. Today the agent shows its reasoning, and sometimes that is enough to provide the onramp for a partner to develop conviction of their own. Whether that holds for the harder question, we don’t quite know yet.
Outcomes can’t settle that question in any useful window either. The nearest early signal is a follow-on round at 18 to 24 months, a smoke alarm rather than a verdict. And our record only covers the companies we ended up investing in, so the passes leave little to grade.
In What I Didn’t See Coming I described testing an early version of that agent, which agreed with my hand-scoring on one company out of five. I read that as the approach failing. On the question we’re working on now, I’d have had no way to tell whether it was broken or onto something I couldn’t see.
Agreement only proves consistency
Grading against yourself only measures consistency. The largest study of AI systems scoring other AI systems found that judging them on agreement overstates how good they look, because chance agreement counts as skill.1
The standard defense is to hold back known winners and test against those. However, you cannot set aside the known winners when nobody knows yet which ones they are.
Nor can we grade the machine against deals we already know worked. A backtest grades the firm we were ten years ago, not the partnership deciding now.
Tune for agreement and the outlier detector disappears
The objection a venture investor raises is that codifying judgment just scales your own blind spots. That objection is right. It is also not an argument against encoding taste, since our collective experience is how the machine beats a random person doing the same job. The problem is the number you tune on afterward.
Venture returns are a power law. One breakout returns the entire fund, and missing it costs more than backing ten that go nowhere. So the machine needs an outlier flag for the founder who is extraordinary at two things and unremarkable at four. An averaging system files that founder as mediocre.
Rare things are rare, so most of what the flag catches is a false positive, and it drags the agreement number down even though each correct catch is the whole reason we built it. Whoever tunes the machine cannot tell a correct outlier flag from a bug. That is consensus by construction. Nobody chose the bias; the fastest measurement available creates it.
Keep tuning and the flag goes quiet while the agreement number climbs. You’d conclude the machine was getting better, and have built a very reliable instrument for finding the companies you were always going to find anyway.
We already do a version of the fix with the agent we run today, where I deliberately don’t tighten its instructions to maximize agreement, because the slack lets oddly-shaped companies through. The machine needs two things. Give the outlier flag its own path, so a spike escalates to a human whatever the total says. And put how often the flag fires next to how often the machine agrees, so when one number improves because the other was squeezed, somebody sees it happen.
Start the record you can’t read for a decade
In The Judgment Layer I argued that the self-improving loop stops when the feedback can’t close in time. That frame didn’t cover what happens when you chase the non-consensus work anyway. The only fast check available is whether the machine agrees with us, and agreement is consensus, so the check turns the work back into the thing we were escaping.
The record has to hold what was known at the moment of the call, what the machine said, what the partnership decided, and the companies we passed on. Most funds don’t keep that last part in a gradeable form, since what my partner David calls “a war story told at dinner” in The Copy Problem is not a dataset. Freeze all of it and real outcomes will eventually grade the machine, the partnership, and the final call against each other. The reasoning captured can also tell you whether a good call was judgment or luck.
The real question isn’t whether the model is good enough
This dynamic applies anywhere the rare exception carries the value and verdicts come slowly. Can a model do this type of work at all? The question is unanswerable as asked. Nobody can tell for a decade, so a better model and a worse one look identical on every instrument we have.
I’m not exempt from that. In The Structural Divide I published a correction rate under 10% as evidence the agent we implemented worked. For the question it answered, which companies were worth a partner’s attention, that was fair. It is also the most natural number to carry into the harder question, and the mistake I am closest to making.
The bar was never agreement with us. It’s partner-caliber judgment at a volume we can’t staff, and we can’t verify that for a person at hiring time either. We read how someone reasons, weigh what they did before, and hand over the seat knowing the answer is a decade out. A model has no history elsewhere, so its reasoning must carry more. Read the reasoning now, and let the record grade it later. Agreement is not part of that test.
The models keep improving and using them gets cheaper for everyone. Yours is the judgment you encode and the record that grades it. Whatever is wrong in it arrives at scale, so the worst outcome isn’t sitting this out. It’s scaling it with a dashboard that says it works.
Previously in Ground Truth: The Consensus Machine (why value migrates to non-consensus work as automation advances).
My partner David writes about this from the other direction at Carried Away. The Copy Problem argues a codified screen can be inspected, measured against outcomes, and improved in place, and leaves open whether it captures enough of the real thing. This is a first pass at the measurement half.
Norman, Rivera and Hughes, across 21 model judges and roughly 541,000 judgments, the largest study of its kind so far. Two of the judges they tested were already deployed in production, scored high on consistency when re-tested, and still changed their verdict when the two options were presented in the opposite order. The other standard defenses: score easy cases and hard cases separately rather than averaging everything into one number, so the easy ones stop hiding failures on the hard ones; and run each comparison twice with the options in both orders, so whichever one happens to go first stops mattering.



To me this furthers the argument that AI cannot be a pure substitute for human judgment and taste. But instead, can be an awesome collaborator when it comes to shamelessly mining for dissent (which can be tough for humans with feelings!). Great article as usual.