Decoupled Judgment
An architecture review. Six engineers around a table. The proposal: migrate the primary database from on-premises to a managed service. The author presents a careful analysis. Peak query patterns over the last eighteen months. Failure modes observed in production. Capacity projections based on extrapolated growth. A phased rollout plan with rollback gates.
The analysis is competent. Everyone in the room has seen worse plans approved. The query patterns are documented in detail, the capacity numbers carry confidence intervals, the rollback gates are sensible. Discussion focuses on a few edges: the consistency model of the target service, the cost projection for the first eighteen months, the cutover sequence for the secondary indexes.
Every engineer at the table formed their judgment on the system the proposal leaves behind. None has operated the target service. What a capacity projection has to cover, they know from a system where load and capacity related in ways the managed service does not preserve. The proposal is approved.
A year after cutover, the system behaves nothing like the projection. Load patterns have shifted. Queries that the analysis treated as representative are now rare; new patterns dominate that the original system would not have permitted. Applications were rewritten in that year, in ways the plan did not anticipate, because the managed service allowed shapes the old system had constrained. The plan was calibrated against a system that ceased to exist in the act of being migrated. What the room could not price was this: how the new environment would be used once the old constraints were gone.
The analysis was sound.
The shape
What happened in that room is not specific to architecture reviews. The form appears wherever outputs are judged by formal properties meant to signal substantive correctness. As long as the judges could verify the substance themselves, the formal assessment is a shortcut. When they cannot, the form becomes the assessment.
Assessability rests on conditions outside the assessment itself: whether the assessor’s judgment can stay coupled to the reality the output claims to describe. An output can be correct and unassessable at once; correctness belongs to the output, assessability to the relation. The assessor’s competence must have formed under conditions that make the output’s substance legible. The same training, the same tools, the same observable consequences that calibrated their own work. This coupling is rarely built. It arises from the way fields produce their practitioners, or, where judgment is its own practice, from the way judges are produced as judges.
A certificate of competence claims a domain. What earned it was contact with some part of that domain, and the two rarely coincide; the reach of a qualification exceeds the reach of the experience behind it. The gap is durable and mostly harmless: institutions run on the difference.
What makes the gap bite is a substrate shift: the conditions against which competence was calibrated change, while the competence, and the methods it certifies, remain in place. The shift can occur in the object of the methods, the system, the market, the war. It can also occur in the formation of the practitioners themselves, in the training and tools and consequences that once made producers and judges the same people, and in the channel through which that formation passes from one generation to the next. Sometimes the judged decision produces the shift: the thing assessed becomes the thing that invalidates the assessment. And it can run at different speeds. Faster than a career, and someone lives through the transition and can name it. Slower than a career, and no one does.
Incentive, politics, and capture do other work that this essay does not. When a measure becomes a target, the number lies. That is a different failure, though rarely a separate one; the two arrive together. What tells them apart is a test: strip out the incentive and ask whether anyone could then say what the honest number means. Here the answer is no: the numbers can be honest, and still no one can say what they mean. Outputs continue. Their assessability erodes, and nothing in the outputs shows it.
Whether it can be seen at all turns on two things: whether a correction arrives while the judgment is still open, and whether it says which part was wrong. The cases below differ mainly in which of the two was missing.
McNamara
The pattern is old. Robert McNamara came to Ford in 1946 as one of the Whiz Kids, ten young officers who had run statistical control for the US Army Air Forces. At Ford they developed methods that worked precisely in mass production: unit costs, defect rates, throughput, all measurable, all in tight feedback loops between measurement and consequence. The people reading the numbers were the people who would see the cars come back. That is why the correction reached them, and why it pointed at their own figure. McNamara became the first president of Ford from outside the family. Assessment was calibrated against outcomes the substrate made visible. The coupling held.
In 1961 McNamara became Secretary of Defense under Kennedy. He brought the statistical method to the Pentagon. Program budgeting, systems analysis, quantitative evaluation of military options. The transfer was legitimate. The Department of Defense had procurement decisions that quantitative analysis could improve.
As the Vietnam conflict escalated, McNamara turned the same tools on the war itself. Body counts became the measure of military success. The Hamlet Evaluation System quantified pacification. Models projected war duration from attrition rates. The outputs remained formally correct, internally consistent, professionally prepared. What had changed was the substrate. The war in Vietnam was not a closed loop between measurement and consequence, the way Ford production had been. Units reported numbers upward. The counts were often inflated, because the count was also a career metric, but even honest counts would have measured something whose relation to the war no one could verify. Numbers were aggregated, aggregates became the basis for strategic decisions. At no level could anyone close the loop between the number and the reality it claimed to track. The method continued to produce outputs. The relation between output and war had quietly gone.
In Retrospect (1995) has McNamara conceding, three decades after his Pentagon years, that he and his colleagues had failed to grasp the complexity of the conflict. The quantitative method had been misleading on decisive points. What matters about this admission is what it does not say. The Whiz Kids were not incompetent. They had been excellent in their original context. The method was not wrong. It had been right where it was developed. What McNamara describes is a transfer to a substrate where production and assessment no longer coupled, invisible to the people doing it, recognized only in retrospect.
The pattern is not specific to military methodology.
Long-Term Capital
Long-Term Capital Management, founded in 1994, counted Myron Scholes and Robert Merton among its partners. Both had received the Nobel Prize in October 1997 for the theory of option pricing, not the fund’s strategy but the continuous-time framework its models were built in. The fund identified small price discrepancies between closely correlated bonds and amplified the resulting returns through high leverage. In 1995 and 1996 it returned more than forty percent a year. The models were mathematically clean. The calibration drew on extensive historical data. Few quantitative practitioners of their generation were better qualified.
In August 1998 Russia defaulted on its sovereign debt and devalued the ruble. Global markets responded with a flight to quality in a correlation structure the historical data did not contain. Bonds that the models had described as converging diverged. Positions reported as uncorrelated moved in lockstep. The losses arrived immediately; what they meant did not. The models continued to produce recommendations. Their relation to the market had come apart.
The outputs’ assessability rested on a market regime in which the models’ calibration tracked the underlying dynamics. The model-builders had market intuition. What they did not have was a way to detect, in real time, that the regime against which their methods were calibrated had ended. The regime change was a substrate shift the calibration could not represent.
The standard account of the collapse is not about models at all: leverage near twenty-five to one, positions crowded with imitators, funding that vanished when the marks moved. That account explains the size of the loss, and this one does not compete with it. What it does not explain is why, in August, no one inside the fund could say whether the trade had been wrong or the market had changed underneath it. LTCM lost 4.6 billion dollars within four months. In September 1998 a consortium of major banks recapitalized it to prevent systemic disruption. Eleven months passed between the announcement in Stockholm and the rescue in New York.
Replication
In 2015 the Open Science Collaboration published the results of a coordinated replication effort. Independent teams re-ran one hundred psychology studies from leading journals, using the original materials where possible, with protocols reviewed by the original authors. Thirty-six percent replicated by the strictest criterion: statistical significance in the original direction. No criterion put the figure above half. The result was a distribution: across an established empirical discipline, peer-reviewed findings failed replication at a rate the field’s working theory could not accommodate. Ioannidis had argued in 2005, on statistical first principles, that most published findings in biomedicine were likely false.
The demonstration was collective. Ioannidis had reasoned his way to the conclusion; observing it in the field required re-running a hundred studies, at a scale and across clusters no individual vantage point could cover. The original studies had passed peer review at their journals. Competent specialists in the relevant area had conducted each review. The reviewers were not negligent. They were applying the methodological standards they had been trained in. The replication effort revealed that the population of reviewers no longer produced assessments that tracked the substance of the work. Whether the reviewers of 1975 would have caught what the reviewers of 2010 missed cannot be measured. No one re-ran a hundred studies then. The missing baseline is itself an instance of the shift: a change slower than a career leaves no one positioned to have run the measurement before.
The study measures a rate, not a cause. What follows is a conjecture about the mechanism, and nothing in the replication data establishes it: that the shift took three decades and ran through the formation of the practitioners. Empirical psychology in the 1970s had been a smaller field, and doctoral training in it typically meant long apprenticeship with one or two advisors, who had themselves been trained by people who treated statistical inference as craft. Production and judgment were two faces of the same formation.
Apprenticeship is usually credited with contact hours. What it transmits is the abandoned attempt. Sitting beside someone who is uncertain, who runs an analysis and does not trust it, who drops a line of work that will never appear anywhere, teaches what the published record cannot: where the method stops. A publication is a result. The uncertainty has been filtered out of it, not by publication pressure but by the form itself, which has no place to put a discarded approach. A practitioner formed on the published record sees finished inferences and never sees one break off.
By the early 2010s the conditions had changed without anyone deciding to change them. Some of this is on the record. The field had grown, publication pressure with it, and the field’s own methodologists named and measured the analytic flexibility that small studies permit: choosing analyses after seeing data, dropping conditions, redefining outcomes. That cohorts grew and advisors took more students is documented too. That the channel carrying the abandoned attempt thinned as a result is an inference, and no dataset holds it. Statistical software absorbed the derivations a student once had to perform. Whether training stopped producing the intuition that would have made those practices feel wrong cannot be read off a replication rate, and the field’s own account is the better-evidenced one: the incentives changed, and the methods followed. Apply the test from above and the two come apart. Strip out the publication incentive and the flexible analyses go with it. A reviewer formed on finished inferences still has no way to see where a method stopped. The forms of peer review remained intact: blind submission, expert reviewers, qualifications, procedures. What had changed was the population from which reviewers were drawn.
This shift could not be observed from inside a career, because it was slower than one. A reviewer in 2010 had been trained in the mid-1990s, when the erosion was already partway through. Their teachers had been trained in the 1970s under different conditions, but they were a generation removed from the reviewers’ working environment. No single biographical span covered the full transition. The reviewer saw their own work, the work of their immediate colleagues, the studies their journal sent them. Within their cluster, the quality could be recognizably high, because clusters self-select: people work with collaborators they respect, journals send to reviewers they trust, departments hire from networks they know. Local clusters compensated for what the broader population had lost. The cluster looked intact. The discipline did not.
This is what the speed of the shift decides. McNamara could see, in retrospect, what had happened, because the transition occurred within his life. Scholes and Merton could see what had happened, because the regime change had a date. A shift in the formation of practitioners, running slower than any career, produces no comparable witness. The Open Science Collaboration succeeded as a diagnostic only because it sampled across clusters at a scale no individual researcher could match. The conditions of seeing the phenomenon were not the conditions of being inside it.
The precondition
The cases differ in whether anyone could be positioned to see the loss.
Underneath, each is the same gap under a shift, visible in what was certified against what was touched. McNamara’s method was certified as quantitative analysis and coupled to mass production. The Nobel was awarded for option pricing and coupled to a regime that held for four years. The reviewers were qualified as specialists in a discipline and coupled to their cluster. Institutions rarely check whether the reach of the certificate matches the reach of the contact, and the reason is institutional: the unchecked gap is part of what makes certification worth having.
Assessment, as an institutional property, has historically rested on a coupling of production and judgment in the working lives of the people involved. In professions structured around long apprenticeship, shared tools, and observable outputs, those who could produce competently could, by and large, also assess competently. The coupling was not engineered. It arose because both activities drew on the same formation. Institutions resting on the coupling did not have to produce it.
Because the coupling arose without effort, the institutions that relied on it specified the form of assessment without specifying the precondition that made assessment possible. Architecture review, peer review, regulatory oversight: each sets qualifications, procedures, and standards of evidence. None specified the coupling between production and judgment in the people doing the work. As long as the field produced practitioners in whom production and judgment formed together, specification was unnecessary. The precondition stayed invisible, because it was never absent.
The cases above do not prove a general pattern. Four cannot, and neither could thirty. The Therac-25 radiation therapy machine, whose software inherited safety assumptions from a hardware-interlocked predecessor, shows the fast form in a different substrate. What such cases yield is a question: did the coupling between production and judgment still hold when the assessment was made?
The question has to be tested case by case rather than answered in advance. Institutions sometimes try to manufacture the answer: replication, audits, red teams, formal verification. They differ in what they leave behind, and the strongest of them shows what the others share. A proof that an implementation satisfies a specification relocates the judgment. seL4’s functional correctness is proven; whether the specification is the one wanted, and whether the assumed hardware model describes the machine, are not. An audit relocates the same way, checking against criteria somebody had to judge adequate. The proof differs in one respect: it makes the surface it leaves behind explicit. And the substrates above offer nothing to verify against: a war, a market, the methods of a discipline.
The test is harder than it looks in either direction. The conditions whose erosion compromises assessment tend to compromise the ability to diagnose it as well. And the slower the shift, the fewer the witnesses: the cases that most need the question are the cases least equipped to ask it.
The question has a positive form, and it is the one worth asking first. A coupling holds where the assessor could be wrong and find out in time to act, and where the correction says which part of the judgment was wrong. At Ford both conditions held, and they held for the same reason: the people who read the numbers were the people the consequences came back to. Loss alone supplies neither.
The architecture review learned it had been wrong after a year, when the migration was irreversible and the deviation could not say whether the projection method had failed or the usage had moved. Long-Term Capital learned within weeks, at a cost of billions, and the losses did not say at the time whether the trade was mistaken or the regime had changed; that became clear only afterwards, when it no longer mattered. In Vietnam neither condition was ever in play, because the substrate closed no loop at all: the numbers ran upward and nothing ran back, and the war offered no moment at which a judgment was still open. The reviewers were denied the same thing by an institution rather than a substrate. No channel runs from a published paper back to the person who approved it, and none was ever built, because the coupling left nothing to carry. That is what the failures had ceased to have: feedback that arrives while the judgment is still open and names what it is about.
What was implicitly carried must become explicit, or it stops being carried.