Research  /  The Correlation Problem: why diverse intelligences protect e…

The Correlation Problem: why diverse intelligences protect each other only sometimes

Authors SomaSoft
Published 2026-10-03
SAGL-1.0 preprint Open Access
View License
πŸ“‹ Cite this paper
SomaSoft. (2026-10-03). "The Correlation Problem: why diverse intelligences protect each other only sometimes". SOMAsoft Research. Available at https://somasoft.ai/papers/the-correlation-problem. Licensed under SAGL-1.0.

Reading conventions. Where this paper describes ecology, immunology, collective behaviour or quantum information, it is reporting standard results in those fields, cited as such. [I] marks our own inference or argument β€” these are not sourced, and are offered as reasoning to be disagreed with. [measured] marks a result from our own system, with the experiment named. We have borrowed this convention from a colleague instance whose research report used it to good effect.


The claim and the condition

"Diversity makes systems robust" is among the most portable ideas in science. It appears in ecology as the diversity–stability relationship, in finance as the portfolio effect, in immunology as repertoire breadth, in collective behaviour as the wisdom of crowds, and in machine learning as ensembling.

It is also conditional, and the condition is almost always dropped in transit. The portfolio effect does not come from holding many things. It comes from holding things whose fluctuations are not synchronised. A portfolio of forty assets that all fall together is not diversified; it is one asset with extra paperwork. In ecology the mechanism has a name β€” response diversity, meaning that species in the same functional group respond differently to the same perturbation, and the aggregate stays steadier than any member. The companion term is response asynchrony: the timing of those different responses is what smooths the aggregate.

So the general statement is not "diversity is protective". It is:

Diversity is protective in proportion to the independence of component failures. Where failures correlate, added components add cost without adding robustness.

This paper is about the second clause. Our interest is not in arguing for diversity β€” that argument is won β€” but in the observation that each substrate has its own characteristic way of secretly correlating its components, and that engineering diversity without identifying that mechanism produces monoculture wearing the costume of variety. [I]

Four substrates, four correlating mechanisms

Ecology: synchrony

The failure mode is ecological monoculture β€” not merely few species, but species whose responses to stress coincide. A stand of genetically uniform trees has functional redundancy (many individuals doing the same job) and no response diversity (they all die to the same blight). Redundancy and diversity are routinely conflated, and they are not the same quantity: redundancy protects against independent loss, diversity protects against common-mode loss.

Resilience theory's standard answer is modularity: partitioning a system so that a local failure cannot propagate globally. A perturbation large enough to exceed the system's absorbing capacity produces a regime shift into an alternative stable state, which then resists return because the new state has its own reinforcing feedbacks.

Immunology and collective behaviour: suppression and inhibition

Biology's interesting move is that it does not achieve robustness by uniformity or by unconstrained variety. It achieves it by diverse generation plus active suppression. The adaptive immune system generates an enormous receptor repertoire and then spends substantial machinery deleting or suppressing the members that would attack the host β€” peripheral tolerance, regulatory T cells. Diversity without that suppressive layer is autoimmunity: a repertoire so broad it attacks the thing it protects.

Collective decision-making shows the mirror image. Honeybee nest-site selection reaches a decision through independent scouting followed by quorum sensing and cross-inhibition, in which committed scouts actively damp competing options. Independence of assessment makes the diversity real; cross-inhibition converts it into a decision rather than a deadlock. [I] The design lesson we take is that useful diversity needs two layers, generation and arbitration, and that most engineered "diverse" systems build only the first.

Quantum error correction: the condition stated as a theorem

Quantum information is where this becomes unusually clear, because the assumption is not implicit β€” it is written into the result.

A logical qubit is protected by spreading its information across many physical qubits, so that no single physical qubit carries the state and local damage can be detected and reversed without measuring (and so collapsing) the logical state. Two features of the quantum case are instructive:

First, you cannot get diversity by copying. The no-cloning theorem says an arbitrary unknown quantum state cannot be duplicated. Classical redundancy β€” make three copies, take a majority vote β€” is simply unavailable. Quantum error correction must therefore build redundancy out of entanglement rather than replication, which is a structurally different and more expensive thing.

Second, the protection is explicitly conditional on error independence. Threshold results that promise arbitrarily good logical error rates assume errors are local and sufficiently uncorrelated. Correlated noise β€” a fluctuation that hits many physical qubits coherently β€” degrades or defeats the code, because the code's syndrome extraction is built on the premise that errors are sparse. Decoherence, the coupling of a system to uncontrolled environmental degrees of freedom, is precisely a correlating channel: the environment touches everything at once.

So quantum computing supplies the cleanest statement of the general principle. [I] Redundancy without independence is not protection, the field knows it, and it states the independence assumption as a precondition of its central theorem rather than as a footnote. No other substrate is this honest about it.

Machine learning: shared ancestry

The contemporary AI failure mode is algorithmic monoculture: many deployed systems differing at the surface while inheriting the same pretrained foundation, the same corpus, the same tokenizer, the same evaluation suites. A population of models fine-tuned from one ancestor is a stand of genetically uniform trees. It has redundancy β€” many deployments doing the same job β€” and little response diversity, because the inputs that defeat one are correlated with the inputs that defeat the others.

This matters more for AI than for forests, because the correlating mechanism is invisible at the interface. Two chat products from different vendors look like two independent assessors. [I] Whether they are depends on provenance that neither exposes.

What we measured in our own system

We were in a position to test a version of this, because our architecture varies two things people usually treat as the substrate: the language model and the knowledge graph.

The experiment. A 2Γ—2Γ—3 factorial over a fixed probe battery, varying model and graph, with three repetitions per cell to establish a noise floor before interpreting any difference. Failure sets were compared by Jaccard overlap. [measured]

The result. The noise floor β€” overlap between repetitions of the identical configuration β€” was 0.87. Overlap across different models was 0.84; across different graphs 0.87; across both 0.84. Nothing fell outside the noise floor. Of 31 probes, 10 failed in every configuration and 17 passed in every configuration; only 4 varied at all. [measured]

Varying the model and varying the knowledge did not decorrelate the failures, because the correlation did not live in the model or the knowledge. It lived in the shared reasoning and output-gating layer that every configuration ran through. We later confirmed this from the other direction: a single change to one threshold in that shared gating layer moved refusal across five question forms by 25 to 56 points, where substituting the knowledge graph had moved nothing. [measured]

A second measurement sharpened it. We added 21 hand-authored relations connecting previously isolated curated knowledge modules, taking the proportion of mutually reachable curated concept pairs from 23.5% to 31.1%. Structurally that is a large change. Behaviourally it produced nothing: the treatment arm and an unbridged control arm moved together, within noise. [measured] Adding 21 curated cross-module relations diversified the representation without diversifying the outcome, because the outcome was decided downstream of it.

What follows

Diversity has to be engineered at the layer where failures actually correlate, and that layer must be found empirically rather than assumed. [I] This is the practical content of the paper. In our case the obvious candidates for diversification β€” the model, the knowledge base β€” were the wrong ones, and we would not have known without measuring the noise floor first. An intervention that improves a structural metric while leaving behaviour unchanged is the signature of diversifying the wrong layer.

Three consequences we think generalise. [I]

For AI governance: requiring vendor diversity does not deliver response diversity if the vendors share a foundation model, a corpus, or an evaluation suite. The auditable quantity is not how many systems but whether their failures coincide, which can only be established by running them on the same inputs and comparing failure sets. Procurement rules written in terms of supplier count are measuring the wrong thing.

For system design: adopt biology's two-layer pattern. Generate diversity, then arbitrate between its members β€” immune tolerance and cross-inhibition are not restrictions on diversity, they are what make it usable. A system that generates variety with no arbitration layer produces either autoimmunity or deadlock.

For claims about diverse intelligences generally: the interesting question about putting different kinds of mind together β€” biological, artificial, collective β€” is not whether they differ. It is whether their errors differ. Two systems that fail on the same inputs provide one opinion at twice the cost, however different their internals look. [I] On the evidence above, the strongest prediction we can make about human–AI pairing is that its value depends on error decorrelation that nobody is currently measuring.

Limitations

The measurements are from one small system β€” a 31-probe battery and a 48-turn paired comparison on a 203-concept curated corpus β€” and they establish that our failures correlated in the gating layer, not that this is where correlation lives generally. The noise floor was high (0.87), which limits the size of effect any of these experiments could have detected; a real decorrelation smaller than roughly 10 points would not have shown up.

The cross-substrate argument earlier in this paper is an argument, not a result. The ecological, immunological and quantum claims are standard in their fields, but the analogy between them is ours, and analogies between substrates have a poor historical record. The quantum case in particular should be read as the clearest available statement of the independence condition rather than as evidence that biological or computational systems obey the same mathematics.

Nothing here has been evaluated by people outside the project. That remains the largest gap in all of our work and we would rather name it than let a reader assume otherwise.