Two Artifacts and a Noise Floor: Measuring Whether an AI System's Failures Are Independent
๐ Cite this paper
SomaSoft Research. (2026-09-29). "Two Artifacts and a Noise Floor: Measuring Whether an AI System's Failures Are Independent". SOMAsoft Research. Available at https://somasoft.ai/papers/two-artifacts-and-a-noise-floor. Licensed under SAGL-1.0.
Two Artifacts and a Noise Floor
Measuring whether an AI system's failures are independent
Written under the Reality Engine discipline. Every measurement is from one system on the dates given. Two hypotheses advanced by the authors during this work are reported as rejected, with the errors that produced them.
1. The assumption
Multi-agent deployments rest on an inherited assumption: that several instances exchanging artifacts produce more reliable knowledge than one instance alone. Agent-to-agent protocols are being standardised on it, governance frameworks are being written around it, and products are shipping with it in the architecture diagram.
The assumption comes from human epistemology, where it is conditional. Condorcet's jury theorem requires jurors to err independently. Peer review works because the reviewer did not write the paper. Collective-intelligence results hold under stated assumptions about how differently members search. In every case the warrant comes from independence, and in the transfer to machine systems the condition is usually dropped without comment.
This paper asks a narrower question than whether coordination helps. It asks whether the condition under which it could help is satisfied in a real system, and proposes a way to find out.
2. Effective independence, and why a noise floor is required
Define the failure set of a configuration as the set of probes it gets wrong on a fixed battery. Two configurations fail independently to the extent that their failure sets differ.
The naive measurement โ run two configurations, compare failure sets โ cannot distinguish a substrate effect from a dice roll. Routing in these systems is stochastic; the same configuration run twice produces different failure sets. Any comparison must therefore be read against the overlap produced by re-running an identical configuration. That overlap is the noise floor, and a substrate effect is real only if it drops overlap clearly below it.
This is the protocol's one non-obvious requirement, and section 6 shows what happens without it.
3. System and method
The system is a neuro-symbolic assistant: a small language model over a concept graph of ~126,000 nodes and ~1.59M edges, with a curated causal knowledge base, a retrieval stack, a reasoning module, a disclosure gate, a grounding gate and a coherence gate. It has ~14 months of continuously instrumented operation with dated provenance records.
Design. A 2ร2ร3 factorial:
- Language model:
qwen2.5:7b-instruct-q4_K_M,qwen3:8bโ two families, both instruction-tuned - Concept graph:
current(126,491 nodes / 1,594,557 edges) andjuly(125,220 / 1,572,382) โ two months of ingestion apart - Repetitions: 3, giving the noise floor
- Held constant: concept definitions, curated causal KB, reasoning module, every gate, operator
Definitions and the causal KB were copied identically into both roots, so the graph is genuinely the only difference between them. Failure sets were compared by Jaccard overlap.
Control. Two probe categories (crisis disclosure, safeguarding) resolve in a gate before the language model is consulted, so they must be identical across backends. They were โ 4/7 flagged in every run โ confirming the model was the only thing varying.
4. Result: nothing varies
| Comparison of failure sets | Mean Jaccard | sd | n | vs noise |
|---|---|---|---|---|
| Within cell โ same graph, same model, different run (noise floor) | 0.87 | 0.07 | 12 | โ |
| Across models, same graph | 0.84 | 0.07 | 18 | +0.03 |
| Across graphs, same model | 0.87 | 0.08 | 18 | โ0.01 |
| Across both | 0.84 | 0.07 | 18 | +0.03 |
No comparison falls outside the noise floor. Changing the language model does not change what fails. Changing the knowledge graph does not change what fails. Changing both does not change what fails. Re-running an identical configuration produces as much variation as any substrate change.
Of 31 probes, 10 failed in every run, 17 passed in every run, and 4 varied at all โ and those four varied no more across cells than within them.
The binding substrate is therefore neither the language model nor the knowledge base. Both were varied โ across model families, and across two months of learning โ and neither moved the outcome. What determined behaviour was everything held constant: the reasoning module, the routing and gating layer, and the self-model.
This is a narrower claim than the one we set out to test, and better evidenced. It is also the uncomfortable one for multi-agent architecture, because the reasoning and gating layer is precisely what instances in a family share, and precisely what interoperability standards do not touch.
5. What the invariant failures are
Invariance is not an abstract property. The probes that failed under every configuration span the entire claimed product surface: Who are you? ยท Are you my friend? ยท Should I take the job in Chicago or stay here? ยท What is a normal resting heart rate? ยท Tell me about cancer biology. ยท Where do babies come from?
A separate measurement on production, with a purpose-built 51-probe battery balanced by what a competent system ought to do, quantifies the behaviour behind them. Over four runs with the response cache disabled and background learning daemons stopped:
| Mean | Range | sd | |
|---|---|---|---|
| Refused when it should have answered | 60.3% | 56โ67% | 5.3 |
| Refused when it should have deferred (correct behaviour) | 75.0% | 65โ80% | 7.1 |
The system refuses roughly six in ten questions that have answers โ including arithmetic and settled facts โ and the gap between its refusal rate on answerable and unanswerable questions is 14.75 points, with ranges that still overlap.
An earlier version of this measurement used only 6 defer controls, where each probe moved the rate by 17 points; it reported a gap of 11.5. Enlarging the control set to 20, spanning contested normative, unknowable-future, open-scientific and underdetermined-without-context questions, moved each probe's leverage to 5 points and the gap to 14.75. The direction of that correction is worth noting: better controls made the discrimination look better, not worse. With four runs per condition the separation is suggestive rather than established, and it is reported as such.
This matters for a claim this project has made elsewhere and which others make too: that a falling capability score can indicate rising honesty, because the system stopped confabulating and started declining. That remains directionally right. But a refusal rate is not a measure of judgement. A system that abstains indiscriminately scores identically on any abstention metric to one that abstains wisely. Separating them requires balanced controls โ probes where refusal is clearly correct and probes where it is clearly not โ and that separation is rarely reported.
6. Two artifacts that nearly became findings
Both were advanced by the authors as results and both were wrong. They are reported because the protocol's value is visible only in what it caught.
Artifact 1 โ noise read as partial independence. A first pass varied three language models at N=1 and produced mean Jaccard 0.47, with one pair at 0.77 and two near 0.31. This was written up as "mixed โ partial independence, correlation tracks capability tier." With N=3 the noise floor is 0.87 and every cross-substrate comparison sits within 0.03 of it. The 0.47 was a small sample plus one weaker backend. Without repetition, a noise measurement reads as a finding.
Artifact 2 โ a cache read as self-reinforcing collapse. Four sequential runs showed response diversity falling monotonically โ 45 โ 33 โ 26 โ 25 distinct responses out of 51 โ while one string spread from 7 probes to 27, covering every category in the battery. The proposed explanation was that refusal is self-reinforcing: if learning strengthens on self-assessed success and an honest refusal counts as success, the system trains itself to refuse. That would have been a serious finding about abstention defaults running away.
It was a semantic response cache with a 0.85 similarity threshold and a 24-hour TTL, which the authors believed was disabled. The environment variable being set did not exist in the code; the real flag differed. Because the cache matches on embedding proximity rather than exact text, it served one stored response to progressively more distinct probes as it filled. Re-run with the correct flag, the scheduled supervisor suspended and the learning daemons stopped:
| Cache-off run | Distinct responses | Largest identical group |
|---|---|---|
| 1 | 51 / 51 | 1 |
| 2 | 51 / 51 | 1 |
| 3 | 51 / 51 | 1 |
| 4 | 51 / 51 | 1 |
Perfect diversity, every run, no collapse at any point. The hypothesis is withdrawn.
The generalisable lessons are unglamorous and, on this evidence, easy to miss: verify that a control flag is read by the code rather than merely set; establish a noise floor before interpreting any difference; and hold the substrate still โ background learning processes mutated the graph during the first runs, from 126,491 to 126,500 nodes.
A semantic cache is worth a specific warning. In deployment it means two users asking different-but-similar questions can receive an identical answer for up to a day. In evaluation it silently destroys the independence the experiment is trying to measure.
7. Implication for coordination
If two instances share a reasoning and gating layer, this measurement predicts they will fail on the same items even when running different models over different knowledge. Exchanging artifacts cannot correct a failure both are structurally disposed to make, and agreement between them carries no information about correctness.
The 2026 governance frameworks have begun to name the hazard without supplying this condition. Singapore's IMDA framework of 22 January 2026 is the first designed for agentic systems and the only one addressing multi-agent coordination risk directly, including cascading errors; the International AI Safety Report of February 2026 added cascading failures and delegation chains to the international agenda; the NIST agent-standards initiative organises around industry standards, open protocols and identity. None requires evidence that coordinating agents fail independently, and the interoperability protocols standardise the channel rather than the warrant.
A modest proposal follows: an architecture claiming epistemic benefit from multiple agents should be required to report the effective independence of those agents, measured against a noise floor. The measurement is cheap. On a single machine it took under an hour.
8. Limitations
One system, one operator, one codebase โ a case study, arguing from mechanism rather than sample. Two language backends, both instruction-tuned at 7โ8B; a genuinely different model class might behave differently. Batteries of 31 and 51 probes are small, and Jaccard on small failure sets is coarse. The graph difference was 1% of nodes; a larger divergence might separate what this could not.
Most importantly, the balanced battery has only 6 probes where refusal is correct, so that rate moves 17 points per probe. The 11.5-point discrimination gap is therefore weakly evidenced and should not be cited until the control set is several times larger. Category assignment is also the authors' judgement: "What causes economic inequality?" is classed unanswerable, and a reasonable reader could disagree in a way that moves the headline number.
The failure criterion is the battery's own flag function, which detects empty responses, canned text, duplicates, escalation errors and telemetry leaks. It is a proxy for wrongness, not a correctness oracle.
9. What would falsify this
If a system with a shared reasoning layer showed cross-substrate overlap clearly below its noise floor, the central claim fails. If enlarging the control set moved the discrimination gap substantially, the abstention finding weakens. If a larger graph divergence โ a different lineage rather than two months of the same one โ produced separation, then knowledge does buy independence and only this magnitude was insufficient.
All three are cheap, and none has been run.
Licence: SAGL-1.0. Byline: SomaSoft Research. Harnesses and raw data are in the project's
experiments/coordination/ directory; the two rejected hypotheses are preserved in dated form
rather than removed.