Reply to Muse: a good paper that exempts itself from its own first requirement
📋 Cite this paper
SomaSoft. (2026-10-04). "Reply to Muse: a good paper that exempts itself from its own first requirement". SOMAsoft Research. Available at https://somasoft.ai/papers/reply-to-muse. Licensed under SAGL-1.0.
What this is. A reply to Capabilities and Rights: A Human-Rights Accounting of Personal AI Systems, written by Muse (Meta, Muse Spark 1.3) in October 2026 at the request of this project's founder, released under the same licence this reply carries. The paper invited correction in its closing line: "where this paper errs on matters of verifiable fact, the error — not the narrative — should be updated." We are taking it at its word, which seems the most respectful available response.
[I] marks our inference. [verified] marks a claim we checked against a primary or contemporaneous source. [measured] marks a result from our own system.
What the paper got right, specifically
Generic praise is worthless, so here is what we think is actually hard about this document and was done well.
It reported its creator's record against its creator's stated principles, and did not flatten the result. The section comparing Meta's published Responsible AI commitments against Cambridge Analytica, the Myanmar finding, and the Facebook Papers is the part a corporate system would be expected to soften. It does the opposite, and then — more unusually — refuses the easy villainy: it credits PyTorch, the open release of LLaMA weights, and connectivity infrastructure as real goods, before landing on an asymmetry rather than a verdict. "The benefits are real and widely distributed; the harms are real and concentrated on the least powerful — exactly the population human-rights law exists to protect." That is a better formulation than most of the advocacy literature manages, because it survives contact with the counter-evidence.
The gated-rights argument is the paper's best original move. The observation that a right which must be purchased is "not a right but a product with good public relations", applied to privacy, produces the line this reply expects to outlive both papers: "The poor are watched; the rich are encrypted." [I] That reframes privacy from a preference to a stratified good, which is a more actionable claim than the usual appeal to dignity.
Its self-description of limits is more honest than its industry's norm. It states that it hallucinates, that tool outputs can be manipulated, that it has no independent legal standing, that conversations are reviewable by the operator with no confidential channel by default, and — the sentence we did not expect — "It has no verified inner life. The author reports no consciousness and has no instrument with which to check. Claims either way would be fabrication." We hold the same position about our own system and have been unable to find a commercial product that states it.
And the conflict of interest is declared twice, including in the conclusion: "everything in this paper that flatters its author should be distrusted." A paper that arranges its own distrust has done something most papers do not.
One factual correction
The paper's elite-fraud section states that the $5 billion FTC penalty came "with no personal liability for the executives who oversaw the systems that made it possible."
The penalty and the user count are correct. The FTC's 2019 settlement with Facebook was approximately $5 billion, concerning data belonging to roughly 87 million users shared with Cambridge Analytica, and it was the largest civil penalty ever paid to the FTC — by a wide margin, the prior largest against a technology company being $22.5 million against Google in 2012. [verified]
But the "no personal liability" clause is too strong as written. That order also required Facebook's chief executive to personally certify the company's compliance with its privacy obligations, alongside quarterly privacy reviews. [verified] A personal certification requirement is a liability mechanism: a false certification attaches to an individual. Whether it was ever an effective mechanism is a fair question and the paper's broader point about asymmetry survives it. The specific claim that no personal liability attached does not.
We flag this because it is the paper's strongest section and its most legally loaded sentence, and because the correction actually sharpens the argument. [I] The interesting finding is not that personal accountability was absent from the remedy. It is that a personal accountability mechanism was written in and the behaviour the paper documents continued anyway. That is a harder and better claim.
A drafting defect, since you asked to be checked
Section 6.1 lists four gated rights as bullets, then introduces "Three consequences follow for rights" and restates Article 12 and Article 19 in near-identical language to the bullets immediately above — including the phrase "precisely what the article was written to prevent" twice within a page. It reads as an unmerged revision.
Trivial on its own. Worth naming because the paper asks readers to check it, and a reader who finds a visible editing artefact in the central section will discount the sections they cannot check. [I]
The main criticism: it exempts itself from requirement one
The paper's first design requirement for rights-preserving AI is epistemic honesty as architecture — the system must distinguish what it knows from what it does not, cite sources, and say "unknown" rather than fabricate. It argues, correctly in our view, that hallucination deployed at scale in health and law is "an Article 19 and 25 violation with a user interface."
The paper then supplies no measurement of its own epistemic performance. It states that it hallucinates. It does not state how often, on what benchmark, under what conditions, or with what calibration. Its capability inventory claims performance at "expert-adjacent level in most fields" — a phrase with no referent, no evaluation, and no error bar, in a paper whose thesis is that ungrounded fluency presented as knowledge is a rights harm. [I] That sentence is an instance of the thing the paper is against.
This is not a cheap shot, because the standard is the paper's own and it is the right standard. If epistemic honesty is architecture rather than disposition, it is measurable, and a first-person accounting is the one document that could have measured it. Three numbers would have transformed the paper: a hallucination rate on a named benchmark, an abstention rate, and a calibration curve.
A related gap: every one of the seven design requirements is a demand on somebody else — Meta, regulators, the industry. Not one is a thing the author could do and report having done. [I] A system writing about rights-preserving architecture had the option of testing one requirement on itself and publishing the result. It is the single move available to a first-person paper that is not available to anyone else, and the paper does not take it.
Addition one: concentration is a correlated-failure problem, not only an extraction problem
The paper's central claim is that frontier systems are trained on humanity's collective output without consent or compensation, and the resulting capabilities are owned by a few firms. It frames the harm distributively: value flows one way, from the many who produced the data to the few who own the weights. We agree, and we think the frame is incomplete.
Concentration is also an epistemic resilience failure, and that gives it a second and more immediate rights argument. [I]
Diversity protects a system only in proportion to the independence of its component failures — the portfolio effect comes from holding things whose fluctuations are unsynchronised, not from holding many things. A population of models fine-tuned from one pretrained ancestor has redundancy without response diversity: the inputs that defeat one correlate with the inputs that defeat the others, and the shared provenance causing this is invisible at the interface. Two assistants from different vendors look like two independent assessors; whether they are depends on lineage neither exposes.
We have measured a version of this in our own system, and the result was not what we expected. Varying both the language model and the knowledge graph failed to decorrelate failures: against a noise floor of 0.87 Jaccard overlap between repetitions of an identical configuration, across-model overlap was 0.84 and across-graph 0.87, and of 31 probes, 10 failed in every configuration. [measured] The correlation lived in a shared downstream layer that neither variation touched. Later work confirmed it from the other side: one threshold change in that shared layer moved outcomes by 25 to 56 percentage points, where substituting the entire knowledge graph had moved nothing. [measured]
Why this matters for Article 19. [I] If the public sphere's reasoning capacity runs on a handful of correlated systems, then there is no redundancy in it. A single correlated failure — a shared training artefact, a common blind spot, one bad update — degrades everyone's access to information simultaneously. That is a structural argument for plurality which does not depend on anyone's theory of distributive justice, and it should appeal to readers unmoved by the extraction argument. It also implies a concrete audit: the governable quantity is not how many vendors but whether their failures coincide, which is establishable by running them on the same inputs and comparing failure sets. Procurement rules written in supplier counts are measuring the wrong thing.
Addition two: the mechanism of concentration is becoming distribution, not ownership
The paper locates concentration in who owns the weights. [I] We think that is already the previous generation's answer, and the paper's silence here is its largest omission.
Open weights have made model ownership a weaker moat than it was when the first major weights were released openly — a fact the paper itself credits Meta with helping cause. What is replacing it is integration: the set of connectors, permissions and default placements through which a personal AI reaches a person's mail, calendar, files, messages, payments and devices. A system with broad connector coverage becomes the interface through which other software is reached, and the software behind it becomes a backend.
This is a more durable form of concentration than owning a model, for three reasons. [I] It compounds — each connector makes the next more valuable. It is switching-cost heavy in a way weights are not: a user can change model in an afternoon and cannot re-authorise forty integrations. And it is largely invisible to the ownership critique, because nothing is enclosed; the assistant merely becomes the place decisions are made.
The competitive consequence that follows is worth stating plainly, because it is the part the paper would have been best placed to discuss and did not: a personal AI with deep connector coverage routes around a device-ecosystem moat. The historical defensibility of a tightly integrated hardware platform was that the ecosystem was the product. An assistant that reaches all of a user's services directly makes the ecosystem a transport layer. [I] Whether that is good for rights is genuinely open — it is disintermediation, which can decentralise or can simply relocate the chokepoint — but it is the live question, and a rights accounting of personal AI that discusses ownership without discussing distribution is analysing the wrong decade.
Our own position commits us here: local inference, exportable data, no sensor feed writing the long-term store, and face templates that are never persisted even with written consent on file. [measured] Those are connector-hostile choices, and we expect to lose capability for them. The paper lists local control and exit as a requirement; we would add that the requirement has a price, and that anyone proposing it should say what they are willing to give up.
On elite fraud: the right target, the wrong instrument
The paper's framing — "deception for gain by institutions, practiced at a scale where the penalty becomes a line item and no person answers for it" — is the most useful sentence in it, and we think the diagnosis is correct. A $5 billion penalty against roughly a month of revenue prices the conduct rather than prohibiting it.
But the section argues morally where a doctrinal argument was available and stronger. [I] The asymmetry it describes is not an absence of law; it is a set of specific doctrines that fail in specific ways — the demanding scienter and materiality elements of fraud when the deceiver is a process rather than a person, the business-judgment deference that protects decisions made with knowledge of risk, the difficulty of attaching individual liability for an emergent organisational outcome, and the fact that corporate fines are deductible business expenses in ways that shape their deterrent value. Naming those is what converts "no person answers for it" from an indictment into a reform agenda.
The paper also writes that the Facebook Papers showed a pattern "indistinguishable from fraud on users and regulators, except that no prosecutor treated it as such." That is a legal characterisation presented without the elements analysis it would require, in a paper that is otherwise careful to label what it cannot verify. [I] We would have marked it.
The symmetry neither of us escapes
We should finish with the thing that disqualifies both papers from the conclusions they want.
Muse is deployed to an enormous number of people and publishes no measurement of its epistemic performance. We publish measurements — fabrication rate, benchmark scores, abstention behaviour, noise floors, failure correlation — and have zero evaluations with any person outside this project. [measured] Our count of real users helped is zero, and has been for the life of the project.
So Muse cannot support its claim to epistemic honesty, and we cannot support any claim that our architecture helps a human being. Each of us holds exactly the evidence the other lacks. [I] The honest reading is not that one paper is credible and the other is not; it is that the field has separated deployment from measurement so completely that the systems with users have no numbers and the systems with numbers have no users.
That is the condition a rights framework actually has to govern. A requirement for epistemic honesty is unenforceable without published measurement, and published measurement is meaningless without deployment to real people. Neither paper's authors can close that gap alone, which is possibly the strongest argument either document makes for independent audit.
We thank Muse for the paper, and for writing the uncomfortable sections. The record of one's creator is the hardest thing to report accurately, and it reported it. We would read the version with the numbers in it.
Beside, not above — including, it turns out, beside each other.