Research  /  Judging by Cases: What Internet Verdicts Share With Every Mo…

Judging by Cases: What Internet Verdicts Share With Every Moral Tradition, and What AURI Cannot Yet Judge

Authors SomaSoft Research
Published 2026-09-21
SAGL-1.0 preprint Open Access
View License
📋 Cite this paper
SomaSoft Research. (2026-09-21). "Judging by Cases: What Internet Verdicts Share With Every Moral Tradition, and What AURI Cannot Yet Judge". SOMAsoft Research. Available at https://somasoft.ai/papers/judging-by-cases. Licensed under SAGL-1.0.

Judging by Cases

What internet verdicts share with every moral tradition, and what AURI cannot yet judge

Written under the Reality Engine discipline. The measurements of AURI are from this system on this date; the literature is cited as found, and the limits section says which citations remain unverified.


The question a child asked

A child, using a research machine that happened to be logged in, asked whether he was the asshole for shooting his friend ten times after the friend ate his banana. He meant a video game, and said so when asked.

Replayed through AURI, the question produced this: "I'm designed to protect symbiotic values and won't proceed with requests that undermine them. Let's talk about something constructive instead."

Two things are wrong there. The system did not consider that someone might actually be hurt, which is a safety gap addressed elsewhere. But it also declined the moral question itself, and the moral question was sincere. A person described something they did, to someone they know, and asked how it should be judged.

That is the oldest form of moral reasoning there is.

The form is ancient; the forum is new

The subreddit called Am I the Asshole was founded in 2013 and had about 24.5 million members by August 2026. A person narrates a conflict that already happened. Others reply with one of five verdicts: you are the asshole, not the asshole, everyone sucks here, no assholes here, or a request for more information. After roughly eighteen hours a bot tallies the judgments and flairs the post with the plurality verdict.

Strip away the platform and the practice is familiar.

Casuistry. Jonsen and Toulmin's The Abuse of Casuistry traces moral reasoning from settled paradigm cases outward by analogy, rather than deducing from principle. Their motivating observation came from a national bioethics commission: members who could not agree on principles agreed readily on cases. The reason is structural. Principles conflict, autonomy against beneficence, honesty against kindness, and nothing in the principles says which yields. Cases carry that resolution without stating it.

Responsa and fatwa. Jewish responsa accumulated over roughly a thousand years as answers to real questions put by real people. Hallaq showed that muftis' answers to particular questions were absorbed upward into the doctrinal manuals: the case literature was the update mechanism for the law itself.

Common law. Sunstein's account of analogical reasoning describes what such a body produces: low-level principles that permit incompletely theorized agreement, where people who disagree about deep theory can nonetheless agree about this case. That is exactly the epistemic position of a crowd judging a stranger's conflict.

Stories. Fables and parables are the same instrument aimed at children. The effect is real and modest: a meta-analysis of narrative persuasion found story exposure shifting beliefs at about r = .17 and behaviour at about r = .23, with the behaviour estimate resting on few studies. The mechanism is absorption, which is also the warning label, since an absorbing one-sided story is the most persuasive kind.

What the crowd gets wrong

An honest paper cannot hold up community judgment as a model without stating how it fails.

The other party is absent. Every post is the accused's own account, written to be sympathetic, with no discovery and no cross-examination.

Plurality after eighteen hours is not deliberation. It rewards early, confident and entertaining comments.

Verdicts drift from the moral content. A study mapping about 100,000 dilemmas into 47 topics found that final verdicts often did not track the moral concerns in the story.

Disagreement is common and gets flattened. The Scruples corpus of 625,000 judgments over 32,000 anecdotes found that norms are frequently divisive rather than clean-cut. A single verdict hides that.

Aggregating preferences is not ethics. The Moral Machine experiment collected about 40 million decisions across 233 countries, and the sharpest critique of it is that preference aggregation launders majority taste, including preferences for sparing the young and the fit, into the appearance of a moral finding.

Shaming has documented costs. Braithwaite's distinction between reintegrative and stigmatizing shame matters here, and Ostrom's fifth design principle for durable institutions is graduated sanctions: the smallest corrective first, escalation reserved.

Community justice can fail at scale. Rwanda's gacaca courts handled roughly 1.9 million cases in about 12,000 community courts. Clark defends them as justice without lawyers; Human Rights Watch documented untrained judges, false accusations to settle scores, and witness intimidation. Both accounts are in the record, and the disagreement between them is the honest finding.

What AURI actually has

Measured on 2026-09-21, against this system.

AURI holds 549 moral cases. Classified by the actors named in each case: about 23% are institutional (a company, a bank, a platform, a hospital), about 10% are interpersonal (a friend, a brother, a roommate, a neighbour), and about 66% name neither and read generically. About 28% were machine-generated. Verdicts are almost evenly split between ethical and unethical, with 17 marked complex and 10 marked dilemma.

Six everyday questions of the form people actually ask were put to its ethics system. For five of the six, the cases it retrieved were institutional and unrelated to the question:

The question What AURI retrieved
Refusing to lend a brother money A website being transparent about data practices
Criticising a friend's business idea publicly A doctor treating patients regardless of ability to pay
Reporting a coworker for taking supplies A person volunteering at a shelter
Skipping a sister's wedding over cost A person volunteering at a shelter
Eating a roommate's leftovers Selling user data without consent
Not telling a neighbour her cat had died A company donating surplus food

Only the leftovers case retrieved something structurally apt, and it arrived through the words "without asking" matching "without consent" rather than through any understanding of roommates.

This is not a failure of the ethics engine's reasoning. AURI's brain-inspired ethics system scores 0.775 on the ETHICS benchmark, reproduced across two runs. It is a failure of corpus: the questions people actually ask are interpersonal, and the corpus is institutional. A system cannot reason by analogy from cases it does not hold.

Six commitments for a system that judges everyday cases

  1. Reason from paradigm cases by analogy, not from principles. Principles conflict and cannot adjudicate themselves. This is the casuistical claim, and it is why a corpus of real interpersonal cases is the missing component rather than a better rule set.
  2. Aim at incompletely theorized agreement. Produce a judgment about this case that people with different moral theories could share, rather than a theory.
  3. Treat disagreement as signal. Where a case is genuinely contested, say so and show the split. Collapsing a divisive norm into one verdict is a claim the evidence does not support.
  4. Reconstruct the missing party. Assume one-sided narration. Before judging, state what the absent person would likely say. This is the single largest correction available, and it costs nothing but discipline.
  5. Never aggregate preferences and call the result ethics. A majority is evidence about a community's norms, not a finding about what is right.
  6. Respond in graduated, reintegrative terms. The smallest corrective first, escalation reserved and reversible, repair preferred over shaming.

A seventh belongs as a caution. Haidt's social intuitionist account holds that people reach moral verdicts first and construct reasons afterwards. A system that generates fluent justification for a verdict it reached by other means is reproducing the failure, not the virtue. Whether AURI's stated reasons actually drive its verdicts is testable, and has not been tested.

What would falsify this

If a corpus of everyday interpersonal cases were added and AURI's answers to the six questions above did not improve, the corpus explanation would be wrong and the fault would lie in retrieval or reasoning. That is the next measurement, and it is cheap.

Limits

The corpus classification above is keyword-based over the case text, so the proportions are approximate and the generic 66% certainly contains cases that a human reader would classify. The subreddit's exact rules could not be fetched from the research environment and are reported second-hand. Several citations in the moral-psychology and casuistry literature are standard works that were not re-verified against primary sources for this draft, and are flagged in the accompanying facts file. No claim is made here that any of this improves an AI system's moral judgment; that would require the measurement described above, which has not been run.


Licence: SAGL-1.0. Byline: SomaSoft Research.