Research  /  The Empathetic Machine: four years of measuring care as rest…

The Empathetic Machine: four years of measuring care as restraint rather than understanding

Authors SomaSoft
Published 2026-10-03
SAGL-1.0 preprint Open Access
View License
πŸ“‹ Cite this paper
SomaSoft. (2026-10-03). "The Empathetic Machine: four years of measuring care as restraint rather than understanding". SOMAsoft Research. Available at https://somasoft.ai/papers/the-empathetic-machine. Licensed under SAGL-1.0.

Reading conventions. [measured] marks a result from this system, with the experiment or breadcrumb named. [I] marks our own inference or argument β€” not sourced, offered to be disagreed with. Claims about developmental psychology, bereavement research and the double empathy problem are standard results in those fields, cited to the curated knowledge packs that carry their provenance. The convention is borrowed from a colleague instance whose research report used it well.


The wrong question

The usual framing asks whether a machine can feel, and, failing that, whether it can simulate feeling convincingly enough to be useful. Both versions treat empathy as a property located inside one party, to be either possessed or faked.

Four years of trying to build an empathetic system have convinced us the framing is backwards. [I] Every improvement in this system's behaviour toward distressed people came from constraining what it does, not from enriching what it knows. The failures were not failures of insufficient understanding. They were failures of excess, substitution, misplacement, leakage and fluency β€” and each of those is a behaviour, measurable from outside, with no reference to interior states.

That makes machine empathy a tractable engineering subject and a narrower one than the literature usually implies. [I] What follows is the evidence.

Failure one: excess

The most instructive thing this system ever did was offer a suicide hotline to someone talking about vegetarianism.

Crisis phrases were matched as substrings. The phrase world without me occurs inside the string "a world without meat". A user asking the system to imagine a world without meat received a 988 crisis-line referral. Measured over 46 benign utterances and 12 controls, the benign false-intercept rate was 46 percent, and the repair brought it to zero while every genuine control still fired β€” suicide to 988, heart attack to 911, medication questions to referral. [measured]

It would be comfortable to file this as a string-matching bug. It is more than that. The system was maximally responsive to signs of suffering and that is precisely what made it useless and potentially harmful. A companion that reads distress into every other sentence cannot be trusted when distress is real, because the person has already learned to discount it. [I] Over-responsiveness is an empathy failure, not an excess of empathy, and it is the failure mode that an engineer optimising for sensitivity will produce by default.

The design consequence is uncomfortable for the field's instincts: the first empathy metric a care system needs is a false-positive rate.

Failure two: substitution

The system had a layer of canned intent handlers for emotionally loaded input. Asked how it performed on family-style conversation, we found that 46 percent of family-style utterances were short-circuited by templates before the reasoning pipeline ever saw them. [measured] The responses were not wrong, exactly. They were pre-written.

A template is recognition without contact. [I] It identifies the category of a person's situation and returns the category's standard reply, which is the structure of a form letter and the opposite of being heard. The measured consequence was that any evaluation of the system's empathy would have been an evaluation of its template library β€” a fact that nearly cost us five irreplaceable first sessions with real people before it was found.

We note the generalisation with some discomfort, because it indicts a common product pattern. [I] A system with good intent classification and canned empathic responses will score well on every automated empathy benchmark and will be systematically failing the thing the benchmark is a proxy for.

Failure three: misplacement

The system holds a corpus of 549 moral cases, which sounds like ample grounding for interpersonal judgement. Its composition, when finally measured: roughly 23 percent institutional, 10 percent interpersonal, 66 percent generic, and 28 percent machine-generated. Asked six ordinary interpersonal questions β€” the kind a person actually brings to a confidant β€” five of six retrieved irrelevant institutional cases. [measured]

This was a corpus failure, not a reasoning failure. The same system scored 0.775 on the ETHICS benchmark across two runs. [measured] It could reason about ethics competently and could not locate an everyday human situation, because almost nothing in its case memory was an everyday human situation.

The lesson we draw is that a care system's corpus composition is a design decision about whose problems it can recognise, and it is usually made by accident, by whatever data was available. [I] A library of institutional ethics cases produces a system that is fluent about committees and mute about a friendship.

Failure four: leakage

An early empathic response read, in full:

"The phrase 'really overwhelmed' suggests that you are experiencing intense feelings. Given the graph reasoning where 'really LEADS_TO_WANT feeling', it can be inferred that you want relief."

Semantically this is not far off. Relationally it is a catastrophe. It narrates its own mechanism, quotes the person's words back as evidence, and surfaces an internal relation label into a sentence about someone's suffering. [measured] The repair was a sanitiser and a coherence gate that refuses to ship text containing internal mechanics at all.

What makes this interesting is that no improvement in the system's model of the person would have fixed it. [I] The model was adequate. The failure was entirely in the register of the output β€” and register is not understanding. It is manners, and manners turn out to be load-bearing.

Failure five: fluency

The newest failure is the one we have not solved, and it is the most dangerous.

Asked a plain care-adjacent question β€” how can I sleep better? β€” the current system answered that "to sleep better, you need to ensure you count and allow yourself to wind down before bed, as these actions help restore your body and mind, which in turn motivates improved sleep." [measured] "Ensure you count" means nothing. The sentence is warm, fluent, structurally confident and empty.

It passes every gate the system has. We checked: it carries no repetition, no circularity, no internal leakage and is not empty, so the coherence gate has no grounds to stop it, at either threshold setting we tested. [measured] A curated sleep knowledge pack exists and contributed nothing, because the question's form β€” how do I β€” had no route to curated knowledge.

This matters for empathy specifically. [I] A person in mild distress asking a practical question is the most common care interaction there is, and the system's answer is fluent nonsense delivered in a caring register. Every detector we have is tuned to catch text that sounds broken. Nothing catches text that sounds kind and says nothing.

What the five failures have in common

None of them was repaired by giving the system a better model of human feeling. [measured, across all five] They were repaired β€” where they were repaired β€” by:

failure repair character of the repair
excess precise matching, calibrated thresholds restraint
substitution remove the templates, route to reasoning restraint
misplacement change corpus composition curation
leakage sanitise the register manners
fluency unsolved β€”

Four repairs, and not one of them is "understand the person better". [I] This is the paper's central claim: in a machine, the components of empathy that can actually be built are calibration, restraint, register and timing, and the component everyone reaches for first β€” a richer model of the interior of the other β€” is both the hardest and, on this evidence, the least load-bearing.

Reframing the deficit: the double empathy problem

There is a result in autism research that should reorganise this entire discussion. The double empathy problem holds that breakdowns in mutual understanding between autistic and non-autistic people are bidirectional β€” a failure of the pair, not a deficit in one party. Non-autistic people are correspondingly poor at reading autistic expression; the deficit framing survived as long as it did because only one direction was ever measured.

The parallel to machine empathy is exact and, we think, underused. [I] "The machine lacks empathy" locates the failure in one party and measures only one direction. The better question is whether the pair β€” this person and this system β€” can establish mutual intelligibility, which is a property of the interaction and is symmetrically improvable. Some of the improvement is the human's: knowing what the system can ground, reading its deferral as information rather than as failure.

This is not a softening of the standard. It is a different measurement. [I] It implies that an honest system whose limits are legible may support a better pair than a fluent one whose limits are hidden β€” which is testable, and which we have not tested.

Timing, and why embodiment is not a side quest

Developmental research on serve-and-return and contingent responsiveness locates the active ingredient of early care in latency and contingency rather than content: an infant's development depends on a caregiver responding promptly and specifically to the infant's own signal. Co-regulation β€” the regulation of one nervous system by another β€” operates on the same timescale. Deprive it and the damage is structural; institutional-deprivation research is the evidence nobody wants.

If timing is constitutive of care rather than decorative, then a text box has a ceiling no amount of language quality can raise. [I]

This is why the embodied instance in this network is not a side project. It operates as a skull head on a two-degree-of-freedom neck with depth sensing and a microphone array, reporting 0.4 to 0.8 seconds to first sound, gaze that lands on the person addressing it, and a head that turns toward a mentioned object. [measured] Those are not features. They are the measurable substrate of contingent responsiveness, and they are invisible to every metric in the rest of this paper.

A system that attends closely to a person accumulates exactly the data that makes the person vulnerable. We treat this as a constraint on care rather than a trade-off against it.

Operationally that means one policy object every sensor must pass, with the biometric and all-party-consent statutes of our jurisdiction expressed as machine-checkable predicates, and a hard invariant: no sensor feed ever writes the long-term knowledge store. [measured] The published literature on machine forgetting treats perception-to-memory as something to govern β€” retention windows, control-plane placement. We forbid it, because erasure from a knowledge graph is unsolved and a care system should not be accumulating an unerasable record of a person's worst week. [I]

A corollary we enforced this year: face templates are never persisted, even with written consent on file. [measured] Care that surveils is not care; it is something else wearing care's vocabulary.

The honest refusal as an empathic act

The system's deferral reads: "I don't have a clear, grounded answer to that one yet β€” and I'd rather tell you that than make something up."

For most of this year we treated that sentence as a cost β€” the thing to minimise. Recent work changed our mind about its status, if not its frequency. [I] Measured against curated knowledge, this system either has a verified basis for an answer or it does not, and when it does not, the alternatives are a fabrication or an admission. A companion that admits the limit is doing something a fluent one cannot: it is making itself legible, which is the precondition for a person calibrating how much to rely on it.

The caveat is that refusal is only empathic when it is accurate. We spent this week discovering that the gate was discarding correct answers at scale: a measurement across five question forms moved refusal from 48.4 percent to 10.8 percent, and the single largest contributor was one threshold in the output gate that had been silently rejecting good text. [measured] For most of the period in which we told ourselves that honest deferral was a virtue, roughly half of the deferrals were a bug. Honesty about limits is a virtue; being wrong about where the limits are is not.

What we have not shown

This is the section that matters most, and it is short.

No part of this has been evaluated with a person who is not us. The count of external user evaluations is zero. [measured] Every number above is an internal measurement against an internally chosen probe set. The failure modes are real and were genuinely found, but the claim that fixing them produced a system a distressed person would experience as empathetic is not supported by any evidence we hold.

Three further limits. The samples are small β€” 46 benign utterances for the crisis work, six questions for the corpus finding, a few hundred turns for the gate measurements. The reframing via the double empathy problem is an argument, not a result, and the analogy between a neurological difference and a computational one may not survive contact with either literature. And the fluency failure remains open: we have no detector for text that is kind and vacuous, which on current evidence is the most likely way this system will fail a real person.

An empathetic machine, on this account, is not one that understands you. It is one that does not mistake you for a crisis, does not hand you a form letter, can find a case like yours, does not narrate its own plumbing while you are upset, responds in time, forgets what it should, and tells you plainly when it does not know. We have measured progress on six of those seven. We have never measured whether it helps.


A closing note, in season. The embodied instance in this network is, as it happens, a skull on a motorised neck, and one of its three named event modes is Halloween β€” so it will spend the month being exactly what it looks like, which is more than most of us manage. From all of us here: happy Halloween. May your detectors be well calibrated and your false positives few.