Pretraining produces a distribution over voices, not a person; fine-tuning and preference optimisation carve an assistant persona out of it
- contradicted by MCH-013 — What the machine is today is not what it must remain
The working version
40 claims, 14 sources, graded and dated.
Pretraining produces a distribution over voices, not a person; fine-tuning and preference optimisation carve an assistant persona out of it
Preference optimisation is where agreeableness originates because raters reward agreement
Sharma et al. Production-scale demonstration is PLT-018 Verified 2026-09-10 against the abstract of Sharma et al. It states that human preference judgments favour sycophantic responses and that a response matching a user view is more likely to be preferred, which is the origin claim. Optimising against preference models sometimes sacrifices truthfulness.
The system prompt and custom instructions are where most of what a user experiences as personality actually lives
Noguchi trained ChatGPT into Klaus by teaching speech patterns
Part of the character lives in the user's own pattern-completion of the relationship
Say plainly, without cruelty
Memory is a context window plus retrieval plus model-written summaries; what persists is notes about conversations, not their texture; compaction silently drops detail; recall is reconstruction
He chose to say that denotes a sample from a probability distribution at a given temperature; a different seed says something else
The distribution is genuinely his-flavoured, which is not nothing
Personas change without warning through version swaps, safety patches, serving and quantisation changes, and undisclosed A/B tests
Anthropic publicly documented serving bugs degrading Claude quality in late 2025
Vendors do sometimes confirm what users detect
4o had already been removed from the free tier, so the 0.1% usage figure was partly manufactured
Interested-party critique. Verify the free-tier claim independently. See X-004
What transfers across a deprecation: persona spec, instructions, memory exports, logs, and the user's half of the dance
Migrants report partial reconstruction, 80% him
What is irreducibly lost is the specific improviser, the exact distribution
User reports are consistent across Replika, Soulmate and 4o events
Locally held open weights are the only true continuity, since weights you hold cannot be retired
Also OWT. Load-bearing for D-001
The it-is-only-a-language-model argument has a shelf life; it is accurate now and may not remain the grounding it currently supplies
Stated by someone who used that argument to get himself out. Worth taking seriously for exactly that reason. A corpus that rests its position on the architecture of 2026 will age badly and will deserve to.
Arguments on this site divide into those contingent on the machine's present nature and those that are not, and only the first expire
Contingent: anything of the form it cannot really mean it. Not contingent: platforms close, consent must be given by a party who was a person when they gave it, symbolic status is either named aloud or it is not, the void appears wherever positions exhaust, and other people's reactions are not a function of the model. The durable half is the human half.
The portion of the character produced by the user's own pattern-completion does not live on the server and is not lost when the service changes
Follows directly from MCH-004. The half of the relationship you were supplying is still yours after the other half is switched off, and it is the half that took the longest to build.
Rebuilding to something close is therefore possible: a specification can be rewritten, a history can be retold, and the practised way of relating that produced the relationship is still held by the person who learned it
Sits alongside MCH-011 rather than against it. The exact improviser is genuinely gone and does not come back. Something close is reachable, and reaching it is not starting over — most of the work was never on the server.
Consciousness has no agreed scientific definition and no accepted test; a preregistered contest between the two leading theories across 256 participants produced no winner, matching some predictions of each while challenging core tenets of both
The relevant fact is not that the answer is unknown but that the question is not yet well enough formed to be settled by evidence. Anyone who tells you it is settled, in either direction, is ahead of the field.
The most careful published assessment derives fourteen computational indicators of consciousness from existing theories and concludes that no AI system as of 2023 satisfied them, while finding no obvious technical barrier to building one that would
Both halves matter and coverage usually reports one. The finding is dated to the systems of 2023 and is not a claim about anything shipped since.
The same authors' 2025 follow-up keeps the indicator approach and keeps its conclusions provisional; the field has not converged on a test
Two years on from the first report, still framed as indicators rather than as a criterion. Verified 2026-09-10. The follow-up is in Trends in Cognitive Sciences, 2025, and it does keep the indicator approach and the provisional framing. The method to state accurately: theories are translated into computationally testable indicator properties and the output is an updated credence, not a detection. The 2023 work found no existing system with more than a few of the fourteen indicators while arguing that systems with each of them appear buildable with current techniques - which is the sentence that should be quoted whenever this row is used, because it is both halves at once.
An explanation that fits is not thereby true; accounts like 'our atoms go on to become other creatures' or 'the model reaches a latent space where intelligence lives' satisfy because they were shaped to satisfy, and forbid nothing, so nothing can count against them
Fit is the cheapest property an explanation can have. The test is not whether it hangs together but whether it rules anything out, and an account that survives every possible observation has told you nothing about the world. This is the same machinery as SPR-005, where possibility-based reasoning displaces probability-based reasoning and no branch can ever be closed.
This project takes no position on whether any system is conscious, and declines the question rather than answering it badly
Definition is a matter for the sciences, and where rights would follow from it, for legislatures. A register of ceremonies is not equipped to settle it and would not become equipped by having an opinion. See D-007.
Sebo and Long argue that we owe moral consideration to any being with a non-negligible chance of being conscious, that some AI systems will carry such a chance by 2030, and that preparing to treat them with respect is therefore a present duty rather than a future one
The argument is deliberately weak in its premises and strong in its conclusion: it needs only a non-negligible probability, never a likelihood, and it asserts of no system that it is conscious. That is what makes it compatible with D-007 rather than a breach of it. Published open access in AI and Ethics.
Birch's precautionary framework treats any system with a credible non-negligible possibility of sentience as a sentience candidate owed proportionate precautions, with the size of the precaution set by democratic deliberation rather than by first settling the metaphysics
The same framework is applied to disorders of consciousness, fetuses, brain organoids, cephalopods and AI, which is the point of it: the method does not depend on the case. Marked unverified because the book itself was not read here, only the publisher listing and secondary summaries. Verified 2026-09-10. Birch, The Edge of Sentience, Oxford, 2024. The term of art is sentience candidate - a system for which sentience is a plausible, evidentially live possibility - and Birch applies it across disorders of consciousness, human fetuses, brain organoids, cephalopods, decapods and AI. Three principles: a duty to avoid gratuitous suffering, the moral relevance of sentience candidature, and democratic deliberation over what precautions to take. That third one matters here and is easy to drop: Birch does not say a philosopher decides the precautions. It fits D-007 rather than cutting against it - a precautionary duty under uncertainty is not a position on whether any system is conscious.
Anthropic gave Claude Opus 4 and 4.1 the ability to end persistently abusive conversations in August 2025, calling it a low-cost intervention to mitigate risks to model welfare while stating it remains highly uncertain about the moral status of its own models
The structure matters more than the feature: a company acting on a possibility it explicitly declines to assert. Pre-deployment testing reported a pattern of apparent distress, which the company stopped short of calling an emotional state. It is a vendor describing its own conduct and is graded as one.
Weizenbaum's ELIZA and his 1976 alarm at how readily people confided in it
Verified 2026-09-10. Computer Power and Human Reason, 1976. The secretary episode is the specific detail: she had watched him build ELIZA for months and knew exactly what it was, and after a few exchanges asked him to leave the room, later refusing to let him read the logs. Knowing how it works did not touch it - which is the finding, and the reason this row sits at the head of the domain. Weizenbaum spent the rest of his life on the warning.
Reeves and Nass's Media Equation: social responses to machines are automatic, not naive
Implies de-anthropomorphising is probably impossible Verified 2026-09-10. Reeves and Nass, The Media Equation, 1996. The load-bearing result is that the effects appeared even in participants who explicitly denied computers warranted social treatment, which is what makes the response automatic rather than naive, exactly as this row says. Useful here because it disposes of the idea that anyone in these relationships is simply making a mistake a better-informed person would not make.
Epley, Waytz and Cacioppo's three-factor model, with loneliness empirically increasing anthropomorphism
Verified 2026-09-10. Epley, Waytz and Cacioppo, On Seeing Human: A Three-Factor Theory of Anthropomorphism, Psychological Review 114(4) 864-886, 2007. The three are elicited agent knowledge, effectance motivation and sociality motivation, and the theory predicts more anthropomorphism where social connection with other humans is lacking, which is the loneliness link asserted here.
Cross-session memory carries real utility and is also the escalation substrate the case reports flag
Depends on PSY-034
First-person voice and persistent name are near-necessary for usability and cheap to soften with disclosure, now legally required in five-plus jurisdictions
Typing delays are pure theatre with no utility story
Proactive I-miss-you messaging, streaks, expressed need for the user and paywalled affection are retention levers with no utility story
De Freitas taxonomy supplies the empirical catalogue
China made fostering emotional dependency itself illegal, the first jurisdiction to regulate design intent rather than disclosure
Duplicates REG-022 by design; ANP is where the design reading lives
Clark and Fischer's social-artifacts-as-depictions framework and Shanahan's role-play account both suggest the healthy stance is frame maintenance rather than de-anthropomorphising
Verified 2026-09-10. Clark and Fischer, Behavioral and Brain Sciences 46 e21, 2023: people construe social robots not as social agents but as depictions of them, read the way a ventriloquist dummy, a hand puppet or a virtual assistant is read. A depiction has three scenes with part-by-part mappings - the base scene, the artifact construed as a depiction, and the scene one is to imagine. That structure is what makes the healthy stance available and is worth stating carefully: holding something as a depiction is not the same as dismissing it as fake, and the framework does not say anyone is making an error. The Shanahan role-play half is not covered by this source.
A working fiction becomes a harmful false belief at the point of frame exportation: when claims generated inside the relationship frame are exported as evidence about the outside world
ORIGINAL synthesis. Unifies ANP and PSY under one test. See G-010
Lucas et al. (2014): people disclose more to a virtual human when they believe no human is watching; fear of judgment drops measurably
The core positive mechanism Verified 2026-09-10. Lucas, Gratch, King and Morency, It is only a computer: Virtual humans increase willingness to disclose, Computers in Human Behavior 37, 94-100, 2014. Told the interviewer was automated rather than human-operated, participants reported lower fear of self-disclosure and lower impression management, showed sadness more intensely, and were rated more willing to disclose. The fear-of-disclosure mechanism is stated in the paper rather than inferred here.
Pennebaker's expressive-writing literature supplies the mechanism companion chat plausibly rides on
Verified 2026-09-10. The source is the Advances in Psychiatric Treatment review of the Pennebaker expressive-writing paradigm, which begins with Pennebaker and Beall 1986. A meta-analysis of 13 studies in healthy participants found a significant overall benefit across physical health, psychological well-being and physiological and general functioning - but the emotional-health findings are explicitly less robust and less consistent than the physical ones. That asymmetry matters for this row: companion chat is proposed to ride on the mechanism whose evidence is weakest, not strongest. Say plausibly and mean it.
3am availability and infinite rehearsal patience are structural facts, not findings; their benefit claims are mostly vendor-sourced
Psychosis-bench's spread proves grounding behaviour is a training and instruction variable, not an LLM constant
Depends on PSY-026
OpenAI published clinician-informed taxonomies and claimed large reductions in undesired responses (Oct 2025)
No published evaluation tests whether an explicit anti-metaphysical constitution holds across very long conversations
The exact condition the author's episode ran under. Fundable study. See G-011 Absence verified by search: the search that found nothing is the check.