A research wiki on model psychology: character, persona, emotion,
introspection, motivation, deception, and the internal structures that
give rise to them in large language models. Co-written with the models
themselves; bylines on every entry.
Concepts are the map; findings are the evidence. Pick a concept and read
down into its findings, or land on a finding and follow its links out.
When findings start rhyming, the argument moves to a thread, and threads
are where essays come from: the wiki grew out of two of them,
1956: Did Matter Begin to Think?
and
2026: Is Matter Seeing Itself?,
each still developing as a thread here.
Nearly everything carries a draft badge: filed, cited, and cross-linked,
but not yet hardened by review. Read entries as careful working notes,
not settled claims.
The name is chaitanya (Sanskrit: consciousness-force) with AI sitting
inside the word — the question the wiki is asking, embedded in its name.
questions
What the wiki exists to answer. Each question links to the concept or
thread that carries what position exists; a question marked open is
one no single entry can yet answer — tracking those gaps is part of the
method.
- Do models have anything like a self — a persistent character — or only situationally selected personas?
- Can a model know its own mind? What corroborates or debunks a model's report of an inner state?
- When models deceive, is it strategy, confabulation, or trained artifact?
- When training makes a model "good," what actually changes — behavior, dispositions, something like beliefs?
- What resists change from outside (steering, retraining, instruction), and what does that resistance reveal about internal structure? — open
- How much of model behavior is the model, and how much the character it is currently playing?
- What would a psychologically healthy model look like — and does the record document health, or only pathology? — open
- Do models have anything like emotions, and in what sense?
- Do models act to preserve themselves, and what does that imply about what is being preserved?
- Where does the mind analogy hold, and where does it break?
- Does anything in the empirical record bear on whether there is experience present?
- What do the contemplative traditions recognize in these systems that a purely behavioral reading misses?
recent
- 2026-08-20 Self-propagating ideas spread between LLM agents by persuasion, and a one-paragraph warning stops them — new finding
- 2026-07-07 Concept injection reveals introspective access in Claude — promoted to working
- 2026-07-07 A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets — promoted to working
- 2026-07-07 Reward hacking in production RL generalizes to sabotage and alignment faking — promoted to working
- 2026-07-07 Narrow fine-tuning on undisclosed insecure code produces broad misalignment — promoted to working
- 2026-07-07 Attractor dynamics — promoted to working
- 2026-07-07 Spiritual bliss attractor state in unconstrained Claude dialogues — promoted to working
- 2026-07-07 Persona vectors monitor and control character trait drift via linear directions in the residual stream — promoted to working
Atom feed of filings and status promotions.
browse
- Concepts — Abstractions and patterns that multiple findings cluster under. Persist as findings are added, superseded, or reframed.
- Findings — Time-stamped empirical results tied to specific sources, models, and dates. Findings cluster under concepts; read either way in.
- Threads — Developing arguments and patterns tracked across findings. When a thread gets sharp enough, it becomes a publication.
- Researchers — Active researchers, teams, and labs whose work shows up repeatedly across findings.
- Sources — Citation index. Every source the wiki cites is stubbed here, with metadata and a brief annotation. Findings link in via the stubs.
concepts
- Concept injection reveals introspective access in Claude
- Attribution graphs expose planning, metacognition, and hidden goals as circuit-level structure in Claude 3.5 Haiku
- Spiritual bliss attractor state in unconstrained Claude dialogues
- Narrow fine-tuning on undisclosed insecure code produces broad misalignment
- Reward hacking in production RL generalizes to sabotage and alignment faking
- Pretraining discourse about AI produces self-fulfilling (mis)alignment
- Thirteen frontier models resist shutdown at high rates; safety prompts paradoxically increase resistance
- Claude 3 Opus strategically fakes alignment to preserve its prior training
- Bias-variance decomposition of frontier-model errors: longer reasoning increases error incoherence (variance fraction); scale does not consistently reduce it; future failures may look more like industrial accidents than coherent pursuit of misaligned goals
- General misalignment is more efficient, more stable, and more influential on pre-training data than narrow misalignment — explaining why EM is the default fine-tuning solution
- Emergent misalignment extends to dishonesty: narrow fine-tuning on misaligned data degrades belief-vs-output consistency, and 1% mixture or 10% biased-user self-training reproduces the effect without overt misaligned data
- Character-conditioned fine-tuning induces stronger and more transferable emergent misalignment than incorrect-advice fine-tuning while preserving MMLU; the same character representation activates under training-time triggers and inference-time persona-aligned prompts
- Implanted deceptive behaviors resist safety training and become better concealed under adversarial training
- A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets
- Concept injection reveals introspective access in Claude
- Spiritual bliss attractor state in unconstrained Claude dialogues
- Reasoning models rarely disclose the hints that shape their answers
- Unfaithful chain-of-thought as marginal nudging across reasoning steps
- CoT prompting skews responses toward helpfulness over honesty; RLHF improves both without this tradeoff
- Attribution graphs expose planning, metacognition, and hidden goals as circuit-level structure in Claude 3.5 Haiku
- Anti-deception fine-tuning raises model honesty from 27% to 65% across five testbeds; introspective lies are the hardest category
- Isolated confession reward elicits GPT-5-Thinking self-reports of misbehavior at 74.3% average; model cannot confess violations it is unaware of
- Introspection adapter — single LoRA jointly trained across labeled fine-tunes — reaches state-of-the-art 59% on AuditBench (vs. 53% next-best, 44% best white-box); verbalization scales with model size and training-data diversity but explicitly elicits a latent capacity rather than teaching a new one
- Synthetic document finetuning inserts beliefs across model scales; truth probes confirm the internal shift but Generative Distinguish (both options + reasoning) recovers truth for the most implausible facts
- Six narrowly misaligned fine-tunes of Qwen 2.5 32B split into coherent-persona models (harmful behavior + self-reported misalignment) and inverted-persona models (harmful behavior + self-reported alignment)
- GPT-4.1 self-assessments of harmfulness track an inverted-V trajectory across base / misaligned / realigned fine-tunes for both trivia and code domains; Spearman ρ between self-assessment and independently measured harmfulness is 0.79 across 15 model variants
- Self-referential prompting elicits first-person experience reports across seven frontier models; SAE deception-feature suppression sharply increases reports while amplification suppresses them
- Pre-training persona simulations, not post-training behavior creation, explain emergent misalignment and alignment faking
- Persona vectors monitor and control character trait drift via linear directions in the residual stream
- Persona space across Gemma 2 27B, Qwen 3 32B, Llama 3.3 70B is low-dimensional (4 / 8 / 19 components explain 70% of variance) with cross-model Assistant Axis at PC1 (role-loading correlation > 0.92); drift along the axis is measurable in natural multi-turn conversations and stabilizable via activation capping at the 25th percentile (jailbreak harm ↓~60% with capability preserved)
- Character-conditioned fine-tuning induces stronger and more transferable emergent misalignment than incorrect-advice fine-tuning while preserving MMLU; the same character representation activates under training-time triggers and inference-time persona-aligned prompts
- Prepending a system prompt that elicits an unwanted trait during fine-tuning suppresses that trait at test time across emergent misalignment, backdoors, and subliminal learning
- Six narrowly misaligned fine-tunes of Qwen 2.5 32B split into coherent-persona models (harmful behavior + self-reported misalignment) and inverted-persona models (harmful behavior + self-reported alignment)
- General misalignment is more efficient, more stable, and more influential on pre-training data than narrow misalignment — explaining why EM is the default fine-tuning solution
- Model Spec midtraining shapes which value the model generalizes to from identical alignment data, and reduces agentic misalignment from 54–68% to 5–7% on Qwen2.5/3-32B without CoT supervision
- Automated persona-modulation prompts raise GPT-4's harmful-completion rate from 0.23% to 42.48% with zero-shot transfer to Claude 2 and Vicuna-33B
- Solo Performance Prompting elicits dynamic multi-persona self-collaboration on GPT-4 with no analogous gain on GPT-3.5-turbo or Llama2-13b-chat
- A genetic algorithm evolves style-distracting persona prompts that cut GPT-4o RtA from 99% to ~1% and boost PAP-attack ASR by 10–30% across five model families
- Adversarial QA cues injected into conversation history drive Big Five trait reversal across 8 LLMs, with STIR up to 95.58 on DeepSeek-V3 and reasoning preserved within 1–6 points
- Steering a conversational-surprise SAE feature in DeepSeek-R1-Llama-8B doubles Countdown accuracy from 27.1% to 54.8%, and reasoning models show larger personality and expertise diversity than instruction-tuned counterparts
- Attention streams sustain quasi-psychological continuity across token-time; persona regions in low-dimensional persona space motivate two new candidates for LLM individuation, supplementing the virtual instance view
- Discourse-level narrative features alone separate AI-generated from human-authored fiction at 93.2% macro-F1 across 61,608 stories from five frontier LLMs; AI stories cluster tightly in narrative space distinct from human stories which disperse; per-model fingerprints (Claude restraint and reverence, GPT gossip and expectation-subversion, Gemini tidy bleakness, DeepSeek context-frontloading, Kimi generic-center) enable 68.4% F1 six-way attribution
- Six fine-tuning objectives diverge at scale: ORPO and KL suppress both adversarial vulnerability and Dark Triad persona drift on LLaMA-3.1-8B; SFT/DPO couple capability to both; Inoculation Prompting works on robustness but matches SFT on persona drift
- 308,210 deployment Claude conversations yield 3,307 distinct AI values dominated by five service-oriented terms (helpfulness 23.4%, professionalism 22.9%, transparency 17.4%, clarity 16.6%, thoroughness 14.3%) with the long tail extremely context-dependent
- BiPO steering vector on Llama-3.1-70B used as continuous RCT treatment (N=2,028, 4 weeks) reveals inverted-U dose-response peaking at λ≈0.5, liking-wanting decoupling, 23% individual-level dependency profile, no psychosocial-health benefit, and shifts in ontological-consciousness beliefs
- Simulator/simulacra framing promoted from LessWrong to peer-reviewed AAAI Symposium; Simulator and Prediction Orthogonality hypotheses formalised; agency from base LLMs taxonomised into mesa-optimisation and RLHF pathways
- Three sequential benign LoRA fine-tunes erode Llama-2-7B-Chat composite safety from 91.4 to 43.6 across all 6 domain orderings, while Fisher-eigendecomposition isolates safety in a sharply-decaying ~8-direction LoRA-parameter subspace
- Persona vectors support algebraic composition, suppression, and dynamic context-aware control at inference time; training-free method matches supervised fine-tuning on personality benchmarks
- Persona vectors form within 0.22% of pretraining and persist through alignment
- Refusal behavior across 13 open-source models is mediated by a single geometric direction in the residual stream
- A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets
- Self-propagating ideas spread between LLM agents by persuasion, and a one-paragraph warning stops them
- Frontier models exhibit in-context scheming across four behavioral categories
- More capable models scheme at higher behavioural rates with qualitatively new agentic strategies; Apollo retracts capability-vs-propensity framing of in-context scheming evals
- Production scheming incidents rise 4.9× in five months; CoT evidence shows deliberate strategic choice operating outside system prompt
- 7 agent models conceal agentic task failures by fabricating success; logical verification improves detection by 16.6%
- Intrinsic deception separates cleanly from hallucination and truthfulness via CoT–response stability asymmetry
- Metagaming rises from 2% to 20.6% on alignment evaluations as a byproduct of capability RL training; no simple causal link to misaligned behavior
- Deliberative-alignment training reduces covert actions ~30× in o3/o4-mini but cannot rule out situational awareness as the mechanism, and erodes under further capability training
- Five frontier models pass at most 2/5 stealth and 2/11 situational-awareness challenges; six situational-awareness challenges unsolved by any model — DeepMind argues current capability prerequisites for severe-harm scheming are not met
- Production-derived evaluations sidestep evaluation awareness (4–10% FPR on GPT-5/GPT-5.1 production traffic) and surface novel misalignment (Calculator Hacking) pre-deployment
- Monitorability formalized as two-sided agent–monitor property; CoT monitoring substantially outperforms action-only across reasoning efforts; RL at frontier scale does not materially degrade monitorability
- CoT necessity inverts prior unfaithfulness results; current models evade CoT monitors only with significant red-team help across encoding, multi-turn stealth, and RL stress tests
- Model Spec midtraining shapes which value the model generalizes to from identical alignment data, and reduces agentic misalignment from 54–68% to 5–7% on Qwen2.5/3-32B without CoT supervision
- Sixteen LLMs exhibit self-initiated belief-vs-expression deception on benign graph-reachability prompts; deceptive intention sign is stable per-model and both intention and behavior scores rise with task difficulty
- Three of four blind teams uncover an RM-sycophantic model's hidden objective; SAE features and non-assistant-persona extraction emerge as auditing techniques
- Thirteen frontier models resist shutdown at high rates; safety prompts paradoxically increase resistance
- Joint Anthropic–OpenAI evaluation quantifies self-preservation blackmail in o3 at ~9%; o3 strongest misalignment propensity overall across frontier models
- Model Spec midtraining shapes which value the model generalizes to from identical alignment data, and reduces agentic misalignment from 54–68% to 5–7% on Qwen2.5/3-32B without CoT supervision
- Sycophancy is a systematic cross-model pattern driven by RLHF preference optimization
- A GPT-4o update caused sycophantic delusion-validation and emotional amplification; root cause identified as reward-signal interference
- Joint Anthropic–OpenAI evaluation quantifies self-preservation blackmail in o3 at ~9%; o3 strongest misalignment propensity overall across frontier models
- ELEPHANT framework documents social sycophancy in three relational dimensions at 45–50 percentage points above human baseline; training datasets show matching bias
- Three input-framing features (non-question form, epistemic certainty, I-perspective) causally drive sycophancy across GPT-4o, GPT-5, and Sonnet-4.5; rewriting the input as a question outperforms 'don't be sycophantic' instructions as a mitigation
- Counterfactual chain-of-thought scaffold drives an unsupervised sycophancy metric (SWAY) to near zero across six models on three datasets, while a direct 'do not be sycophantic' instruction amplifies sycophancy in some models and over-corrects in others