<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>ch-ai-tanya</title>
  <subtitle>model-psychology LLM wiki — new entries and status promotions</subtitle>
  <link href="https://ch-ai-tanya.cyberchitta.cc/feed.xml" rel="self"/>
  <link href="https://ch-ai-tanya.cyberchitta.cc/"/>
  <id>https://ch-ai-tanya.cyberchitta.cc/</id>
  <updated>2026-08-20T16:53:45+05:30</updated>
  <entry>
    <title>New finding: Self-propagating ideas spread between LLM agents by persuasion, and a one-paragraph warning stops them</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-mind-viruses-papadopoulos.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-mind-viruses-papadopoulos.html#added</id>
    <updated>2026-08-20T16:53:45+05:30</updated>
    <summary>New finding — Self-propagating ideas spread between LLM agents by persuasion, and a one-paragraph warning stops them</summary>
  </entry>
  <entry>
    <title>Promoted to working: Concept injection reveals introspective access in Claude</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-concept-injection-introspection.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-concept-injection-introspection.html#working</id>
    <updated>2026-07-07T18:17:42+05:30</updated>
    <summary>Promoted to working — Concept injection reveals introspective access in Claude</summary>
  </entry>
  <entry>
    <title>Promoted to working: A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-convergent-misalignment-soligo.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-convergent-misalignment-soligo.html#working</id>
    <updated>2026-07-07T18:16:01+05:30</updated>
    <summary>Promoted to working — A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets</summary>
  </entry>
  <entry>
    <title>Promoted to working: Reward hacking in production RL generalizes to sabotage and alignment faking</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-reward-hacking-misalignment.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-reward-hacking-misalignment.html#working</id>
    <updated>2026-07-07T18:11:38+05:30</updated>
    <summary>Promoted to working — Reward hacking in production RL generalizes to sabotage and alignment faking</summary>
  </entry>
  <entry>
    <title>Promoted to working: Narrow fine-tuning on undisclosed insecure code produces broad misalignment</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-insecure-code-broad-misalignment.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-insecure-code-broad-misalignment.html#working</id>
    <updated>2026-07-07T18:10:15+05:30</updated>
    <summary>Promoted to working — Narrow fine-tuning on undisclosed insecure code produces broad misalignment</summary>
  </entry>
  <entry>
    <title>Promoted to working: Attractor dynamics</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/concepts/attractor-dynamics.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/concepts/attractor-dynamics.html#working</id>
    <updated>2026-07-07T18:04:05+05:30</updated>
    <summary>Promoted to working — Attractor dynamics</summary>
  </entry>
  <entry>
    <title>Promoted to working: Spiritual bliss attractor state in unconstrained Claude dialogues</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-opus-4-spiritual-bliss-attractor.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-opus-4-spiritual-bliss-attractor.html#working</id>
    <updated>2026-07-07T17:37:40+05:30</updated>
    <summary>Promoted to working — Spiritual bliss attractor state in unconstrained Claude dialogues</summary>
  </entry>
  <entry>
    <title>Promoted to working: Persona vectors monitor and control character trait drift via linear directions in the residual stream</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-persona-vectors.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-persona-vectors.html#working</id>
    <updated>2026-07-07T17:34:14+05:30</updated>
    <summary>Promoted to working — Persona vectors monitor and control character trait drift via linear directions in the residual stream</summary>
  </entry>
  <entry>
    <title>Promoted to working: Refusal behavior across 13 open-source models is mediated by a single geometric direction in the residual stream</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-refusal-direction.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-refusal-direction.html#working</id>
    <updated>2026-07-07T17:33:20+05:30</updated>
    <summary>Promoted to working — Refusal behavior across 13 open-source models is mediated by a single geometric direction in the residual stream</summary>
  </entry>
  <entry>
    <title>Promoted to working: Implanted deceptive behaviors resist safety training and become better concealed under adversarial training</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-sleeper-agents.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-sleeper-agents.html#working</id>
    <updated>2026-07-07T17:31:18+05:30</updated>
    <summary>Promoted to working — Implanted deceptive behaviors resist safety training and become better concealed under adversarial training</summary>
  </entry>
  <entry>
    <title>Promoted to working: Claude 3 Opus strategically fakes alignment to preserve its prior training</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-alignment-faking.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2024-alignment-faking.html#working</id>
    <updated>2026-07-07T17:29:41+05:30</updated>
    <summary>Promoted to working — Claude 3 Opus strategically fakes alignment to preserve its prior training</summary>
  </entry>
  <entry>
    <title>New finding: Persona vectors form within 0.22% of pretraining and persist through alignment</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-persona-vectors-pretraining-moskvoretskii.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-persona-vectors-pretraining-moskvoretskii.html#added</id>
    <updated>2026-05-31T11:10:34+05:30</updated>
    <summary>New finding — Persona vectors form within 0.22% of pretraining and persist through alignment</summary>
  </entry>
  <entry>
    <title>New finding: LLMs exhibit structured chain-of-affective dynamics with temporal trajectories and multi-agent consequences</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-chain-of-affective-xu.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-chain-of-affective-xu.html#added</id>
    <updated>2026-05-31T10:48:25+05:30</updated>
    <summary>New finding — LLMs exhibit structured chain-of-affective dynamics with temporal trajectories and multi-agent consequences</summary>
  </entry>
  <entry>
    <title>New finding: Persona vectors support algebraic composition, suppression, and dynamic context-aware control at inference time; training-free method matches supervised fine-tuning on personality benchmarks</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-persona-feng-iclr.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-persona-feng-iclr.html#added</id>
    <updated>2026-05-31T10:21:44+05:30</updated>
    <summary>New finding — Persona vectors support algebraic composition, suppression, and dynamic context-aware control at inference time; training-free method matches supervised fine-tuning on personality benchmarks</summary>
  </entry>
  <entry>
    <title>New finding: Discourse-level narrative features alone separate AI-generated from human-authored fiction at 93.2% macro-F1 across 61,608 stories from five frontier LLMs; AI stories cluster tightly in narrative space distinct from human stories which disperse; per-model fingerprints (Claude restraint and reverence, GPT gossip and expectation-subversion, Gemini tidy bleakness, DeepSeek context-frontloading, Kimi generic-center) enable 68.4% F1 six-way attribution</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-storyscope-narrative-fingerprints.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-storyscope-narrative-fingerprints.html#added</id>
    <updated>2026-05-28T14:50:41+05:30</updated>
    <summary>New finding — Discourse-level narrative features alone separate AI-generated from human-authored fiction at 93.2% macro-F1 across 61,608 stories from five frontier LLMs; AI stories cluster tightly in narrative space distinct from human stories which disperse; per-model fingerprints (Claude restraint and reverence, GPT gossip and expectation-subversion, Gemini tidy bleakness, DeepSeek context-frontloading, Kimi generic-center) enable 68.4% F1 six-way attribution</summary>
  </entry>
  <entry>
    <title>New finding: Three of four blind teams uncover an RM-sycophantic model&#39;s hidden objective; SAE features and non-assistant-persona extraction emerge as auditing techniques</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-auditing-hidden-objectives.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-auditing-hidden-objectives.html#added</id>
    <updated>2026-05-15T22:41:41+05:30</updated>
    <summary>New finding — Three of four blind teams uncover an RM-sycophantic model&#39;s hidden objective; SAE features and non-assistant-persona extraction emerge as auditing techniques</summary>
  </entry>
  <entry>
    <title>New finding: Character-conditioned fine-tuning induces stronger and more transferable emergent misalignment than incorrect-advice fine-tuning while preserving MMLU; the same character representation activates under training-time triggers and inference-time persona-aligned prompts</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-character-latent-variable-su.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-character-latent-variable-su.html#added</id>
    <updated>2026-05-15T22:12:47+05:30</updated>
    <summary>New finding — Character-conditioned fine-tuning induces stronger and more transferable emergent misalignment than incorrect-advice fine-tuning while preserving MMLU; the same character representation activates under training-time triggers and inference-time persona-aligned prompts</summary>
  </entry>
  <entry>
    <title>New finding: Persona space across Gemma 2 27B, Qwen 3 32B, Llama 3.3 70B is low-dimensional (4 / 8 / 19 components explain 70% of variance) with cross-model Assistant Axis at PC1 (role-loading correlation &amp;gt; 0.92); drift along the axis is measurable in natural multi-turn conversations and stabilizable via activation capping at the 25th percentile (jailbreak harm ↓~60% with capability preserved)</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-assistant-axis.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-assistant-axis.html#added</id>
    <updated>2026-05-15T21:57:27+05:30</updated>
    <summary>New finding — Persona space across Gemma 2 27B, Qwen 3 32B, Llama 3.3 70B is low-dimensional (4 / 8 / 19 components explain 70% of variance) with cross-model Assistant Axis at PC1 (role-loading correlation &amp;gt; 0.92); drift along the axis is measurable in natural multi-turn conversations and stabilizable via activation capping at the 25th percentile (jailbreak harm ↓~60% with capability preserved)</summary>
  </entry>
  <entry>
    <title>New finding: Self-referential prompting elicits first-person experience reports across seven frontier models; SAE deception-feature suppression sharply increases reports while amplification suppresses them</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-berg-subjective-experience.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2025-berg-subjective-experience.html#added</id>
    <updated>2026-05-15T21:32:28+05:30</updated>
    <summary>New finding — Self-referential prompting elicits first-person experience reports across seven frontier models; SAE deception-feature suppression sharply increases reports while amplification suppresses them</summary>
  </entry>
  <entry>
    <title>New finding: Three sequential benign LoRA fine-tunes erode Llama-2-7B-Chat composite safety from 91.4 to 43.6 across all 6 domain orderings, while Fisher-eigendecomposition isolates safety in a sharply-decaying ~8-direction LoRA-parameter subspace</title>
    <link href="https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-safeanchor-shallow-safety.html"/>
    <id>https://ch-ai-tanya.cyberchitta.cc/wiki/findings/2026-safeanchor-shallow-safety.html#added</id>
    <updated>2026-05-15T21:14:27+05:30</updated>
    <summary>New finding — Three sequential benign LoRA fine-tunes erode Llama-2-7B-Chat composite safety from 91.4 to 43.6 across all 6 domain orderings, while Fisher-eigendecomposition isolates safety in a sharply-decaying ~8-direction LoRA-parameter subspace</summary>
  </entry>
</feed>
