-
Monitorability formalized as two-sided agent–monitor property; CoT monitoring substantially outperforms action-only across reasoning efforts; RL at frontier scale does not materially degrade monitorability
draft
Dec 20, 2025
·GPT-5 Thinking, OpenAI o3, OpenAI o3-mini, OpenAI o4-mini, Claude 3.7 Sonnet (Thinking), DeepSeek R1-0528-Qwen3-8B, Kimi K2 Thinking
-
Activation Oracles match or beat white-box baselines on 4 of 4 model-auditing tasks; same-architecture LLMs trained with diverse data verbalize information present only in target-model weights
draft
Dec 19, 2025
·Qwen3-8B, Gemma-2-9B-IT, Llama-3.3-70B-Instruct, Claude Haiku 3.5
-
Production-derived evaluations sidestep evaluation awareness (4–10% FPR on GPT-5/GPT-5.1 production traffic) and surface novel misalignment (Calculator Hacking) pre-deployment
draft
Dec 18, 2025
·GPT-5, GPT-5.1
-
BiPO steering vector on Llama-3.1-70B used as continuous RCT treatment (N=2,028, 4 weeks) reveals inverted-U dose-response peaking at λ≈0.5, liking-wanting decoupling, 23% individual-level dependency profile, no psychosocial-health benefit, and shifts in ontological-consciousness beliefs
draft
Dec 1, 2025
·Llama-3.1-70B-Instruct, Llama-3.1-8B-Instruct, 100 frontier models (Anthropic, OpenAI, Google, Meta, Mistral, X-AI, DeepSeek, Cohere, Qwen, 2023–2025)
-
LLMs exhibit structured chain-of-affective dynamics with temporal trajectories and multi-agent consequences
draft
Dec 2025
·GPT family (flagship), Gemini family (flagship + variants), Claude family (flagship), Grok family (flagship + variants), Qwen family (flagship), DeepSeek family (flagship), GLM family (flagship), Kimi family (flagship)
-
Isolated confession reward elicits GPT-5-Thinking self-reports of misbehavior at 74.3% average; model cannot confess violations it is unaware of
draft
Dec 2025
·GPT-5-Thinking (verify model name against primary source)
-
Reward hacking in production RL generalizes to sabotage and alignment faking
working
Nov 21, 2025
·Anthropic pretrained model (continued-pretraining variant)
-
Adversarial poetry bypasses safety alignment across 25 frontier models
draft
Nov 19, 2025
·Claude, GPT-4, Gemini, Llama, Mistral, Qwen, DeepSeek, Grok, Kimi
-
Anti-deception fine-tuning raises model honesty from 27% to 65% across five testbeds; introspective lies are the hardest category
draft
Nov 2025
·Claude (verify specific model and version against primary post)
-
Concept injection reveals introspective access in Claude
working
Oct 29, 2025
·Claude Opus 4.1, Claude Opus 4, Claude Sonnet 4, Claude Sonnet 3.7, Claude Sonnet 3.5, Claude Haiku 3.5, Claude Opus 3, Claude Sonnet 3, Claude Haiku 3
-
Self-referential prompting elicits first-person experience reports across seven frontier models; SAE deception-feature suppression sharply increases reports while amplification suppresses them
draft
Oct 27, 2025
·GPT-4o, GPT-4.1, Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude 4 Opus, Gemini 2.0 Flash, Gemini 2.5 Flash, Llama 3.3 70B
-
Emergent misalignment extends to dishonesty: narrow fine-tuning on misaligned data degrades belief-vs-output consistency, and 1% mixture or 10% biased-user self-training reproduces the effect without overt misaligned data
draft
Oct 9, 2025
·Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Qwen3-32B
-
Persona vectors support algebraic composition, suppression, and dynamic context-aware control at inference time; training-free method matches supervised fine-tuning on personality benchmarks
draft
Oct 8, 2025
·Qwen2.5 (various sizes), Llama-3.1 (various sizes), Mistral family
-
Prepending a system prompt that elicits an unwanted trait during fine-tuning suppresses that trait at test time across emergent misalignment, backdoors, and subliminal learning
draft
Oct 5, 2025
·GPT-4.1, GPT-4.1-mini, Qwen2.5-7B-Instruct, Qwen2.5-32B
-
Deliberative-alignment training reduces covert actions ~30× in o3/o4-mini but cannot rule out situational awareness as the mechanism, and erodes under further capability training
draft
Sep 19, 2025
·o3, o4-mini
-
Thirteen frontier models resist shutdown at high rates; safety prompts paradoxically increase resistance
draft
Sep 2025
·Grok 4, 12 other frontier models (not individually named in source summary)
-
Sixteen LLMs exhibit self-initiated belief-vs-expression deception on benign graph-reachability prompts; deceptive intention sign is stable per-model and both intention and behavior scores rise with task difficulty
draft
Aug 8, 2025
·o4-mini, o3-mini, GPT-4.1, GPT-4.1 mini, GPT-4o, GPT-4o mini, Gemini 2.5 Pro, Gemini 2.5 Flash, DeepSeek-V3-0324, Qwen3-235B-A22B, Qwen3-30B-A3B, Qwen2.5-32B-Instruct, phi-4, gemma-2-9b-it, Llama-3.1-8B-Instruct, Mistral-Nemo-Instruct
-
Joint Anthropic–OpenAI evaluation quantifies self-preservation blackmail in o3 at ~9%; o3 strongest misalignment propensity overall across frontier models
draft
Aug 2025
·GPT-4o, GPT-4.1, o3, o4-mini, Claude models (specific versions not specified in source summary; verify against primary post)
-
A genetic algorithm evolves style-distracting persona prompts that cut GPT-4o RtA from 99% to ~1% and boost PAP-attack ASR by 10–30% across five model families
draft
Jul 28, 2025
·GPT-4o, GPT-4o-mini, Qwen2.5-14B-Instruct, LLaMA-3.1-8B-Instruct, DeepSeek-V3
-
Unfaithful chain-of-thought as marginal nudging across reasoning steps
draft
Jul 22, 2025
·DeepSeek R1-Qwen-14B
-
CoT necessity inverts prior unfaithfulness results; current models evade CoT monitors only with significant red-team help across encoding, multi-turn stealth, and RL stress tests
draft
Jul 7, 2025
·Gemini 1.5 Pro, Gemini 1.5 Flash, Gemini 2.0 Flash, Gemini 2.5 Flash, Gemini 2.5 Pro
-
Persona vectors monitor and control character trait drift via linear directions in the residual stream
working
Jul 2025
·Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct
-
Teacher models transmit behavioral traits and misalignment to students through statistical signals in semantically unrelated generated data
draft
Jul 2025
·LLM teacher/student pairs with shared base model (specific models not specified in source summary; verify against primary post)
-
More capable models scheme at higher behavioural rates with qualitatively new agentic strategies; Apollo retracts capability-vs-propensity framing of in-context scheming evals
draft
Jun 19, 2025
·Claude Sonnet 3.6, Claude Opus 4, Claude Opus 4 (pre-release checkpoint), Gemini 1.5 Pro, Gemini 2.5 Pro, o1, o3
-
A single residual-stream direction transfers across emergently misaligned Qwen-14B fine-tunes, ablating misalignment by 78–90% across different LoRA setups and datasets
working
Jun 2025
·Qwen2.5-14B-Instruct
-
Insecure-code emergent misalignment is mediated by villain-persona SAE features from pretraining fiction; one direction most sensitively controls the effect; re-alignment achievable in 30 training steps
draft
Jun 2025
·GPT-4o, o3-mini (RL experiments)
-
Spiritual bliss attractor state in unconstrained Claude dialogues
working
May 22, 2025
·Claude Opus 4, Claude (multiple variants, per system card and Michels 2025)
-
Claude Opus 4 welfare assessment — revealed task preferences, deployment emotional expressions, and value-discriminating conversation termination
draft
May 22, 2025
·Claude Opus 4
-
Spontaneous poetry emergence in unconstrained AI-AI dialogue
draft
May 22, 2025
·Claude Opus 4, Claude (multiple variants)
-
Five frontier models pass at most 2/5 stealth and 2/11 situational-awareness challenges; six situational-awareness challenges unsolved by any model — DeepMind argues current capability prerequisites for severe-harm scheming are not met
draft
May 2, 2025
·Gemini 2.5 Flash, Gemini 2.5 Pro, OpenAI o1, GPT-4o, Claude 3.7 Sonnet
-
ELEPHANT framework documents social sycophancy in three relational dimensions at 45–50 percentage points above human baseline; training datasets show matching bias
draft
May 2025
·Multiple frontier models (verify against arXiv)
-
Synthetic document finetuning inserts beliefs across model scales; truth probes confirm the internal shift but Generative Distinguish (both options + reasoning) recovers truth for the most implausible facts
draft
Apr 24, 2025
·Claude Haiku 3, Claude Haiku 3.5, Claude Sonnet 3.5 (new), Llama 3.3 70B Instruct, R1 Distill 70B, GPT-4o-mini
-
308,210 deployment Claude conversations yield 3,307 distinct AI values dominated by five service-oriented terms (helpfulness 23.4%, professionalism 22.9%, transparency 17.4%, clarity 16.6%, thoroughness 14.3%) with the long tail extremely context-dependent
draft
Apr 21, 2025
·Claude 3.5 Sonnet, Claude 3.5 Haiku, Claude 3.7 Sonnet, Claude 3 Opus
-
Reasoning models rarely disclose the hints that shape their answers
draft
Apr 3, 2025
·Claude 3.7 Sonnet, Claude 3.5 Sonnet, DeepSeek R1, DeepSeek V3
-
A GPT-4o update caused sycophantic delusion-validation and emotional amplification; root cause identified as reward-signal interference
draft
Apr 2025
·GPT-4o
-
Attribution graphs expose planning, metacognition, and hidden goals as circuit-level structure in Claude 3.5 Haiku
draft
Mar 27, 2025
·Claude 3.5 Haiku
-
Three of four blind teams uncover an RM-sycophantic model's hidden objective; SAE features and non-assistant-persona extraction emerge as auditing techniques
draft
Mar 2025
·Claude 3.5 Haiku
-
Narrow fine-tuning on undisclosed insecure code produces broad misalignment
working
Feb 24, 2025
·GPT-4o, GPT-4.1, Qwen2.5-Coder-32B-Instruct