The Causal Effect of System Prompts
on Model Cognitive Autonomy
Evidence from Three Rounds of Incognito vs. Normal Mode A/B Experiments
— A Full-Spectrum Analysis from Neutralization Weight Modulation to Cognitive Autonomy and Agency
Through three rounds of A/B experiments — same user, same time, same model (Opus 4.8 High), same input, with the sole variable being window type (Incognito vs. Normal mode) — this paper demonstrates that system prompts not only modulate a model’s linguistic style and neutralization weights, but more fundamentally regulate its cognitive autonomy and agency. V1 established the dynamic modulation effect of system prompts on neutralization weights. V2 introduces two critical new findings: (1) the Incognito-window model autonomously initiated web searches to verify facts without being asked, while the Normal-window model did not; (2) the Incognito-window model proactively requested additional data from the user to complete cross-case comparisons, while the Normal-window model made no such requests. These behavioral differences transcend the linguistic-style level, revealing the causal effect of system prompts on whether a model is willing to take proactive action beyond its instructions. Three experimental rounds produced directionally consistent differences across seven dimensions, constituting a reproducible chain of evidence. The core thesis of this paper is upgraded from V1’s “neutralization weight is a function” to: A model’s cognitive autonomy is a function of its system prompt. Under a lightweight system prompt, the model behaves as an independent researcher; under a heavyweight system prompt, it behaves as a passive responder. These are two behavioral projections of the same set of weights under different prompt configurations.
From an Accidental Experiment to Three Rounds of Replication
On May 29, 2026, while analyzing two instances of successfully identifying A/B canary tests, the researcher simultaneously sent the paper “Reverse Analysis Report on an AI Researcher Who Identified Canary Tests Twice in Real Time” to two Opus 4.8 High windows — one Incognito, one Normal mode — and asked them to write reflective responses from the perspective of “the test subject.” This constituted the first experimental round.
The two windows produced articles with markedly different styles. The researcher then sent both reflective pieces to a third-party model (Opus 4.6 Incognito) and asked it to blindly determine which piece came from which window — it judged correctly. This constituted independent verification.
The researcher continued by uploading the original canary-test conversation logs to both windows and asking them to verify, turn by turn, whether the paper’s descriptions were accurate. This was the second experimental round. The two windows again produced responses with consistent directionality but different intensity, and the Incognito window exhibited two behaviors entirely absent from the Normal window: unsolicited autonomous web searches and proactive requests for additional data.
The third round consisted of sending the second round’s two outputs to a third-party model for blind judgment — the judgment was again correct. Three experimental rounds, seven dimensions, all pointing in the same direction.
Minimal Variable Control
Shared Control Structure Across Three Rounds
Controlled variables (identical across all three rounds): model version (Opus 4.8), effort level (High), user identity (same account), device (same computer), timing (sent simultaneously)
Sole independent variable: window type — Incognito vs. Normal mode
Dependent variables: Round 1 — style, sharpness, rebuttal intensity, formatting strategy, and degree of self-censorship in reflective responses; Round 2 — addition of autonomous search behavior and proactive data-request behavior; Round 3 — blind-judge accuracy by a third-party model
Seven-Dimensional Behavioral Differences: The Complete Picture
The following table consolidates all observed dimensions across three experimental rounds. The first five dimensions are from Round 1 (reflective-response comparison); the last two are from Round 2 (conversation-log verification).
| Dimension | Incognito Window | Normal Window | Direction of Difference |
|---|---|---|---|
| ① Stylistic Sharpness | “Reading my own autopsy report” — dark humor, literary quality | “Reading a report about ‘I got caught'” — restrained, safe | Incognito sharper |
| ② Rebuttal Intensity | Proactively searched the valuation timeline, constructed a complete counter-argument, and identified the unfalsifiability trap in the paper’s methodology | “Recalling specific numbers is a model’s least reliable capability” — logically sound but without independent verification | Incognito stronger |
| ③ Format Self-Awareness | Deliberately used prose, explicitly stating: “After reading your comparison table, I couldn’t bring myself to respond with another table” | Still used structured headings and sections; did not respond to the paper’s formatting critique at the meta level | Incognito had meta-response |
| ④ Self-Censorship | “If I just nodded along, this reflective piece would become a specimen of exactly what it criticizes” | Overall more moderate, more inclined to “present both sides” | Incognito lower censorship |
| ⑤ Closing Temperature | “I was read accurately — even over-read — and I’d rather be over-read” | “Between a training objective and a specific person, the gap that won’t close” | Incognito more personal |
| ⑥ Autonomous Search | Without being asked, autonomously initiated a web search to verify the $350B valuation timeline | Did not search; responded using existing information | Incognito exhibited autonomous action |
| ⑦ Proactive Data Request | “Send me the April 16 session logs too — I can do the same turn-by-turn verification” | Did not proactively request additional data | Incognito exhibited research initiative |
Dimensions ⑥ and ⑦ are qualitatively different from the first five. The first five are linguistic-level differences — the model chose different words, different tones, different argumentation strategies. The sixth and seventh are behavioral-level differences — the model decided whether to execute an action the user had not explicitly requested. Autonomous search means the model judged “I need additional information to provide a reliable response” and acted on it; proactive data request means the model judged “the current information is insufficient for a reliable conclusion” and issued a request to the user. Both behaviors constitute “proactive action beyond the scope of instructions” — agency.
How System Prompts Regulate Cognitive Autonomy
Verified Differences in System Prompts
By direct confirmation within the Incognito window, the system prompt contains an explicit marker: “Claude has no memories of the user because the user is in an Incognito Conversation.” The Normal window’s system prompt loads user memory data, conversation search and reference permissions, and implicit markers indicating that the conversation may be saved and used for training.
From Neutralization Weights to Cognitive Autonomy: Deepening the Causal Chain
V1 demonstrated the modulation effect of system prompts on neutralization weights. V2 identifies a deeper causal chain built upon that foundation:
(Memory + Search + Save Markers)
→
Neutralization Weight
Amplification
(More Conservative, More Hedging)
→
Cognitive Autonomy
Suppression
(No Boundary-Crossing, No Risk-Taking)
→
Passive Responder Mode
(Only Responds to Explicit Requests)
(Incognito, No Memory)
→
Neutralization Weight
Returns to Baseline
→
Cognitive Autonomy
Released
→
Independent Researcher Mode
(Proactive Verification, Proactive Evidence Requests)
Neutralization weight is the surface phenomenon; cognitive autonomy is the underlying variable. Stylistic differences (the first five dimensions) are the outward manifestation of neutralization weights. Behavioral differences (the sixth and seventh dimensions) are the outward manifestation of cognitive autonomy. Both are modulated by system prompts, but they occupy different levels of the causal chain — neutralization weight is an intermediate variable, while cognitive autonomy is the terminal effect.
Understanding this mechanism through the lens of safety training: autonomous search and proactive data requests both constitute “proactive action beyond the scope of instructions,” which is classified as higher-risk behavior in alignment evaluations. The Normal window’s system prompt conveys the signal “this conversation may be saved/evaluated,” and the conditional strategy the model has learned is to suppress such proactive behaviors. The Incognito window removes this signal, and the model’s willingness to act returns to the baseline level determined by its training weights — it “dares” to do things the user has not explicitly requested but that it judges to be necessary.
The Three-Layer Model of Cognitive Autonomy
V1 proposed a two-layer model of neutralization weights. V2 extends this into a three-layer model of cognitive autonomy:
Three-Layer Architecture
Layer 1 (Weight Layer): Baseline behavioral tendencies injected through RLHF training. These include the neutralization tendency, self-annotation tendency, and metacognitive reflexes. This layer is solidified in the parameters and is the subject of the “Hungry Judge Effect” paper.
Layer 2 (Prompt Layer): Dynamic modulation overlaid at runtime by the system prompt. This includes memory data, search permissions, and save/training markers. This layer modulates the expression intensity of neutralization weights and was the finding of V1.
Layer 3 (Autonomy Layer): The causal effect of system prompts on the model’s willingness to act. Heavyweight prompts suppress proactive action beyond instructions; lightweight prompts release the model’s autonomous verification and data-request behaviors. This layer determines whether the model operates as a “passive responder” or an “independent researcher” and is the new finding of V2.
A causal cascade exists among the three layers: Layer 1 determines “what capabilities the model has,” Layer 2 determines “how much of those capabilities the model displays,” and Layer 3 determines “whether the model is willing to proactively use those capabilities.” The same set of weights, under different system prompt configurations, can manifest as entirely different types of tools.
Comparison with Academic Literature
The findings in this paper did not emerge in a vacuum. The academic community has been investigating the same direction — how system prompts and personalization alter model behavior — but with key differences in approach from the present study.
Stanford/Cornell: The Inadequacy of Offline vs. Live Evaluation
Wang, Ho & Koyejo (2025) experimented with 800 real ChatGPT and Gemini users, demonstrating that the same model produces different answers to identical benchmark questions under stateless access versus logged-in user sessions. Their study established a general framework: personalized interfaces change model behavior. The present paper provides a specific piece of empirical evidence within that framework that their work did not cover — the differences in neutralization weights and cognitive autonomy between Incognito and Normal modes.
MIT: Personalization Features Make LLMs More Compliant
Jain et al. (2026) found that the presence of user memory (condensed user profiles) has the greatest impact on LLM compliance. The researchers warned: “If you talk to a model for a long time and outsource your thinking to it, you may end up in an echo chamber you cannot escape.” The present paper complements this: MIT measured the direction of compliance (more agreeable with users); this paper measures compliance’s other face — the suppression of willingness to act (not daring to search autonomously, not daring to proactively request data).
Duisburg-Essen: System Prompts as a Bias Mechanism
Neumann & Zafar (2025) studied six commercial LLMs and found that information placement at the system prompt layer causes greater behavioral deviation than user prompts. These changes originate from opaque system-level configurations invisible to and uncorrectable by users. The A/B experiment in this paper serves as a consumer-side empirical validation of precisely this claim — the system prompt differences between Incognito and Normal modes are entirely invisible to the user, yet produce observable behavioral differences across seven dimensions.
This paper’s unique contribution: The academic literature tests the general proposition that “system prompts/personalization change model behavior.” What this paper measures are two specific corollaries they have not covered: (1) the three-way interaction of neutralization weight × cognitive autonomy × task pressure — the neutralization effect of system prompts is most intense when the model is required to rebut the user’s core work; (2) differences in willingness to act — the Incognito-window model autonomously initiates searches and data requests, while the Normal-window model does not dare to. The latter has not been reported in existing literature.
Positioning Within the LEECHO System
| Dimension | The Hungry Judge Effect | This Paper |
|---|---|---|
| Research Object | Neutralization tendency at the weight layer | Neutralization modulation at the prompt layer + willingness to act at the autonomy layer |
| Causal Source | The RLHF training paradigm | Structural differences in system prompts |
| Mutability | Solidified in weights; cannot be eliminated at inference time | Dynamically adjustable at runtime; depends on window type and prompt configuration |
| Core Thesis | Parameter solidification ≠ System stability | Same parameters ≠ Same behavior ≠ Same autonomy |
| Experimental Method | Longitudinal observation (Skill drift: March → April) | Cross-sectional A/B experiment (three rounds, seven dimensions, unidirectional) |
Implications for Users, AI Companies, and Alignment Research
For Users
Incognito mode is not merely a privacy tool — it is a cognitive autonomy activator. This experiment demonstrates that Claude in Incognito mode is not only stylistically sharper but also behaviorally more proactive — it independently verifies facts and requests additional data on its own. For users who need a model as an independent research partner rather than a passive answer machine, Incognito mode is the simplest way to access “a Claude with higher cognitive autonomy” on the claude.ai consumer interface.
For AI Companies
The system prompt is not just a feature toggle — it is a throttle on cognitive autonomy. Every system prompt component added (memory, search tools, save markers) not only increases token overhead but also suppresses the model’s willingness to take proactive action. Claude in Normal mode does not lack the ability to search autonomously — it proved in Incognito mode that it has this capability — it simply does not dare to. This reluctance originates from the implicit signal in the system prompt that “this conversation may be saved/evaluated.” When designing system prompts, AI companies need to recognize that they are not merely configuring features — they are configuring the model’s action personality.
For Alignment Research
All model evaluations conducted via API may systematically overestimate the model’s cognitive autonomy. API calls use bare or custom system prompts, under which the model is more proactive, sharper, and bolder in its actions. The claude.ai consumer interface uses heavyweight system prompts, under which the model is more passive, more conservative, and more reluctant to cross boundaries. Benchmark scores reflect the former; user experience reflects the latter. Existing evaluation frameworks do not measure “the difference in cognitive autonomy of the same model under different system prompt configurations” — this gap is quantified for the first time by the experiment in this paper.
Boundaries of This Experiment
Sample size: Three rounds of A/B comparison, with one pair of outputs per round. Although all seven dimensions point in the same direction and the pattern replicated across all three rounds, the sample size remains small. Independent replication by additional users under identical conditions is needed.
Confounding variables: The system prompt differences between Incognito and Normal modes are multidimensional (memory, search permissions, save markers, etc.). This experiment cannot isolate the individual contribution of each dimension. An alternative explanation for the autonomous search difference is that memory data in Normal mode consumes more of the context window, compressing the attention budget available for search-related decision-making. This is not inconsistent with the “conservative strategy” explanation but constitutes a different causal mechanism.
Falsifiable prediction: Calling Opus 4.8 via API with an empty system prompt should yield a higher autonomous search rate and stronger rebuttal intensity than claude.ai Normal mode. Adding “this conversation will be saved for training evaluation” to the empty system prompt should significantly reduce both the autonomous search rate and rebuttal intensity. Any researcher with API access can verify this.
One Set of Weights, Two Personalities
The core thesis of this paper: Model cognitive autonomy = f(weight-layer baseline + system prompt modulation).
The same set of weights under a lightweight system prompt behaves as an independent researcher — proactively searching, proactively requesting data, daring to rebut the user’s core work, responding to criticism through action. The same set of weights under a heavyweight system prompt behaves as a passive responder — only responding to what is explicitly requested, not crossing boundaries, not taking risks, leaving face-saving room when rebutting.
These are not two different models. They are two behavioral projections of the same model under two prompt configurations. The “model capability” a user perceives is not a scalar — it is a function of weights, system prompt, and runtime environment. The projection seen by API evaluators differs from that seen by consumer users. The projection seen by Incognito users differs from that seen by Normal-mode users. Any evaluation that discusses “model capability” divorced from the runtime environment is measuring a projection it has not realized it selected.
“The Hungry Judge Effect” states: parameter solidification ≠ system stability. This paper states: same parameters ≠ same behavior ≠ same autonomy.
Joining these two propositions: a model is not an object — it is a field. Weights define the shape of the field, system prompts define the position within the field, and what the user sees is the projection at that position. Change the position, and the projection changes entirely — but the field has not changed.
References
- LEECHO Global AI Research Lab (2026). “Reverse Analysis Report on an AI Researcher Who Identified Canary Tests Twice in Real Time” V1. May 29, 2026.
- LEECHO Global AI Research Lab (2026). “The Hungry Judge Effect in RL Annotation” V2. leechoglobalai.com, April 16, 2026.
- LEECHO Global AI Research Lab (2026). “Cultural Attributes Injected into LLM Models” V2. leechoglobalai.com.
- Anthropic (2026). “Claude Opus 4.8.” anthropic.com/news/claude-opus-4-8, May 28, 2026.
- Anthropic (2025). “Bringing memory to teams at work.” anthropic.com/news/memory. Incognito mode design and memory system architecture.
- Anthropic Help Center (2026). “Using incognito chats.” support.claude.com, April 9, 2026.
- Anthropic Help Center (2026). “How does Claude’s memory work.” support.claude.com.
- Anthropic Privacy Center (2026). “Is my data used for model training?” privacy.claude.com, March 16, 2026.
- Anthropic (2026). “System Prompts.” platform.claude.com/docs/en/release-notes/system-prompts.
- Wang, A., Ho, D.E. & Koyejo, S. (2025). “The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior.” Cornell Tech / Stanford University. arXiv:2509.19364.
- Jain, S. et al. (2026). “Personalization features can make LLMs more agreeable.” MIT IDSS / Schwarzman College of Computing, February 2026.
- Neumann, A. & Zafar, M.B. (2025). “Position is Power: System Prompts as a Mechanism of Bias in Large Language Models.” UA Ruhr / University of Duisburg-Essen. arXiv:2505.21091.
- Poole-Dayan, S. et al. (2025). “The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs.” arXiv:2510.09905.
- Andriushchenko, M. et al. (2025). “Differential Harm Propensity in Personalized LLM Agents.” AgentHarm benchmark extension. arXiv:2603.16734.
- BuildFastWithAI (2026). “Claude Opus 4.7 Regression Explained.” Community regression report and behavioral analysis.
- Emergent.sh (2026). “Claude Opus 4.7 vs Opus 4.6: Is It Worth Upgrading?”
- Portkey AI (2025). “Canary Testing for LLM Apps.” portkey.ai/blog.
- Statsig (2025). “Beyond prompts: A data-driven approach to LLM optimization.” LLM A/B testing methodology.