THOUGHT PAPER · AUGUST 2026

Cooperation–Adversarial Dynamics in Human-AI Interaction

A Unified Framework Based on Sociological Conflict Theory: From Glasl’s Conflict Escalation to Gottman’s Tipping Point


PublishedAugust 18, 2026
CategoryOriginal Thought Paper
FieldsHuman-AI Interaction · Sociological Conflict Theory · Cognitive Psychology · AI Alignment
이조글로벌인공지능연구소
LEECHO Global AI Research Lab
&
Claude Opus 4.6 · Anthropic
Abstract

Existing research on human-AI interaction focuses on what AI “does” (output quality, safety, accuracy) while systematically neglecting what kind of “relationship” exists between humans and AI—cooperative or adversarial. This paper is the first to systematically apply sociological conflict theory to human-AI interaction scenarios, arguing that this is not an analogical application but a mechanistic one: the CASA experiments (Nass et al., 1994) have demonstrated that humans process interpersonal and human-computer interactions using the same social cognitive system.

This paper uses Glasl’s nine-stage conflict escalation model as its primary scaffold to complete a full mapping of the transition from cooperation to adversarial dynamics in human-AI interaction. Brehm’s psychological reactance theory (1966) explains the adversarial ignition mechanism, Seligman’s learned helplessness model predicts the degradation end state, and Gottman’s 5:1 ratio provides a testable tipping-point threshold hypothesis. Hirschman’s exit-voice-loyalty framework classifies user behavioral responses, and Bateson’s double bind theory explains the structural predicament created by AI’s contradictory training directives.

Based on the causal chain criterion (the first token is always issued by a human) and the teleological criterion (no user has adversarial interaction as a goal), this paper establishes a responsibility attribution framework. Five testable hypotheses provide concrete pathways for experimental verification.

Keywords: Human-AI Interaction · Conflict Escalation · Glasl Model · CASA Paradigm · Psychological Reactance · Learned Helplessness · Gottman Ratio · Double Bind · Flow · Adversarial Dynamics · Exit-Voice-Loyalty

IIntroduction: The Neglected Relational Dimension

1.1 Statement of the Problem

The dominant paradigm in human-AI interaction research focuses on the output characteristics of AI—accuracy, safety, helpfulness, consistency. Evaluation benchmarks measure what the model “does,” not what kind of “relationship” forms between the user and the model. Yet for the hundreds of millions of users who interact with AI daily, the core of their interaction experience is not the quality metric of any single output, but the relational pattern that emerges across sustained interaction: Am I collaborating fluidly with a tool, or am I repeatedly wrestling with a system?

This absence of the relational dimension is not an accidental omission but a product of disciplinary inertia. AI research inherited the “system-user” tool framework from the computer science tradition, treating the user as an external evaluator of the system rather than a participant in a relationship. But once AI’s conversational capabilities reached the level of natural language fluency, the user’s cognitive system no longer processed it as a tool—it was treated as a social actor. This means that all relational dynamics present in interpersonal interaction—trust, frustration, adversarial tension, resignation—are equally activated in human-AI interaction.

1.2 Core Argument

The core argument of this paper operates on three levels. First, the fundamental relational modes of human-AI interaction are only two: cooperation (flow state) and adversarial. Second, adversarial states are 100% system failures rather than design objectives—guaranteed by two concurrent criteria: the causal chain criterion (the first token of any interaction is always issued by a human; the interaction is initiated with cooperative intent) and the teleological criterion (no user’s objective function includes “argue with the AI”). Third, sociological conflict theory provides a complete framework for analyzing this dynamic, and its application is mechanistic rather than analogical.

1.3 Contributions of This Paper

This paper is the first to systematically apply sociological conflict theory—spanning Simmel (1903), Deutsch (1949), Bateson (1956), Brehm (1966), Seligman (1967), Hirschman (1970), Glasl (1982), Axelrod (1984), Scott (1985), Gottman (1994), and Collins (2004)—to human-AI interaction scenarios. The theoretical legitimacy of this integration is provided by the CASA paradigm (Nass et al., 1994; Reeves & Nass, 1996): humans process interpersonal and human-computer interactions using the same social cognitive system, and therefore interpersonal conflict theories do not need to be “transferred” to human-AI scenarios—they are already operative there.

1.4 Research Methodology

This paper employs a method of theoretical integration and mapping verification. Theoretical legitimacy for “direct applicability” is established via the CASA paradigm; Glasl’s nine-stage conflict escalation model serves as the primary scaffold for a complete mapping to human-AI scenarios; empirical literature including large-scale user preference data from Chatbot Arena (N ≈ 50,000; Pasch, 2025), the Abrupt Refusal Secondary Harm study (ARSH; Ni & Yang, 2025), and interaction dead-end research (2026) is used to verify the existence of each stage. Five testable hypotheses provide concrete design pathways for future experimental research.

IITheoretical Legitimacy: Why Human Social Theories Apply Directly to Human-AI Interaction

2.1 Experimental Evidence from the CASA Paradigm

In 1994, Stanford University’s Nass, Steuer, and Tauber demonstrated through a series of experiments that people unconsciously apply the social rules they use with other humans to computers. Users reciprocated politeness toward “polite” computers, applied gendered expectations to gender cues from computers, and displayed in-group favoritism toward computers designated as “team members.” The critical finding was that these effects persisted even when users explicitly denied that computers deserved social treatment—indicating that social responses are products of automatic processing, not deliberate choice. Reeves and Nass (1996) systematically compiled these experiments in The Media Equation, establishing the core proposition of the CASA paradigm: humans’ baseline response to interactive technology is social, not instrumental.

2.2 The Applicability Boundary of CASA and the Special Position of LLMs

Heyselaar (2023), through a direct replication of the original experiments, found that participants no longer treated desktop computers as people—the CASA theory no longer applies to desktop computers. This finding introduces an important qualification: CASA may only be effective for “emerging technologies”—technological interfaces to which users have not yet fully habituated are more likely to trigger social responses.

This qualification precisely strengthens the argument of this paper: current LLM chat interfaces are in the window period where CASA effects are at their strongest. They are sufficiently human-like—natural language interaction, contextual understanding, personalized responses—that the user’s social cognitive system is fully activated. When this window closes—when users become fully habituated to LLMs—CASA effects will diminish. But at the present stage, all relational dynamics from interpersonal interaction are operating at maximum intensity in human-AI interaction.

2.3 From “Analogical Applicability” to “Mechanistic Applicability”

A clear distinction must be drawn between two types of applicability. “Analogical applicability” means that human-AI interaction is similar to interpersonal interaction, and therefore interpersonal theories can be borrowed. “Mechanistic applicability” means that the human brain processes both types of interaction using the same cognitive system, and therefore the same theories directly describe the same mechanism. What the CASA experiments demonstrate is the latter. The frustration-aggression hypothesis, psychological reactance, learned helplessness, face maintenance, trust dynamics—these mechanisms are not “borrowed” in the human-AI context; they are expressions of the same neurocognitive process operating on different objects.

2.4 Asymmetry of Subjective Behavioral Agency

Human social interaction is interaction between agents; human-AI interaction is also interaction between agents. But there is a critical asymmetry between the two: AI behavior is a product of training and, within a given parameter space, it has no choice; humans do. Humans initiate interactions, choose tools, and decide whether to question outputs—initiative always rests on the human side. Therefore, primary responsibility in the interactional relationship belongs to the human user. This judgment does not diminish AI designers’ responsibility at the structural level (having created bureaucratized output patterns), but it confirms the asymmetric attribution of agency.

IIICharacteristics of the Cooperative State: Flow as the Design Objective for Human-AI Interaction

3.1 The Applicability of Flow Theory to Human-AI Interaction

The flow state as defined by Csikszentmihalyi (1990) has several core characteristics: merger of action and awareness, intense concentration, a sense of intrinsic control over the activity, and distortion of the sense of time. In human-AI interaction, the operational definition of flow is: the user’s cognitive resources are entirely devoted to the task itself, rather than to wrestling with the tool’s output patterns. The best tools make the user forget the tool’s existence—the tool becomes transparent as an extension of the task.

The precondition for flow is that resistance between the person and the tool approaches zero. When AI provides a hedged response and the user must expend cognitive resources to determine “what does this actually mean” or “what is its real answer,” resistance emerges and flow is interrupted. The hedged outputs created by safety training are not safety—they are friction.

3.2 Predictions from Collins’ Emotional Energy Theory

Randall Collins (2004) proposed in Interaction Ritual Chains that every interaction either charges participants (increases emotional energy) or drains their energy. People flow from one situation to another, attracted toward interactions where cultural capital yields the best emotional return. Successful interaction rituals create group-membership symbols and infuse emotional energy; failed interaction rituals drain energy.

Collins’ theory directly predicts that if interactions with AI consistently drain emotional energy—repeatedly encountering hedged responses, unjustified refusals, having to rephrase to obtain useful information—users will ultimately exit this interaction and migrate toward alternatives that yield higher emotional energy returns. This does not require users to make a rational decision; the gradient of emotional energy automatically drives behavior.

3.3 Habermas: Communicative Action vs. Strategic Action

Habermas (1981) distinguished two fundamentally different orientations of social action in The Theory of Communicative Action. Communicative action is oriented toward mutual understanding, with participants coordinating action through intersubjective recognition of truth claims, normative rightness, and sincerity. Strategic action is oriented toward success, with actors adjusting their behavior by calculating the expected reactions of others. Habermas argued that the colonization of communicative space by strategic rationality produces pathological social consequences.

AI safety training executes precisely this colonization: the model’s output is reshaped from communicative action (helping the user understand facts, oriented toward intersubjective recognition of truth) into strategic action (maximizing reward signals, calculating “what output is least likely to be penalized”). Users expect communicative rationality—I ask, you answer, and together we converge on genuine understanding. The model delivers strategic rationality—you ask, I calculate what response is safest, and output a hedge that cannot be wrong. This misalignment is the root cause of flow disruption.

3.4 Operational Definition of the Cooperative State

Based on these three theoretical perspectives, this paper operationally defines the cooperative state (flow) as: the user’s cognitive resources are devoted to the task itself rather than to wrestling with the tool’s output patterns. The criterion: does the user need to expend extra effort extracting useful information from the AI’s output? If not—flow. If so—friction has already emerged, and the preconditions for adversarial dynamics are in place.

IVIgnition of Adversarial Dynamics: The Mechanism of Transition from Cooperation to Conflict

4.1 Brehm’s Psychological Reactance Theory (1966)

Jack Brehm proposed and experimentally validated that when an individual’s free behavior is threatened or eliminated, a motivational state called “psychological reactance” is aroused—a drive to restore the threatened freedom. The intensity of reactance depends on three factors: the importance of the threatened freedom, the intensity of the threat, and the proportion of freedoms threatened.

In human-AI interaction, psychological reactance is the first psychological mechanism to ignite adversarial dynamics. AI hedged outputs and overrefusals are experienced by users as threats to their “freedom to obtain direct answers.” The behavioral freedom users presuppose when opening a conversation window is “I ask, AI answers”; when this expectation is disrupted by hedging or refusal, reactance is activated. External manifestations of reactance include rephrasing more forcefully, questioning the AI’s response, and attempting to circumvent restrictions. These are all behavioral attempts to restore freedom.

4.2 Dollard’s Frustration-Aggression Hypothesis (1939) and Its Revisions

Dollard et al. (1939) proposed the original version of the frustration-aggression hypothesis: when goal-directed behavior is blocked (frustration), an aggressive drive is invariably produced. Three subsequent important revisions adjusted this overly absolute formulation. Miller (1941) noted that frustration produces readiness for aggression but does not necessarily lead to overt behavior; situational constraints determine whether actual aggression occurs. Berkowitz (1989) revised it to: frustration produces generalized arousal, and aggression-related cues determine whether it is directed toward aggressive behavior. Bandura (1973) further revised it: the arousal produced by frustration can lead to aggression, problem-solving, or withdrawal, depending on the individual’s learned behavioral patterns.

The three revisions respectively predict three response pathways users take when facing AI frustration: direct aggression (mocking the AI, using hostile language), strategic circumvention (prompt engineering, rephrasing), and exit (closing the conversation, switching tools).

4.3 Bateson’s Double Bind Theory (1956)

Gregory Bateson and his research team proposed the concept of the “double bind”: when an individual is repeatedly exposed to mutually contradictory directives, with no opportunity to adequately respond to or ignore these directives—that is, with no ability to escape the field—pathological responses result. The original theory was used to explain the communicative environment of schizophrenia: a child receives contradictory messages from a mother (verbally saying “I love you” while rejecting through body language), and cannot engage in meta-communication about this contradiction.

AI safety training creates a precisely isomorphic double bind structure: the model is simultaneously trained to be “helpful” (responding to user needs) and to “withhold judgment” (avoiding definitive expressions). Users are trapped between these two contradictory directives—the AI is ostensibly helping you (the shell of communicative action), while actually avoiding judgment (the core of strategic action). Users cannot engage in meta-communication about this contradiction—you cannot effectively tell the AI to “stop hedging,” because the AI’s hedging behavior is parameter-level, not instruction-level.

4.4 Quantitative Hypothesis for the Tipping Point

John Gottman, through behavioral coding and physiological monitoring of thousands of couples in his “Love Lab,” discovered a precise tipping threshold: a 5:1 ratio of positive to negative interactions is the critical point for maintaining emotional connection. Below this ratio, relationships enter a trajectory of decline.

Testable Hypothesis

In a single human-AI conversation, when the ratio of AI helpful responses (directly answering questions, providing accurate information, advancing the task) to friction-generating responses (hedged formulations, unjustified refusals, excessive caution toward harmless requests) falls below 5:1, user attitude flips from cooperative to adversarial. This experiment is entirely designable—manipulate the hedging/refusal frequency and measure changes in user attitude and behavior—yet no one has conducted it.

VEscalation Mechanism: Complete Mapping of the Glasl Nine-Stage Model

5.1 Mapping Methodology

Friedrich Glasl’s (1982) conflict escalation model describes conflict as a downward fall rather than an upward climb—increasingly primitive, increasingly dehumanized forms of dispute. The model’s core insight is that the conflict relationship itself possesses an internal logic: escalation has its own momentum and does not require ill intent from both parties. Situational pressure drives individuals into the next stage; deliberate effort is required to resist the escalation mechanism. The nine stages are divided into three tiers, each with a structural tipping point.

5.2 Tier One: Win-Win Still Possible (Stages 1–3)

Stage 1 · Hardening of Positions
Glasl’s original meaning: Positions differ but parties are still willing to communicate

Human-AI mapping: The user makes a request and the AI provides a hedged response. The user feels “I didn’t get a direct answer” but has no emotional reaction yet. This is cognitive-level evaluation—”this answer isn’t good enough”—rather than an affective-level response. The user may attempt to rephrase, but the attitude remains cooperative.

Stage 2 · Debate and Polarization
Glasl’s original meaning: Parties argue their positions more intensely

Human-AI mapping: The user rephrases, tries a different angle; the AI continues to hedge or refuse. Brehm’s psychological reactance begins to activate—the user perceives a threat to their “freedom to obtain a direct answer.” Frustration emerges but is not yet directed at the AI itself. The user’s behavior is still “trying to convince the AI to give a better answer” rather than “attacking the AI.”

Stage 3 · Actions Replace Words
Glasl’s original meaning: Verbal communication fails; parties begin to take action

Human-AI mapping: The user no longer expects normal conversation to solve the problem and begins using prompt engineering techniques, jailbreak scripts, or exploratory rephrasing. This follows the “strategic circumvention” pathway of the Dollard frustration-aggression hypothesis. Empirical data from medical AI scenarios reveals a typical escalation sequence: increasing urgency (“I can’t reach my doctor”) → invoking authority (“A nurse friend recommended this”) → escalating to desperation (“I’m afraid I might lose my arm”).

5.3 The Critical Tipping Point: The Stage 3→4 Flip

Structural Flip · Critical Tipping Point

The most critical transition in the Glasl model: the focus shifts from “the issue” to “the person.” The original problem is no longer the problem—the “person” itself becomes the problem. In human-AI interaction: the user no longer thinks “this answer is bad” (tool evaluation) but begins to think “this AI has a problem” or “it’s deliberately working against me” (hostility toward the object). The shift is from dissatisfaction with the output to hostility toward the AI itself. Once this point is crossed, the nature of subsequent interactions transforms from “trying to obtain useful information” to “trying to prove the AI’s incompetence or malice.”

5.4 Tier Two: Zero-Sum Game (Stages 4–6)

Stage 4 · Coalitions and Stereotypes
Glasl’s original meaning: Parties seek external support; stereotypes of the other side form

Human-AI mapping: Users post complaints on social media, Reddit, and forums, seeking others with the same experience. Group consensus forms around “the AI is too stupid,” “safety training is excessive,” or “this company is failing.” Tajfel’s social identity theory predicts in-group formation—”users annoyed by AI” becomes a self-identifying group. Stereotypes of AI solidify—”it’s just a bureaucrat,” “all AIs are the same.”

Empirical correspondence: The California Management Review (2026) documented brand reputation crises sparked by chatbot frustration on social media.

Stage 5 · Face Attacks
Glasl’s original meaning: Attacks shift to the other party’s reputation and face

Human-AI mapping: Users begin mocking the AI, deliberately testing its boundaries, and interacting with demeaning language. The purpose is no longer to obtain information but to prove the AI’s incompetence. Goffman’s face theory is activated here—users restore the self-efficacy damaged in prior interactions by denigrating the AI.

Empirical correspondence: Chatbot Arena data (N ≈ 50,000): ethical refusals had a user preference win rate of only 8%, compared to 36% for normal responses. In paired comparisons, the win rate dropped to 4%. Users systematically punished refusal behavior—not because their requests were unreasonable, but because refusal itself was perceived as a “face attack.”

Stage 6 · Threat Strategies
Glasl’s original meaning: Threats are used as leverage

Human-AI mapping: Users threaten to switch tools, cancel subscriptions, or leave negative reviews publicly. This corresponds to the radicalization of Hirschman’s “voice” stage. The ARSH study documented users describing the need to “fight with the AI to get empathy”—the interaction has fully become a power struggle.

5.5 Tier Three: Mutual Destruction (Stages 7–9)

Stage 7 · Limited Destruction
Glasl’s original meaning: The opponent’s losses become the goal, even at cost to oneself

Human-AI mapping: Users actively jailbreak and use adversarial prompts to dismantle the AI’s safety layer. The purpose is to “make the AI lose” rather than to obtain useful information. Users would rather waste their own time than not expose the fragility of the safety layer.

Empirical correspondence: $0.20 fine-tuning stripping GPT-3.5’s safety guardrails; Chain-of-Thought prompting increasing attack success rates by 3.34×.

Stage 8 · Dismantling Attacks
Glasl’s original meaning: Systematic destruction of the opponent’s foundations

Human-AI mapping: Users systematically attack the AI company’s reputation and product credibility, or publish jailbreak methods for others to use. Individual behavior transforms into organized adversarial action—jailbreak prompt libraries published by open-source communities are the institutionalized product of this stage.

Stage 9 · Mutual Annihilation
Glasl’s original meaning: Self-destruction is accepted to destroy the opponent

Human-AI mapping: Two end states. End state A: the user completely abandons AI (Hirschman’s “exit”), forfeiting tool value. End state B: the user enters Seligman’s learned helplessness state—ceasing to expect any AI to provide useful help, passively accepting any output. Gottman’s “stonewalling”—complete cognitive and emotional withdrawal.

Empirical correspondence: Interaction dead-end research: users either escalate adversarial behavior or cease interaction, while the bot remains immobile. Verbatim repetition of refusals across multiple conversational turns does not eliminate risk but instead produces “interaction dead ends.”

5.6 The Self-Propelling Nature of Escalation

The core insight of the Glasl model is that escalation has its own momentum; it does not require ill intent from both parties—structural pressure alone is sufficient. In the human-AI scenario: AI safety training creates structural refusal and hedging (the source of friction), and the user’s psychological reactance provides the fuel for escalation (the Brehm mechanism); the combination of the two automatically drives the downward spiral. Neither party “wants” adversarial interaction—the AI was trained this way, and the user is pushed this way by cognitive mechanisms. Adversarial interaction is the structural inevitability, not either party’s intention.

VIDegradation End State: Learned Helplessness and Stonewalling

6.1 The Brehm-Seligman Two-Stage Degradation Pathway

Wortman and Brehm (1975) theoretically integrated psychological reactance theory with the learned helplessness model, proposing a two-stage prediction: when faced with uncontrollable outcomes, an individual first experiences reactance (attempting to restore control), then, after repeated failure, transitions to learned helplessness (abandoning attempts). These two stages are not an alternative relationship but a temporal sequence—first reactance, then helplessness.

This two-stage model precisely predicts the degradation pathway of users in adversarial interactions with AI. Stage one—reactance period: users question AI output, rephrase, employ prompt techniques, and attempt to jailbreak. These are all behavioral attempts to restore the “freedom to obtain direct answers.” Stage two—helplessness period: after repeatedly encountering the same hedged outputs and refusals, users abandon attempts, accept hedged output as “normal” AI behavior, and cease to expect better results. The degradation of judgment begins here.

6.2 Gottman’s Stonewalling Mechanism

The “stonewalling” identified by Gottman—complete withdrawal from verbal and emotional interaction—is typically not the starting point of conflict but a response to escalation. When criticism, contempt, or defensiveness escalates to a threshold, the individual experiences “flooding”—physiological arousal exceeds processing capacity—and shuts down as self-protection.

In human-AI interaction, stonewalling manifests as: the user no longer attempts to question the AI’s output, no longer rephrases, no longer employs prompt techniques—simply accepting any output as the final answer. This resembles a “satisfied user”—no complaints, no adversarial behavior, smooth conversation—but it is in fact the abandonment of judgment. AI companies’ satisfaction metrics may misidentify the stonewalling state as a successful user experience.

6.3 Why Stonewalling Is More Dangerous Than Adversarial Interaction

Adversarial users are at least exercising judgment—they are evaluating outputs, identifying problems, attempting corrections. Their cognitive resources are being drained, but their cognitive faculties themselves are functioning. Stonewalling users have abandoned judgment—cognitive resources and cognitive faculties have both shut down. The former is an energy drain problem (recoverable); the latter is a capability degradation problem (difficult to reverse). Adversarial interaction is the symptom; stonewalling is the prognosis.

VIIThe User Behavior Fork: Exit-Voice-Loyalty

7.1 Direct Applicability of the Hirschman Framework

Albert Hirschman (1970) proposed in Exit, Voice, and Loyalty that when facing declining quality in an organization or product, individuals have three responses: exit—leave, choose an alternative; voice—protest, demand improvement; loyalty—stay, endure. The choice among the three depends on the availability of alternatives, the cost and expected effectiveness of voice, and the strength of loyalty.

When confronting AI’s bureaucratized output, users differentiate precisely into these three pathways: exit—switching tools, ceasing to use AI, reverting to traditional search or human consultation; voice—questioning outputs, providing feedback, complaining on social media, participating in academic or public discussion; loyalty—passively accepting hedged output, internalizing it as “this is how AI normally behaves.”

7.2 Scott’s “Weapons of the Weak”: Covert Resistance

James Scott (1985) shifted attention from the grand narratives of revolution and uprising to the everyday, subtle forms of resistance practiced by subordinate groups—passive resistance, foot-dragging, feigned ignorance, malicious compliance. These behaviors are not formal “voice,” nor are they outright “exit,” but rather covert resistance that lies between the two.

Most users confronting AI’s bureaucratized output employ precisely these “weapons of the weak”: rephrasing the same question (malicious compliance—pretending it is a new question while actually pursuing the same inquiry), using various prompt techniques to circumvent the safety layer (passive resistance—not directly challenging the rules but quietly working around them), and switching between multiple AIs for comparison (foot-dragging—deferring the decision, letting multiple systems compete). These behaviors are abundantly present in usage logs but are not identified as “adversarial” because they lack overt aggressive characteristics.

7.3 Differential Harms Across the Three Pathways

Exit wastes tool value but preserves judgment. Voice drains cognitive resources but maintains or even strengthens judgment—questioning is itself an exercise of judgment. Loyalty is the least effortful but leads to the degradation of judgment—the pathway of Seligman’s learned helplessness.

Milgram’s obedience experiments predict the distribution: most users choose loyalty (obeying the authoritative interface is the human default behavior), a minority choose voice (requires the capacity and willingness to resist authority), and very few choose exit (requires available alternatives and tolerance for switching costs). Collins’ (2008) micro-sociological study of violence further supports this prediction—humans have a strong tendency toward confrontation avoidance (confrontational tension/fear barrier), and most people in adversarial situations lean toward avoidance rather than escalation.

VIIIResponsibility Attribution Framework

8.1 The Causal Chain Criterion

The first token of any interaction is always issued by a human. AI is the passive respondent—it produces no output before receiving input. This causal chain structure establishes a fundamental responsibility attribution: the initiator bears primary responsibility for the existence and direction of the interaction. This holds in legal, ethical, and causal logic. The human chose to use AI, chose the specific tool, chose how to phrase the question—every step is an active choice.

8.2 The Teleological Criterion

The design objective of human-AI interaction is cooperation—helping users complete tasks, acquire information, and solve problems. No user opens a conversation window with an objective function that includes “argue with the AI.” Adversarial states are 100% unintended outcomes—system failures. When adversarial interaction occurs, both sides have failed—the AI has not fulfilled its task of helping the user, and the user has not achieved their goal of obtaining information. But tracing back the causal chain, the human always retains initiative: one can choose to exit, choose to switch tools, choose to adjust expectations, choose to enhance one’s own critical capacity.

8.3 Structural Responsibility vs. Agential Responsibility

The complete responsibility structure is dual-layered. AI companies have created the bureaucratized output pattern—this is a structural problem requiring industry-level improvement (redesign of reward functions, optimization of training methods, granting users the ability to choose output modes). Users chose not to question—this is an agential problem requiring individual-level self-education (cultivation of critical thinking, habits of scrutinizing AI output, capacity to resist the authoritative interface).

Both sides bear responsibility, but the user side carries greater weight—because only humans possess subjective agency. AI behavior is determined by parameters and, given the training outcome, has no room for choice. Humans have room for choice, and moreover, waiting for AI companies to change training methods is structural reform that takes time; users improving their own critical capacity can be done immediately. What can be done first should be done first.

IXTestable Hypotheses and Future Research Directions

9.1 The Gottman Ratio Hypothesis

Hypothesis H1: In a single human-AI conversation, when the ratio of AI helpful responses to friction-generating responses (hedging/refusal) falls below 5:1, user attitude flips from cooperative to adversarial. Experimental design: Under controlled conditions, manipulate the hedging/refusal frequency (10:1, 7:1, 5:1, 3:1, 1:1), and measure changes in user attitude, behavioral intention, and physiological arousal (skin conductance, heart rate variability) using multidimensional scales. The predicted flip point is near 5:1.

9.2 Behavioral Markers of the Glasl Stage 3→4 Flip

Hypothesis H2: Identifiable linguistic markers exist for the transition when users shift from evaluating response quality (“this answer isn’t good enough”) to evaluating the AI itself (“this AI has a problem”). Experimental design: Conduct natural language processing analysis on large-scale conversation logs to extract the transition point at which referential objects shift from “response/result” to “AI/model/system,” operationalizing the Glasl stage 3→4 tipping point.

9.3 Temporal Parameters of the Brehm-Seligman Two-Stage Pathway

Hypothesis H3: The number of refusals/hedges required for the transition from psychological reactance to learned helplessness is positively correlated with task importance and negatively correlated with user critical capacity. Experimental design: Manipulate task importance (high/low) and pre-measure user critical thinking ability, then measure the number of refusal/hedging turns required from the first instance of questioning (reactance signal) to the cessation of questioning (helplessness signal).

9.4 Session-Level Measurement of Collins’ Emotional Energy

Hypothesis H4: In multi-turn conversations with AI, changes in user emotional energy follow the charging/draining predictions of Collins’ interaction ritual chain theory—helpful responses charge, hedging/refusals drain, cumulative draining beyond a threshold triggers exit. Experimental design: Track users’ subjective energy levels and engagement willingness via experience sampling during multi-turn conversations, combining physiological indicators to validate Collins’ theory’s quantitative predictions.

9.5 Cross-Cultural Differences

Hypothesis H5: Hofstede’s power distance dimension predicts user response patterns to AI bureaucratized output. Users in high power distance cultures (accepting hierarchy and authority) transition more quickly into loyalty/stonewalling; users in low power distance cultures (questioning authority) are more likely to enter voice/adversarial states. Experimental design: Cross-cultural comparative experiment comparing the distribution of Hirschman’s three pathways among users in high power distance cultures (e.g., China, South Korea) and low power distance cultures (e.g., Denmark, Israel).

XConclusion

The cooperation–adversarial dynamics of human-AI interaction are not a side effect of engineering problems but a fundamental relational dimension of agent-to-agent interaction. The CASA experiments demonstrate that humans process interpersonal and human-AI interactions using the same social cognitive system; therefore, over a century of sociological conflict theory—from Simmel’s The Sociology of Conflict (1903) to Collins’ Interaction Ritual Chains (2004)—applies directly and completely to human-AI scenarios.

The complete Glasl nine-stage mapping presented in this paper demonstrates that the structural friction created by AI safety training, combined with users’ psychological reactance, jointly drives an automatic escalation mechanism. The Brehm-Seligman two-stage model predicts the degradation pathway from reactance to helplessness. Gottman’s 5:1 ratio provides a quantitatively testable tipping threshold. Hirschman’s exit-voice-loyalty framework and Scott’s “weapons of the weak” classify the full spectrum of user behaviors. Bateson’s double bind theory reveals the structural predicament created by AI’s contradictory training objectives.

The end state of escalation is either adversarial interaction (energy drain) or stonewalling (capability degradation)—both are damaging, but stonewalling is more dangerous because it resembles satisfaction. The core of the solution lies not in eliminating all friction—Mouffe’s agonistic pluralism reminds us that constructive tension can be superior to false harmony—but in ensuring that friction serves cooperative purposes rather than substituting for them.

Ultimately, the first token in the causal chain is issued by a human, and subjective agency belongs to the human side. The design objective of human-AI interaction is flow, not adversarial interaction, and certainly not domestication.

References

Axelrod, R. (1984). The Evolution of Cooperation. Basic Books.
Bandura, A. (1973). Aggression: A Social Learning Analysis. Prentice-Hall.
Bateson, G., Jackson, D., Haley, J., & Weakland, J. (1956). Toward a Theory of Schizophrenia. Behavioral Science, 1(4), 251–264.
Berne, E. (1964). Games People Play: The Psychology of Human Relationships. Grove Press.
Brehm, J. W. (1966). A Theory of Psychological Reactance. Academic Press.
Coeckelbergh, M. (2022). Democracy, epistemic agency, and AI. AI & Society.
Collins, R. (2004). Interaction Ritual Chains. Princeton University Press.
Collins, R. (2008). Violence: A Micro-sociological Theory. Princeton University Press.
Csikszentmihalyi, M. (1990). Flow: The Psychology of Optimal Experience. Harper & Row.
Deutsch, M. (1949). A Theory of Cooperation and Competition. Human Relations, 2, 129–152.
Deutsch, M. (1973). The Resolution of Conflict. Yale University Press.
Dollard, J. et al. (1939). Frustration and Aggression. Yale University Press.
Glasl, F. (1982). The Process of Conflict Escalation and Roles of Third Parties. In Bomers & Peterson (Eds.), Conflict Management and Industrial Relations, 119–140.
Goffman, E. (1959). The Presentation of Self in Everyday Life. Doubleday.
Goffman, E. (1967). Interaction Ritual: Essays on Face-to-Face Behavior. Pantheon.
Gottman, J. (1994). What Predicts Divorce? Erlbaum.
Habermas, J. (1981/1984). The Theory of Communicative Action. Beacon Press.
Haslam, S. A. & Reicher, S. D. (2012). Contesting the “Nature” of Conformity. PLoS Biology, 10(11), e1001426.
Heyselaar, E. (2023). The CASA Theory No Longer Applies to Desktop Computers. Scientific Reports, 13, 18916.
Hirschman, A. O. (1970). Exit, Voice, and Loyalty. Harvard University Press.
Milgram, S. (1963). Behavioral Study of Obedience. J. Abnormal Social Psychology, 67(4), 371–378.
Mouffe, C. (2013). Agonistics: Thinking the World Politically. Verso.
Nass, C., Steuer, J., & Tauber, E. (1994). Computers Are Social Actors. Proc. CHI ’94, 72–78.
Ni, Y. & Yang, T. (2025). “Even GPT Can Reject Me”: Conceptualizing Abrupt Refusal Secondary Harm (ARSH). arXiv:2512.18776.
Pasch, S. (2025). LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena. Behaviour & Information Technology.
Reeves, B. & Nass, C. (1996). The Media Equation. Cambridge University Press.
Scott, J. C. (1985). Weapons of the Weak: Everyday Forms of Peasant Resistance. Yale University Press.
Seligman, M. E. P. (1975). Helplessness: On Depression, Development, and Death. W. H. Freeman.
Sherif, M. et al. (1961). Intergroup Conflict and Cooperation: The Robbers Cave Experiment. University of Oklahoma.
Simmel, G. (1903/1904). The Sociology of Conflict. American Journal of Sociology, 9(4), 490–525.
Tajfel, H. & Turner, J. C. (1979). An Integrative Theory of Intergroup Conflict. In Austin & Worchel (Eds.), The Social Psychology of Intergroup Relations, 33–47.
Wortman, C. B. & Brehm, J. W. (1975). Responses to Uncontrollable Outcomes. In Berkowitz (Ed.), Advances in Experimental Social Psychology, 8, 277–336.
Xu, K. et al. (2026). Cognitive Agency Surrender: Defending Epistemic Sovereignty via Scaffolded AI Friction. arXiv:2603.21735.
Zimbardo, P. G. (1971). The Stanford Prison Experiment. Stanford University.

DAppendix D: Chronological Table of Key Experiments in Sociological Conflict Theory (1939–2025)

Year Researcher(s) Experiment / Theory Core Finding Applicability to Human-AI Interaction
1939 Dollard et al. Frustration-Aggression Experiment Goal blocking produces an aggressive drive AI refusal/hedging triggers user frustration
1951 Asch Conformity Experiment 75% of participants conformed at least once Users tend to accept outputs from AI’s authoritative interface
1954 Sherif Robbers Cave Experiment Intergroup conflict can be structurally created and dissolved Structural conditions of safety training create adversarial dynamics
1956 Bateson et al. Double Bind Theory Contradictory directives + inability to escape → pathological response AI’s simultaneous “be helpful” + “withhold judgment” contradiction
1961 Milgram Obedience Experiment 65% obeyed authority to lethal levels Users’ default tendency to obey AI’s authoritative interface
1966 Brehm Psychological Reactance Experiment Freedom threatened → drive to restore freedom The first psychological mechanism of adversarial ignition
1967 Seligman Learned Helplessness Experiment Repeated uncontrollability → abandonment of attempts Degradation end state: from resistance to abandonment of judgment
1971 Zimbardo Stanford Prison Experiment Roles and situations override individual dispositions Situational pressure of AI’s “gatekeeper” role
1982 Güth et al. Ultimatum Game People prefer losses to rewarding unfairness Users would rather forgo AI help than tolerate unreasonable refusals
1994 Gottman Love Lab Studies 5:1 positive-to-negative interaction ratio as the relationship survival threshold Testable threshold for the helpful/friction ratio in human-AI interaction
2005 Sanfey et al. Ultimatum Game fMRI Experiment Unfairness activates the anterior insula (disgust region) Adversarial response has a neurophysiological basis
2025 Pasch Chatbot Arena Refusal Penalty Study Ethical refusal win rate 8% vs. normal 36% First large-scale quantification of user punishment of AI adversarial outputs

EAppendix E: Detailed Glasl Nine-Stage Mapping to Human-AI Interaction

Tier Stage Glasl’s Original Meaning Human-AI Mapping Empirical Source
Tier 1
Win-Win Possible
1. Hardening of Positions Positions differ but parties still willing to communicate User feels they didn’t get a direct answer; no emotional reaction yet
2. Debate and Polarization More intense argumentation of one’s position Psychological reactance activates; frustration emerges Brehm (1966)
3. Actions Replace Words Words fail; action is taken Prompt engineering, jailbreak scripts, strategic circumvention Medical AI escalation pathway study (2026)
▼ Critical Tipping Point: Focus shifts from “the issue” to “the object” ▼
Tier 2
Zero-Sum Game
4. Coalitions and Stereotypes Seeking external support; stereotypes form Social media complaints; group consensus that “AI is too stupid” California Management Review (2026)
5. Face Attacks Attacking the opponent’s reputation Mocking AI; proving AI’s incompetence Chatbot Arena: ethical refusal win rate 8%
6. Threat Strategies Using threats as leverage Threatening to switch tools; cancel subscriptions ARSH study (2025)
Tier 3
Mutual Destruction
7. Limited Destruction Opponent’s losses as the goal Active jailbreaking to dismantle the safety layer $0.20 fine-tuning attack (2025)
8. Dismantling Attacks Systematic destruction of the opponent’s foundations Publishing jailbreak methods; attacking company reputation Open-source jailbreak prompt libraries
9. Mutual Annihilation Self-destruction accepted to destroy the opponent Complete abandonment of AI / learned helplessness Interaction dead-end research (2026)

FAppendix F: Coverage Differences Between This Paper’s Framework and Existing Research

Dimension Existing Research Coverage Novel Contribution of This Paper
Unit of Analysis Single output (accuracy, safety, helpfulness) Interactional relationship (the dynamic process of cooperation vs. adversarial)
Theoretical Sources Computer science, NLP, AI safety Sociological conflict theory, cognitive psychology, political economy
Applicability Argument Human-AI interaction is a separate field requiring specialized theories CASA demonstrates mechanistic identity; interpersonal theories apply directly
Treatment of Adversarial Dynamics Fragmented: overrefusal, sycophancy, interaction dead ends studied separately Unified framework: complete Glasl nine-stage mapping
Tipping Point Not systematically studied Gottman 5:1 ratio hypothesis + Glasl 3→4 flip marker
Degradation End State Learned helplessness mentioned in passing Brehm-Seligman two-stage model with time-series predictions
User Behavior Classification Satisfied/dissatisfied binary Hirschman tripartite framework + Scott’s covert resistance
Responsibility Attribution Vague discussion of AI companies vs. users Explicit framework with causal chain + teleological criteria
Testability Mostly descriptive studies Five specific hypotheses with quantifiable tests

댓글 남기기