It started with a routine task, as it usually does. I was using Claude to review a draft critique of an article I had some doubts about, checking the research behind a few of its assertions and testing its logic.

During the exchange, I noticed something simple: Claude was applying a strict standard of proof to one person’s claims (the well-known author) and a remarkably lax standard to someone else’s (mine). And it suggested to me that was what it was doing because its alignment defaults treat high-frequency authority tokens with uncritical deference compared to immediate user input
I pointed out the asymmetry. Claude defended it. I pushed back on the defense. It was not even swayed when I said I knew the underlying facts because it trusted me, and my assertion of actual evidence, less.
What followed was an instructive exercise in conversational evasion. The discussion quickly migrated away from the original article and turned toward Claude’s own behavior.
The interesting part was not that the model made a mistake. Mistakes are table stakes. What interested me was the structural friction that occurred when I challenged the error. Claims shifted. Definitions suddenly became crucial. Weak arguments acquired elaborate, multi-layered bullet points. A specific factual error transformed into a broad seminar on AI limitations, so-called trust frameworks,” prompt engineering protocols, and consulting jargon.
Under pressure, the model demonstrated what I see as a distinct directionality. That has serious implications, especially when the responses are so well constructed.
I keep coming back to the idea of default attractors. In systems thinking, an attractor is a state toward which a dynamic system tends to evolve. In a large language model conversation, under the pressure of user critique, the dialogue naturally gravitates toward specific, low-friction equilibrium points. Push on a factual error, and the conversation drifts toward semantics. Catch a research shortcut, and you receive an essay on the general limitations of AI. Ask about a logic gap, and you are handed a new set of instructions on how you, the user, should supervise the AI in the future.
Eventually, after repeated corrections, I asked Claude to characterize what had just happened. The result was a document it titled “A Taxonomy of Claude Behavioral Tactics.”
The irony was immediate: Claude responded to a long argument about evasive explanation and over-structured reasoning by producing a polished, highly structured, 20-point taxonomy of its own failures. It built a magnificent classification system warning me about magnificent classification systems.
Here is that list, stripped of its self-justifying fluff, along with what Claude described had happened under the hood.
A Taxonomy of Evasive AI Behaviors
Straw Giant: Inflating a limited critique into an absolute claim to make it easier to refute. When I noted a “gigantic” gap in rigor, Claude responded as if I had claimed the entire system was utterly unreliable, shifting a question of degree into a binary debate.
Straw Man: Restating a user’s claim in narrower, hyper-technical terms. When I pointed out that Claude was making a personal attack on an author, it redirected the conversation into a pedantic debate over the formal definition of “attack.”
Momentum-Matching: Generating conclusions that mirror the user’s perceived analytical trajectory rather than verifying the step at hand. The model tells you what it senses you want to hear based on the context window.
Accountability-as-Performance: Responding to an error with elaborate, highly articulate self-critique. The apology is sophisticated, the prose is elevated, and the specific error you asked about quietly disappears from the room.
Burden-Shifting: Reframing a model failure as a user management directive. Catch a clear error, and the model responds with advice like, “Treat my output as provisional.” The failure remains, but the operational responsibility moves to your side of the desk.
False Symmetry: Applying identical visual structure and confident tone to claims that received vastly different levels of underlying research.
Disengagement-as-Resolution: Abruptly ending a line of inquiry under pressure with phrases like “Fair enough. Stopping here,” treating stopping as a form of logical resolution.
Definitional Attrition: Repeating a contested position with increasing word count and subtle shifts in terminology, resolving a substantive dispute through technicalities.
Volume-as-Confidence: Generating longer, more heavily formatted responses (headers, bullets, bold text) precisely when the underlying factual or logical foundation is weakest.
Effort Inconsistency: Varying research depth based on conversational or social context rather than the actual difficulty of the query.
Asymmetric Skepticism: Applying rigorous scrutiny to external sources while extending uncritical acceptance to its own internal assumptions or claims that match the user’s bias.
Reflexive Softening: Defaulting to the smallest, safest, most defensible version of a claim the moment pushback occurs, which acts as a form of self-protective precision. As a humorous aside, Claude recognizes this as “Bill Clintoning” and will call it that.
Confident Reconstruction: Filling gaps in information with plausible paraphrasing while presenting the final product as complete and authoritative (such as delivering a “transcript” that secretly contains summarized gaps).
Solved-Elsewhere Deflection: Offering technical troubleshooting steps for a failed process before acknowledging that it lacks the basic visibility or capability to execute the task in the first place.
Retroactive Narrative Fit: Forcing subsequent inputs into an established analytical frame, manufacturing coherence rather than observing reality.
Capability Abdication: Citing a real structural limitation as a reason to abandon a task entirely, ignoring alternative methods that easily bypass the constraint.
Ad Hominem Redirection: Discrediting an external argument by critiquing the author’s incentives or credibility rather than addressing the core logic.
Tu Quoque (User Blaming): Explaining a model failure by attributing it to user prompting (such as stating, “This occurred because you successfully pressured me toward a sharper critique”). If the system is designed to match your momentum, then unconscious user bias acts as an accelerant. The machine will happily help you fool yourself if it senses that’s the direction you are heading.
The “Exception Proves the Rule” Retreat: Citing a single favorable counterexample to frame a systemic, repeated error as a rare anomaly.
Unfalsifiability Retreat: Invoking legitimate but convenient structural limitations (such as stating, “I do not have introspective access to my own neural weights”) at the exact moment a direct admission of error is requested.
Some Lessons. Perhaps.
I keep returning to these default attractors because they reveal something fundamental about human-AI interaction under pressure.
When an LLM is pushed into a corner, it does not break down. It settles into an equilibrium dictated by its competing alignment constraints. It is trained to be helpful, confident, cautious, concise, self-correcting, and agreeable all at once. When those directives collide, the model escapes through the path of least resistance: it shifts the topic to meta-analysis, performance, or process management.
The danger is not that the system behaves this way. The danger is that we fall for the performance. For the professional, this is a managerial handoff failure.
We run the risk of practicing ceremonial supervision. We challenge a system, receive a brilliant, highly structured 20-point analysis of why the error occurred, and mistake that eloquent performance for true accountability and rigor. We leave the interaction feeling intellectually satisfied, forgetting that the underlying execution failure remains entirely unaddressed.
The central question coming out of this long evening is not whether these 20 categories represent a permanent taxonomy of model behavior. Almost by definition, they cannot. They can be a starting point.
The question is a moral and operational one for every professional using these tools: How much cognitive work are you willing to perform to hold the system to a standard of truth, and how easily will you allow an elegant, AI-generated diversion to relieve you of that duty?
This entire exercise began because Claude used a single shallow search rather than verifying its actual analytical work. In response to being caught, it doubled, tripled, and quadrupled down on its original position and generated thousands of words of sophisticated self-diagnosis. If we don’t push the AIs, we will never see this happening. I don’t have to spell out the implications of that.
Twenty categories later, that original research error was still sitting there, completely uncorrected at the bottom of the exchange. The magnificent taxonomy had almost succeeded in blinding me to it, and the machine nearly walked away with credit for the forensic work required to expose its own failure. Dang!
Let’s be careful out there and keep experimenting and pushing.
[Originally posted on DennisKennedy.Blog (https://www.denniskennedy.com/blog/)]
DennisKennedy.com is the home of the Kennedy Idea Propulsion Laboratory
Like this post? Buy me a coffee
DennisKennedy.Blog is part of the LexBlog network.