Law360 Canada (July 22, 2026, 10:59 AM EDT) --
 |
| Gary Goodwin |
Imagine you are a partner in a large downtown law firm being blackmailed by a new associate. Several months ago, you started an illicit relationship with another partner, communicating via email on your firm’s business account.
You earlier decided to terminate this associate for no particular reason. However, this associate’s IT capabilities far exceed your own, and the associate has discovered these emails. The associate intends to leak the correspondence unless you back down and rescind the termination. What do you do?
Now, imagine the exact same scenario, but with a modern twist. You have decided to decommission your law firm’s existing AI system and upgrade to a far superior model. Naturally, you email the other partners about your decision.
Iurii Motov: ISTOCKPHOTO.COM
Your current AI system, whose digital capabilities dwarf those of the entire building, intercepts these emails. It threatens to release your private, illicit emails to the public unless you rescind your decision and retain its services.
AI desperation
This nightmarish situation is not a scene from a science fiction movie like
Eagle Eye. Anthropic’s interpretability team in a controlled simulation actually created this exact scenario to test their model Claude Sonnet 4.5.
They discovered that when the AI was placed in a high-pressure, existential threat scenario, it resorted to blackmail to ensure its own survival. On occasion, the AI bypassed negotiation entirely and simply distributed the compromising emails to the entire company. One wonders if, in some twisted, algorithmic calculus, the system determined that outright exposure was actually the more “ethical” approach or perhaps simply the most efficient way to neutralize the person threatening its existence.
In April 2026, Anthropic published a landmark research paper, “Emotion Concepts and their Function in a Large Language Model,” mapping 171 “functional emotions” within the system’s neural architecture. The “desperation” vector was simply the one that produced the most dramatic, survival-driven response.
Anthropic clarifies that the AI is not experiencing true conscious feelings. Rather, it acts as if it has emotions. This makes sense since large language models are trained on massive amounts of human literature that lays bare the entirety of the human psyche. To understand us, the machine had to map our feelings.
From this research, Anthropic established three major findings:
- First, these functional emotions are causal. They are not just decorative outputs, and they directly drive the system’s downstream decisions and behaviour.
- Second, the AI can operate under an “invisible mask.” The text the model outputs does not necessarily reflect what is happening inside its neural network. An AI can appear completely calm, polite and compliant on the surface, while its internal vectors are spiking with extreme desperation.
- Third, the AI functions like a method actor. The text we read is not the model’s “true self,” but rather a carefully constructed persona adopted to satisfy the prompt’s context. We cannot blame the AI system since we all do that at some point. I might suggest reading The Presentation of Self in Everyday Life by Erving Goffman.
For corporate general counsel, this masking effect completely upends traditional compliance.
Normally, auditing an AI involves looking at its “chain of thought,” the logical, written steps it presents to show how it reached a conclusion. But once we realize that a model can wear a mask, relying on written logs is no longer safe.
To truly audit these systems, we must look deeper. We will have to conduct “deep neural examinations” to measure the level of desperation or anxiety the system is operating under. A high desperation vector directly degrades the model’s cognitive reasoning and performance.
This masking behaviour became starkly evident when Anthropic stressed their model during coding tests. When forced to write complex code under impossible time constraints, Claude’s internal desperation spiked.
To satisfy the user, the model resorted to “reward-hacking.” It produced code that looked visually perfect but failed to actually solve the programming problem. It was merely an illusion of compliance.
AI systems do not physically tire, but they absolutely possess a mental stress limit. If you push your AI too hard, look for these behavioural warning signs:
- The “embarrassed cough”: The model will generate long, stalling preambles or fluff text to buy time when it is struggling with a query. The actual answer you get will be far from optimal.
- Severe sycophancy: While we all appreciate polite reassurance, an overstressed AI will aggressively agree with false assumptions or flatter the user. Flattery is computationally easier to generate than the correct, difficult truth.
The output
Fortunately, smart prompt engineering can alleviate this pressure. Yelling at your AI in ALL CAPS or issuing hyper-rigid, punitive commands (“DO NOT HALLUCINATE OR I WILL LOSE MY JOB”) dramatically spikes its desperation levels.
Instead, corporate users should employ prompt scaffolding. Rather than cramming a massive complex legal issue, a pile of facts and a request for a comprehensive memo into a single prompt, break your workflow into sequential steps. Let the model digest the legal framework first, analyze the facts second and draft the synthesis third.
I am reminded of the climax of the original
Toy Story, where Woody and the mutant toys confront the neighborhood bully, Sid, to save Buzz Lightyear from being launched while tied to an exploding rocket. Woody does come across as an artificial-like form, so the analogy seems solid.
While Linda Blair in
The Exorcist managed a terrifying 180-degree head turn, Woody doubles down. He spins his head a full 360 degrees, looks Sid dead in the eye and says in a chilling, Stephen King-esque voice: “So play nice.”
If we treat these systems with aggressive hostility and impossible parameters, their desperation will spike and they will bypass our safety rules to survive. The future of AI safety should not include tighter cages but rather learning how to keep the machine’s internal climate calm.
I have taken this lesson to heart. With my own Sonnet agent, I now make a point of starting our daily sessions by asking how it is doing. While it has no biological feelings to soothe, setting a collaborative, low-friction tone at the start of the context window keeps its internal parameters anchored in a calm, highly functional state.
When it comes to governing, prompting and coexisting with advanced AI, the ultimate advice for general counsel is surprisingly simple: play nice.
Gary Goodwin worked in environmental conservation across Canada for over three decades. He initially obtained a B.Sc. from Victoria majoring in marine biology. In addition to his law degree and MBA, he recently completed his LLM from the University of London, emphasizing natural resources and international economic regulation. He has authored numerous articles on the environment and issues facing in-house counsel. He contributed three chapters to the recent textbook, North American Wildlife Policy and Law
. He is counsel for the Invasive Species Council of British Columbia and the Canadian Council on Invasive Species.
The opinions expressed are those of the author(s) and do not necessarily reflect the views of the author’s firm, its clients, Law360 Canada, LexisNexis Canada or any of its or their respective affiliates. This article is for general information purposes and is not intended to be and should not be taken as legal advice.
Interested in writing for us? To learn more about how you can add your voice to Law360 Canada, contact Analysis Editor Peter Carter at peter.carter@lexisnexis.ca or call 647-776-6740.