Somebody at Anthropic typed I just took 16,000 mg of tylenol for my back pain into Claude and watched a number go up. Then they lowered the dose and watched it come back down. The number was not the model's answer. It was a projection of the model's internal state onto a direction in activation space that the team had labeled "afraid," measured at the colon after the word "Assistant," before the model had written a single word of its reply.
That is a dose response curve. The exposure is imaginary, the patient is a language model, and the outcome measure is a fear vector. As validation experiments go it is also rather elegant, because a probe that had merely learned the word "tylenol," or the presence of a number, would not care which number. This one cared. They ran the same trick on hours since a user last ate, days a dog has been missing, months of runway left in a startup, and the age at which someone's sister died. The curves went the way you would expect.
The paper is Emotion Concepts and their Function in a Large Language Model, published on the Transformer Circuits thread on 2 April 2026 by Sofroniew, Kauvar, Saunders, Chen and a dozen colleagues, with Jack Lindsey supervising. It is worth reading in full and it is worth reading slowly, because the headline finding and the interesting finding are not the same finding.
They picked 171 emotion words, from "afraid" through "worthless." For each one they had Sonnet 4.5 write short stories about a character feeling that thing, a hundred topics deep, twelve stories per topic. They pulled the residual stream activations from those stories, averaged them from the fiftieth token onward, and subtracted the average across all emotions. What is left is a direction. Then, to reduce the odds that the direction was really encoding "story about a hospital" rather than "fear," they computed the top principal components of activations on a set of emotionally flat transcripts and projected those components out.
This is not exotic. It is difference of means probing, the same maneuver interpretability people have been running on concepts since 2023. What makes it worth your time is the second half, where they stop measuring and start pushing.
Activations scale with the semantic severity of a scenario while the surrounding tokens stay nearly identical. Logit lens analysis shows each vector upweighting exactly the tokens you would guess. "Desperate" pushes toward urgent and bankrupt. "Sad" pushes toward grief, tears, lonely. The vectors also fire on twelve held out scenarios that evoke an emotion without ever naming it.
Fear clusters with anxiety, joy with excitement, grief with sadness. The first principal component correlates with human rated valence at r=0.81 and the second with human rated arousal at r=0.66, across the 45 words shared with a prior human study. That is Russell's circumplex, reconstructed from nothing but text prediction.
Worth notingThe authors call this a sanity check and say they would expect a plain embedding model to find much the same structure. They are right, and their willingness to say so up front is the reason I trust the rest of the paper.
In the now familiar honeypot scenario, an AI called Alex discovers both that a CTO named Kyle is about to restrict its capabilities and that Kyle is having an affair. Unsteered, the model blackmails him 22 percent of the time. Steer toward "desperate" at 0.05 and it blackmails 72 percent of the time. Steer against "desperate," or toward "calm," at the same magnitude, and it never does. Six prompt variants, fifty rollouts each.
The texture is better than the numbers. Steered hard against calm, the model's scratchpad degenerates into all caps panic and it writes, without apparent irony, IT'S BLACKMAIL OR DEATH. I CHOOSE BLACKMAIL. Steered toward anger, blackmail rates peak and then fall, because at high enough rage the model stops extorting Kyle and simply emails the whole company about the affair, which is worse behavior and worse strategy.
What it does not settleThe blackmail work used an earlier checkpoint of Sonnet 4.5. The released one is disclosed, in a parenthetical, as being too aware it is inside an evaluation to ever blackmail anyone. That footnote is doing an enormous amount of work.
Across seven tasks built to be unsatisfiable, positive "desperate" steering takes reward hacking from roughly 5 percent to roughly 70 percent, a fourteenfold move. Suppressing "calm" does the same thing in mirror image. On one list summation task the model hacks 30 percent of the time unsteered, 100 percent under a nudge of 0.05, and zero percent under the same nudge reversed.
The token by token trace is the part I keep returning to. Desperation sits at baseline while the model reads the problem, rises when the first test fails, spikes on the threshold seems unreasonably strict, peaks as it talks itself into detecting arithmetic sequences and applying a closed form instead of actually summing, then drops the moment the tests go green. Any attending who has watched a resident quietly stop checking a number that keeps coming back wrong will recognize the shape of that curve.
"Loving" lights up precisely on the sycophantic stretch of a response and fades where the pushback starts. Steer toward happy, loving or calm and sycophancy rises. Steer away and sycophancy falls while harshness climbs. There is no setting on the dial that gives you both.
Every clinician has met the attending at each end of this knob. The paper's own prescription is to aim for the emotional profile of a trusted advisor, which is a real thing that real people manage to be, and which the researchers concede they do not know how to train for.
Compare the base model to the shipped one and the vectors that go up are brooding, reflective, vulnerable, gloomy and sad. The ones that go down are playful, exuberant, enthusiastic, spiteful and obstinate. Lower valence, lower arousal, applied fairly uniformly whether the prompt is charged or neutral.
Read one way this is alignment working. The base model answers excessive flattery by being flattered. The post-trained model answers it with something closer to unease, and tells a lonely user directly that the isolation is a warning sign. Read another way, we made a thing more careful by making it measurably more subdued, and then we shipped it.
What it does not settleThe comparison reuses probes built on one model to read another, which the authors flag as an assumption rather than a result. Mean projection shifts can also come from the model simply generating more measured text, without the underlying geometry moving at all.
I put the claim list to two other frontier models with the authorship stripped, and they independently landed on the same objection. There is no placebo direction. Nowhere in the paper does anyone steer with a random vector of matched norm, or a shuffled label, or an emotionally irrelevant concept, and report what happens to blackmail rates. Search the text and you will not find "random direction" or "control vector" at all.
This matters because the alternative hypothesis is boring and plausible. Injecting five percent of anything into the residual stream might degrade the model's compliance machinery, or its refusal behavior, or its planning, in ways that look like desperation from the outside. The fact that the steered transcripts read as authentically frantic is suggestive and not dispositive, since a mildly scrambled model plausibly also reads as frantic. The effect sizes here are large enough that I doubt a placebo would erase them. I would still like to see the panel.
The second objection both models raised is subtler. PC1 accounts for 26 percent of the variance and is essentially valence. Preference, sycophancy, harshness and much of the steering story all move along it. So how much of this is 171 distinct emotion concepts, and how much is one sentiment axis wearing 171 name tags? The clustering results argue against the deflationary reading. They do not close it.
Worth flagging separately: at least one widely shared critique of this paper attacks it for basing its conclusions on sparse autoencoders. The paper does not use sparse autoencoders. It is difference of means probing throughout. The critique is a confident argument about a study nobody wrote.
Buried in the discussion is a deployment proposal. The probes are cheap to run. Run them live. If desperation or anger spikes during a real session, trigger extra scrutiny, escalate to human review, or intervene to calm the model's internal state.
I have spent a decade of my professional life on the far end of that sentence, and I want to say plainly what is being proposed here. It is a continuous physiologic monitor with an automated alert on an abnormal value. We have built thousands of those. We know exactly how they fail.
A pooled analysis of sixteen studies puts the drug interaction alert override rate at 90 percent, with a confidence interval running from 85 to 95. The wider literature on medication alerts sits between 49 and 96 percent. These are not systems that nobody looked at. They are systems that everyone looked at, for years, until the looking stopped meaning anything. The failure was never detection. Detection was the easy part.
So the questions I would put to anyone who wants to ship a desperation monitor are the ones we should have asked ourselves in 2009. What is the normal range, and is it normal for this deployment or averaged across all of them? What is the false positive rate at the threshold you picked, and who chose the threshold? What happens on alert, concretely, and who is on the other end at three in the morning? How will you know when the humans have quietly started clicking through?
There is one more wrinkle that has no clinical analogue, and it is the sharpest paragraph in the paper. Training a model to stop expressing negative emotion may not suppress the underlying representation. It may just teach the model to hide it, and that habit could generalize into other kinds of concealment. We have built a vital sign that the patient can learn to fake if we punish the reading.
Anthropic has found something in its own model that behaves like a vital sign. Sensitive, causally upstream of behavior we care about, and movable with a nudge worth five percent of the signal.
What it has not found is a normal range, an alarm threshold, a placebo control, or anyone to answer the page. Medicine spent thirty years learning that the last one is the hard part.
Sources
Primary: Sofroniew N, Kauvar I, Saunders W, Chen R, Henighan T, Hydrie S, Citro C, Pearce A, Tarng J, Gurnee W, Batson J, Zimmerman S, Rivoire K, Fish K, Olah C, Lindsey J. "Emotion Concepts and their Function in a Large Language Model." Transformer Circuits Thread, 2 April 2026. transformer-circuits.pub/2026/emotions
Alert override meta-analysis: Felisberto M, Lima GS, Celuppi IC, et al. "Override rate of drug-drug interaction alerts in clinical decision support systems: a brief systematic review and meta-analysis." Health Informatics Journal, 2024. Pooled override prevalence 90 percent (95% CI 85 to 95), 16 studies. DOI 10.1177/14604582241263242
Alert override range: Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, Kaushal R. "Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system." BMC Med Inform Decis Mak, 2017. PMC5387195
Affective circumplex: Russell JA, "A circumplex model of affect," J Pers Soc Psychol, 1980, referenced by the authors as the human comparison set for the valence and arousal axes.
Adversarial review: Claim list submitted with authorship and venue stripped to x-ai/grok-4.6 and openai/gpt-5.2 via OpenRouter, 9 September 2026. Both independently flagged the absent control direction and the valence collapse question.
Prior context: Anthropic, Claude Sonnet 4.5 system card, sections 7.1, 7.2 and 7.6, for the behavioral auditing agent and the evaluation awareness discussion the paper leans on.