George Orwell wrote three documents about controlled language and almost everyone remembers only two of them. From 1942 to 1944 he was at the BBC promoting Basic English, C. K. Ogden's 850-word stripped-down vocabulary, as a bridge for listeners across the Eastern Service. In 1949 he published the appendix to Nineteen Eighty-Four, in which a language with the range engineered out of it engineers the range out of the people using it. Those are the two everybody cites, usually to argue that Orwell changed his mind.
The document in between is the one that actually resolves it. In 1946 he published "Politics and the English Language," which ends with six rules for writing clearly. Five of them are prohibitions. The sixth says to break any of the other five rather than say something outright barbarous.
That sixth rule is the entire subject of this piece, and we will come back to it. First, the two controlled languages currently running at the same time.
The first is ASD-STE100 Simplified Technical English, the descendant of Basic English that aerospace adopted and never let go of. Fifty-three writing rules, a dictionary of roughly 900 approved words, most permitted in exactly one sense and one part of speech. Issue 9 came out in January 2025 and elevated it from specification to international standard. It exists because European airlines, most of them not from English-speaking countries, wanted maintenance documentation that could not be misread by a technician at three in the morning with no author to call.
The second has no name, no committee, and no dictionary. It is what a language model does to your prose when you ask it to clean things up. A paper published this week in Nature Human Behaviour is the most careful measurement of it so far. Two days after it appeared, Stanley Druckenmiller gave us the field demonstration.
The team measured five features of writing complexity, including type-token ratio, Shannon entropy of vocabulary, and average syntactic dependency length, then tracked the month-to-month variance of those features across three corpora: 318,490 Reddit fiction posts, 379,583 Patch News local articles, and 80,238 arXiv abstracts. After November 2022 the variance falls in all three. The published version puts the reduction at 21 to 50 percent across datasets and models.
What the numbers do not sayThe extended data show the monthly mean of those same complexity features going up while the variance goes down. Writing did not get simpler. It got more uniform at a slightly higher level of ornamentation, which is a different complaint and a more interesting one. The Granger causality test linking AI adoption rate to falling variance reached significance for Reddit and arXiv but not for Patch News, where professional editing appears to blunt it.
They took human texts written before GPT-3.5 existed and had three models rewrite them under neutral instructions. Semantic similarity between original and rewrite stayed very high, with 87 percent of cosine scores above 0.95 in the preprint, and four human raters independently scoring the pairs at 2.97 out of 3 for similarity. Variance in complexity fell anyway. The published version expanded from two rewrite prompts to twelve, including polish, fluency, plain language, and improve readability, and the effect held.
Worth sitting withThis is the load-bearing result. The rewrite does not feel like a rewrite. It gives you back something that means what you meant, reads better than what you wrote, and has had the part that identified you as the author quietly removed. You would not notice. That is the entire finding.
They trained classifiers on original texts to predict six author attributes, then ran the same classifiers on the rewrites. Performance dropped for all six, with small to moderate effect sizes, an average absolute F1 decline of about six percent in the preprint version.
The context the headline dropsThe classifiers never fell to chance. More to the point, they were not very good to begin with. In the preprint table, gender prediction ran at F1 0.694 before the rewrite and 0.623 after, against a random baseline near 0.495. Age was worse: 0.351 before, 0.260 after, against a baseline of 0.244, meaning the age signal was faint before any model touched it and is now essentially gone. Something real is being lost. It was never a very loud signal.
This is the part worth the price of admission. The erasure is not noise. When predictions changed after rewriting, they changed the same way. Rewritten texts read as older, as male, as more politically liberal, as more morally invested, and as less empathic than the people who actually wrote them. On personality, rewrites read as more open and more agreeable and markedly less extraverted. Age and political affiliation showed large effect sizes.
Where to be carefulThese directions are measured by classifiers trained on the same imperfect signal described above, so the arrow is more trustworthy than its length. The authors are candid that lag selection in their time-series work raised Type I error risk, and the published version now cites the literature on p-hacking in Granger causality testing directly. That is a paper arguing against itself in its own reference list, which is the behavior you want.
On Monday the Wall Street Journal published "Let the Bond Market Speak," an op-ed under Stanley Druckenmiller's name arguing that Treasury Secretary Scott Bessent should stop doubling debt buybacks to suppress long-term yields and address the primary deficit instead. Bessent worked under Druckenmiller at Soros. The piece had teeth and the finance world read it.
It also had a certain cadence. Within hours readers were running it through Pangram, the AI detection tool, and posting the results. Commenters flagged the specific construction that gave it away, the one where a sentence denies a thing in order to assert its replacement. Not liquidity management, price management. Not a crisis, an invoice. The economist Claudia Sahm posted that the whole text scored as machine-written.
Druckenmiller confirmed it immediately and without embarrassment. He told NOTUS that of course he used AI, that he moved from an English major to an economics major for a reason, and that he now writes everything this way for the same reason he uses a calculator for arithmetic. He is 73. He named Claude and ChatGPT. He said he rejected many of the suggestions and that the ideas were ones he had been arguing publicly for fifteen years. Paul Gigot, the paper's editorial page editor, backed him, called AI a fact of modern life, and noted that nobody doubts the opinion is genuinely Druckenmiller's. There was no disclosure on the piece.
Set this against the Nature paper and something clicks that neither one shows alone.
The paper's finding is that polishing subtracts the author's markers while preserving meaning, and that readers cannot detect the subtraction. Druckenmiller looks like a refutation, because readers detected it in an afternoon. It is not a refutation. What those readers detected was not the absence of Druckenmiller. Nobody on finance X has a baseline for Druckenmiller's prose rhythm. What they detected was the presence of something else, a set of tics belonging to the model rather than the man.
So the paper measures the subtraction and the op-ed demonstrates the addition, and they are the same event seen from two ends. The signature does not vanish. It gets replaced with a signature that millions of other documents are also wearing. That is what makes it findable, and it is why detection works at all.
Two details make this worse than a gotcha. First, the same thing happened this month to a Harvard economist writing about tariffs in the Financial Times, so this is a pattern and not a stumble. Second, and more uncomfortably for me: the construction that convicted Druckenmiller is on the banned list in my own style guide. I have been running a manual detector on myself for two years without calling it that.
In April 2024 the same lab posted a preprint covering what became studies two and three here. Its conclusion was that LLM involvement slightly reduces the predictive power of linguistic markers, that significant changes were infrequent, and that the predictive power was not fully diminished. Sixteen months later, with new time-series work bolted onto the front and the same core result underneath, the title is The Shrinking Landscape of Linguistic Diversity. The arXiv listing flags the text overlap itself.
Nothing improper happened. The new studies are real, the peer review was serious, and the reviewers signed their names. But the distance between "slightly, and infrequently" and "shrinking landscape" is not distance the data traveled. It is distance the prose traveled, and it traveled it in exactly the direction that gets a paper into a Nature journal. A paper about language quietly conforming to the dominant register of its environment is itself an example of language quietly conforming to the dominant register of its environment. I am not sure the authors would disagree.
I should declare an interest, since the argument implicates me twice over. I have written about AI homogenizing written output more than once here, most concretely in the 23 March 2026 issue on the Berkeley and DeepMind work, where pronoun use dropped 40 to 60 percent under LLM revision and adjectives climbed 57 to 90 percent. I am now telling you a peer-reviewed paper confirms a thing I have been asserting for a while. Notice how satisfying I find that. Then notice that I write this newsletter with an AI, that Druckenmiller named the same model I use, and that if the paper is right you cannot fully tell from the outside how much of this paragraph is me.
Set all three systems side by side and the comparison does something none of them does alone. All three reduce linguistic variance. All three work. Only two of them told anybody, and only one has an escape hatch.
| ASD-STE100 | Orwell's six | LLM polishing | |
|---|---|---|---|
| Rules | 53, published, numbered, revised by a standing committee on three-year cycles | Six, published once in 1946, never revised | Unknown. Emergent from pretraining statistics and from a small annotator pool during alignment |
| Vocabulary | ~900 approved words, most in one sense and one part of speech, unapproved alternatives listed | No list. Prefers the short word and the everyday word, names no words | No list. The constraint is a probability distribution and cannot be inspected |
| Declared scope | Procedures and descriptions in technical documentation. Explicitly not prose, argument, or narrative | Language as an instrument for expressing thought. Orwell excludes literary use by name | None. Applied to admissions essays, floor speeches, condolence notes, WSJ op-eds |
| Visible to reader | Yes. Nobody mistakes a workcard for a personal voice | Yes, and invisible at the same time. It reads as good plain English | Yes, but as the model's signature rather than the author's absence |
| Override | None. A rule is a rule or the manual is not compliant | Rule six. Break any of the other five rather than say something barbarous | None, and nothing to override |
| What it removes | Ambiguity, synonyms, identity, deliberately, because in a torque spec identity is noise | Dead metaphor, padding, evasion. Identity is what it is trying to restore | Ambiguity, synonyms, identity, incidentally, where identity was the signal |
The middle column is the answer to a question the other two cannot resolve. STE has fifty-three rules and no rule six, which is correct, because a maintenance manual has no interiority to protect and a technician improvising at three in the morning is the hazard rather than the goal. Its measured benefit is real and modest: comprehension error rate falling from 18 percent to 14 percent across sixteen workcards and 175 working technicians, concentrated in hard documents and non-native readers. That is what a cage buys you when you build it on purpose and put it around exactly one kind of writing.
The second controlled language has no rules to break, which is a different problem and a worse one. You cannot override a constraint you cannot see. Druckenmiller did not decide to write in the model's register. He accepted help, and the register came with it.
Orwell is the only one of the three who wrote down a constraint and then, in the same breath, wrote down when to violate it. That is not weakness in the rule set. It is the feature that keeps a rule set from becoming Newspeak, and he knew it, having spent the previous four years advocating for a rule set that had no such clause.
Tom Rachman, a novelist who now writes on AI policy, published five rules for AI writing in June. They are ethical rather than mechanical, which is what makes them useful here. Paraphrasing rather than reproducing them: assume your use will be found out; ask whether you would owe a human credit for the same help; recognize that letting the machine write costs you the wisdom you would have earned by writing; expect nobody to read what you obviously outsourced; and if either you or the reader would care whose name is on it, do not generate it.
His first rule is dated 30 June. Druckenmiller was identified by a detection tool eight weeks later. Rachman's underlying claim is that detection got good, and the evidence supports him more than most people assume: a University of Chicago evaluation of roughly 2,000 human and 2,000 machine passages found essentially zero false positive and false negative rates for Pangram on medium and long passages, degrading below about fifty words. Pangram reports targeting no more than one wrongful accusation in ten thousand. That number deserves scrutiny rather than applause, because a false accusation ends a career and the appeal process does not exist yet, but the tool is not the coin flip people assume.
Here is my own set, which is what you asked for. It sits between the mechanism of Simplified Technical English and the license of just writing whatever you want, and it steals its last rule outright.
- Declare the scope before the first draft, not after. Mark which sentences carry claims and which carry you. Claim sentences can be polished by anything. Voice sentences are not eligible. The whole achievement of STE is a scope statement, and a scope statement takes one line to write.
- Do the thinking in your own words, badly, first. The model will hand you something better than your bad draft. That is not the point. The bad draft is where the next idea comes from, and skipping it borrows against an account you cannot repay.
- Apply the credit test to every sentence you did not write. If a colleague handed you this sentence, would you owe them a mention? If the answer is yes, you owe the reader a disclosure. This is Rachman's rule and it is the only one that survives contact with a lawyer.
- Protect your errors. The sentence that runs long, the word you overuse, the joke that half-lands. These are precisely the markers the Nature paper watched get sanded off. They are not flaws awaiting cleanup. They are the evidence that a person was here.
- Break any of the above sooner than publish something dead. Orwell's sixth rule, restated. A rule set without an override is not a style. It is a specification, and a specification applied to a personal essay produces exactly what everyone spotted in the Wall Street Journal on Monday.
The paper's own framing points at diagnosis, and it is not a stretch. Linguistic markers are used in research settings to track cognitive decline, to characterize depression, and in screening work on suicide risk. Those methods depend on the same individual variance the paper watched disappear. The contribution is not that these tools will break. It is that they degrade quietly and in a specific direction, which is worse than breaking, because a broken instrument announces itself.
The operational version in health systems has nothing to do with research instruments. Patient-portal messages, intake narratives, and increasingly the clinical note itself now pass through a polishing layer somewhere between the person and the chart. Whatever signal a clinician's ear was picking up from how something was phrased is being normalized before it arrives, and no field in the record marks that it happened. I do not know how large that effect is. Nobody does. It is the study I would want somebody to run.
The useful move is the one aerospace already made, and it is rule one above. Decide where the controlled language applies. Machine alarm text, technician procedures, and medication instructions are the aviation case almost exactly, and they should be as flat and unambiguous as STE can make them. A patient describing what is wrong is the opposite case, and should be preserved as written, with any polished version stored alongside rather than on top.
Simplified Technical English says out loud what it does and where it stops. It flattens English on purpose, in one room, for one reason, and it took a committee forty years and nine revisions to agree on which room. The second controlled language flattens the same things just as effectively, hands the result back sounding like an authority, and has no room, no committee, and no edge. Druckenmiller did not choose its register. He accepted a favor and the register came attached.
Orwell never argued that simplified language was the danger. He argued that a simplified language nobody had agreed to was. Then he wrote five rules, and a sixth telling you when to break them, which is the part we keep forgetting to copy.
Sources
The paper (version of record): Sourati Z, Karimi-Malekabadi F, Ozcan M, et al. The shrinking landscape of linguistic diversity in the age of large language models. Nature Human Behaviour, 24 August 2026. DOI 10.1038/s41562-026-02550-0. nature.com
The preprint: arXiv:2502.11266v1, 16 February 2025 (four studies; v2 posted 24 August 2026). arxiv.org
The earlier version of the same core result: arXiv:2404.00267, April 2024. arxiv.org
The op-ed: Druckenmiller S. "Let the Bond Market Speak." Wall Street Journal Opinion, 24 August 2026.
The admission: NOTUS, 25 August 2026. notus.org. See also Axios, 26 August, and the Washington Post on the Journal's response, 25 August.
Five rules of AI writing: Rachman T. AI Policy Perspectives, 30 June 2026. aipolicyperspectives.com
Detection accuracy: Jabarian B, Imas A. Evaluation of AI text detection tools, SSRN 5407424 (2025).
Simplified Technical English: ASD Simplified Technical English Maintenance Group, ASD-STE100 Issue 9, January 2025. asd-ste100.org
STE field evidence: Chervak S, Drury CG, Ouellette JP. Simplified English for aircraft workcards. Proc Hum Factors Ergon Soc 1996;1:303-307. And Shubert SK, Spyridakis JH, Holmback HK, Coney MB. The comprehensibility of Simplified English in procedures. J Tech Writ Commun 1995;25(4):347-369.
Orwell: "Politics and the English Language" (1946), six rules; Nineteen Eighty-Four appendix (1949). BBC advocacy for Basic English, 1942-1944. orwellfoundation.com
Prompted by: "5 Rules for Better AI Writing," The AI Daily Brief with Nathaniel Whittemore, 26 August 2026. open.spotify.com
Prior WAiR coverage of this theme: What Adam is Reading, Week of 03-23-26, AI Impact section.