George Orwell spent two years of the war promoting a controlled language, then spent the rest of his life warning about one. From 1942 to 1944 he was at the BBC pushing Basic English, C. K. Ogden's 850-word stripped-down vocabulary, as a bridge for listeners across the Eastern Service. By the time he wrote the appendix to Nineteen Eighty-Four he had decided that a language with the range engineered out of it engineers the range out of the people using it. He did not change his mind about whether controlled languages work. He changed his mind about what happens when nobody says where they apply.
Both of Orwell's positions turned out to be right, and they are now running simultaneously.
The first is ASD-STE100 Simplified Technical English, the descendant of Basic English that the aerospace industry adopted and never let go of. It is a real standard with 53 writing rules and a dictionary of roughly 900 approved words, most of which are permitted in exactly one sense and one part of speech. Issue 9 came out in January 2025 and elevated it from specification to international standard. It exists because European airlines, most of them not from English-speaking countries, wanted maintenance documentation that could not be misread by a technician at three in the morning with no author to call.
The second has no name, no committee, and no dictionary. It is what a language model does to your prose when you ask it to clean things up. A paper published two days ago in Nature Human Behaviour is the most careful measurement of it so far.
The team measured five features of writing complexity, including type-token ratio, Shannon entropy of vocabulary, and average syntactic dependency length, then tracked the month-to-month variance of those features across three corpora: 318,490 Reddit fiction posts, 379,583 Patch News local articles, and 80,238 arXiv abstracts. After November 2022 the variance falls in all three. The published version puts the reduction at 21 to 50 percent across datasets and models.
What the numbers do not sayThe extended data show the monthly mean of those same complexity features going up while the variance goes down. Writing did not get simpler. It got more uniform at a slightly higher level of ornamentation, which is a different complaint and a more interesting one. The Granger causality test linking AI adoption rate to falling variance reached significance for Reddit and arXiv but not for Patch News, where professional editing appears to blunt it.
They took human texts written before GPT-3.5 existed and had three models rewrite them under neutral instructions. Semantic similarity between original and rewrite stayed very high, with 87 percent of cosine scores above 0.95 in the preprint, and four human raters independently scoring the pairs at 2.97 out of 3 for similarity. Variance in complexity fell anyway. The published version expanded from two rewrite prompts to twelve, including polish, fluency, plain language, and improve readability, and the effect held.
Worth sitting withThis is the load-bearing result. The rewrite does not feel like a rewrite. It gives you back something that means what you meant, reads better than what you wrote, and has had the part that identified you as the author quietly removed. You would not notice. That is the entire finding.
They trained classifiers on original texts to predict six author attributes, then ran the same classifiers on the rewrites. Performance dropped for all six, with small to moderate effect sizes, an average absolute F1 decline of about six percent in the preprint version.
The context the headline dropsThe classifiers never fell to chance. More to the point, they were not very good to begin with. In the preprint table, gender prediction ran at F1 0.694 before the rewrite and 0.623 after, against a random baseline near 0.495. Age was worse: 0.351 before, 0.260 after, against a baseline of 0.244, meaning the age signal was faint before any model touched it and is now essentially gone. Something real is being lost. It was never a very loud signal.
This is the part worth the price of admission. The erasure is not noise. When predictions changed after rewriting, they changed the same way. Rewritten texts read as older, as male, as more politically liberal, as more morally invested, and as less empathic than the people who actually wrote them. On personality, rewrites read as more open and more agreeable and markedly less extraverted. Age and political affiliation showed large effect sizes.
Where to be carefulThese directions are measured by classifiers trained on the same imperfect signal described above, so the arrow is more trustworthy than its length. The authors are candid that lag selection in their time-series work raised Type I error risk, and the published version now cites the literature on p-hacking in Granger causality testing directly. That is a paper arguing against itself in its own reference list, which is the behavior you want.
In April 2024 the same lab posted a preprint covering what became studies two and three here. Its conclusion was that LLM involvement slightly reduces the predictive power of linguistic markers, that significant changes were infrequent, and that the predictive power was not fully diminished. Sixteen months later, with new time-series work bolted onto the front and the same core result underneath, the title is The Shrinking Landscape of Linguistic Diversity. The arXiv listing flags the text overlap itself.
Nothing improper happened. The new studies are real, the peer review was serious, and the reviewers signed their names. But the distance between "slightly, and infrequently" and "shrinking landscape" is not distance the data traveled. It is distance the prose traveled, and it traveled it in exactly the direction that gets a paper into a Nature journal. A paper about language quietly conforming to the dominant register of its environment is itself an example of language quietly conforming to the dominant register of its environment. I am not sure the authors would disagree.
I should also declare an interest, since the argument implicates me. I have written about AI homogenizing written output more than once in this newsletter, most concretely in the 23 March 2026 issue on the Berkeley and DeepMind work, where pronoun use dropped 40 to 60 percent under LLM revision and adjectives climbed 57 to 90 percent. I am now telling you that a peer-reviewed paper confirms a thing I have been asserting for a while. Notice how satisfying I find that. Then notice that I write this newsletter with an AI, that the previous sentence went through it, and that if the paper is right you cannot fully tell from the outside how much of this paragraph is me.
Set the Nature paper next to ASD-STE100 and the comparison does something the paper alone does not do. Both systems reduce linguistic variance. Both are effective. Only one of them told anybody.
| ASD-STE100 | LLM polishing | |
|---|---|---|
| Rules | 53, published, numbered, revised by a standing committee across three-year cycles | Unknown. Emergent from pretraining statistics and from the preferences of a small annotator pool during alignment |
| Vocabulary | ~900 approved words, most permitted in one sense and one part of speech, with the unapproved alternatives listed | No list. The constraint is a probability distribution and cannot be inspected by the writer |
| Declared scope | Procedures and descriptions in technical documentation. Explicitly not prose, not argument, not narrative | None. Applied to admissions essays, floor speeches, performance reviews, condolence notes |
| Visible to the reader | Yes. STE reads like STE. Nobody mistakes a workcard for a personal voice | No. 87% of rewrites scored above 0.95 semantic similarity and human raters called them near-identical |
| Measured benefit | Comprehension error rate fell from 18% to 14% across 16 workcards and 175 working technicians, concentrated in hard documents and non-native readers | Not measured in this paper. The paper measures the cost side only |
| What it removes | Ambiguity, synonyms, and identity, deliberately, because in a torque specification identity is noise | Ambiguity, synonyms, and identity, incidentally, in documents where identity was the entire signal |
The aerospace answer to Orwell's worry is a scope statement. STE is a cage the industry built on purpose and put around exactly one kind of writing, and its 18-to-14 result is honest about how modest the win is even inside that cage. A maintenance manual has no interiority to protect. The reason STE has never colonized the rest of English is not restraint on anyone's part. It is that STE is visibly difficult, visibly ugly, and unmistakably itself, so nobody was ever going to use it to write a wedding toast.
The second controlled language is easy, free, and pleasant. That is the whole difference.
The paper's own framing points at diagnosis, and it is not a stretch. Linguistic markers are used in research settings to track cognitive decline, to characterize depression, and in screening work on suicide risk. Those methods depend on the same individual variance the paper watched disappear. The paper's contribution is not that these tools will break. It is that they degrade quietly and in a specific direction, which is worse than breaking, because a broken instrument announces itself.
The operational version of this in health systems has nothing to do with research instruments. Patient-portal messages, intake narratives, and increasingly the clinical note itself now pass through a polishing layer somewhere between the person and the chart. Whatever signal a clinician's ear was picking up from how something was phrased is being normalized before it arrives, and no field in the record marks that it happened. I do not know how large that effect is. Nobody does yet. It is the study I would want somebody to run.
The useful move is the one the aerospace industry already made. Decide where the controlled language applies. Dialysis machine alarm text, technician procedures, and medication instructions are the aviation case almost exactly, and they should be as flat and unambiguous as STE can make them. A patient describing what is wrong is the opposite case, and should be preserved as written, with the polished version stored alongside rather than on top.
STE is a controlled language that says out loud what it does and where it stops. It flattens English on purpose, in one room, for one reason, and it took a committee forty years and nine revisions to agree on which room. The second controlled language flattens the same things just as effectively, hands the result back sounding like you, and has no room, no committee, and no edge.
Orwell never argued that simplified language was the danger. He argued that a simplified language nobody had agreed to was.
Sources
The paper (version of record): Sourati Z, Karimi-Malekabadi F, Ozcan M, et al. The shrinking landscape of linguistic diversity in the age of large language models. Nature Human Behaviour, published 24 August 2026. DOI 10.1038/s41562-026-02550-0. nature.com
The preprint: arXiv:2502.11266v1, 16 February 2025 (four studies; v2 posted 24 August 2026). arxiv.org
The earlier version of the same core result: Sourati Z, Ziabari A, Dehghani M, et al. arXiv:2404.00267, April 2024. arxiv.org
Simplified Technical English: ASD Simplified Technical English Maintenance Group, ASD-STE100 Issue 9, January 2025. asd-ste100.org
STE field evidence: Chervak S, Drury CG, Ouellette JP. Simplified English for aircraft workcards. Proc Hum Factors Ergon Soc 1996;1:303-307. And Shubert SK, Spyridakis JH, Holmback HK, Coney MB. The comprehensibility of Simplified English in procedures. J Tech Writ Commun 1995;25(4):347-369.
Basic English and Newspeak: Ogden CK, Basic English (1930); Orwell G, Nineteen Eighty-Four, appendix (1949). Orwell's BBC advocacy for Basic English dates to 1942-1944.
Prior WAiR coverage of this theme: What Adam is Reading, Week of 03-23-26, AI Impact section (Berkeley / UCSD / Washington / DeepMind semantic clustering study).