What Adam Is Reading
Two Controlled Languages
The aerospace industry spent forty years writing down the rules for stripping identity out of English. A new Nature Human Behaviour paper says we built a second one by accident and forgot to say where it stops.
Single-paper review with comparison · 8 sources · August 2026

George Orwell spent two years of the war promoting a controlled language, then spent the rest of his life warning about one. From 1942 to 1944 he was at the BBC pushing Basic English, C. K. Ogden's 850-word stripped-down vocabulary, as a bridge for listeners across the Eastern Service. By the time he wrote the appendix to Nineteen Eighty-Four he had decided that a language with the range engineered out of it engineers the range out of the people using it. He did not change his mind about whether controlled languages work. He changed his mind about what happens when nobody says where they apply.

Both of Orwell's positions turned out to be right, and they are now running simultaneously.

The first is ASD-STE100 Simplified Technical English, the descendant of Basic English that the aerospace industry adopted and never let go of. It is a real standard with 53 writing rules and a dictionary of roughly 900 approved words, most of which are permitted in exactly one sense and one part of speech. Issue 9 came out in January 2025 and elevated it from specification to international standard. It exists because European airlines, most of them not from English-speaking countries, wanted maintenance documentation that could not be misread by a technician at three in the morning with no author to call.

The second has no name, no committee, and no dictionary. It is what a language model does to your prose when you ask it to clean things up. A paper published two days ago in Nature Human Behaviour is the most careful measurement of it so far.

A note on which paper this is. The arXiv preprint (2502.11266, February 2025) describes four studies. The version of record, published 24 August 2026, describes three, spanning seven datasets and more than 880,000 texts. It was received in February 2025 and accepted in July 2026, so seventeen months of peer review sit between the two documents. Where a number below comes only from the preprint, this piece says so. The distinction matters more than usual here, for reasons that become the subject of the piece.

What the paper found
1
Written English got less varied after ChatGPT
What actually happened

The team measured five features of writing complexity, including type-token ratio, Shannon entropy of vocabulary, and average syntactic dependency length, then tracked the month-to-month variance of those features across three corpora: 318,490 Reddit fiction posts, 379,583 Patch News local articles, and 80,238 arXiv abstracts. After November 2022 the variance falls in all three. The published version puts the reduction at 21 to 50 percent across datasets and models.

What the numbers do not say

The extended data show the monthly mean of those same complexity features going up while the variance goes down. Writing did not get simpler. It got more uniform at a slightly higher level of ornamentation, which is a different complaint and a more interesting one. The Granger causality test linking AI adoption rate to falling variance reached significance for Reddit and arXiv but not for Patch News, where professional editing appears to blunt it.

Solid
2
The meaning survives the rewrite. The fingerprint does not.
What actually happened

They took human texts written before GPT-3.5 existed and had three models rewrite them under neutral instructions. Semantic similarity between original and rewrite stayed very high, with 87 percent of cosine scores above 0.95 in the preprint, and four human raters independently scoring the pairs at 2.97 out of 3 for similarity. Variance in complexity fell anyway. The published version expanded from two rewrite prompts to twelve, including polish, fluency, plain language, and improve readability, and the effect held.

Worth sitting with

This is the load-bearing result. The rewrite does not feel like a rewrite. It gives you back something that means what you meant, reads better than what you wrote, and has had the part that identified you as the author quietly removed. You would not notice. That is the entire finding.

Solid
3
Classifiers get worse at guessing who wrote it
What actually happened

They trained classifiers on original texts to predict six author attributes, then ran the same classifiers on the rewrites. Performance dropped for all six, with small to moderate effect sizes, an average absolute F1 decline of about six percent in the preprint version.

The context the headline drops

The classifiers never fell to chance. More to the point, they were not very good to begin with. In the preprint table, gender prediction ran at F1 0.694 before the rewrite and 0.623 after, against a random baseline near 0.495. Age was worse: 0.351 before, 0.260 after, against a baseline of 0.244, meaning the age signal was faint before any model touched it and is now essentially gone. Something real is being lost. It was never a very loud signal.

Mostly Solid
4
The flattening has a direction
What actually happened

This is the part worth the price of admission. The erasure is not noise. When predictions changed after rewriting, they changed the same way. Rewritten texts read as older, as male, as more politically liberal, as more morally invested, and as less empathic than the people who actually wrote them. On personality, rewrites read as more open and more agreeable and markedly less extraverted. Age and political affiliation showed large effect sizes.

Where to be careful

These directions are measured by classifiers trained on the same imperfect signal described above, so the arrow is more trustworthy than its length. The authors are candid that lag selection in their time-series work raised Type I error risk, and the published version now cites the literature on p-hacking in Granger causality testing directly. That is a paper arguing against itself in its own reference list, which is the behavior you want.

Mostly Solid

The part that is not in the paper

In April 2024 the same lab posted a preprint covering what became studies two and three here. Its conclusion was that LLM involvement slightly reduces the predictive power of linguistic markers, that significant changes were infrequent, and that the predictive power was not fully diminished. Sixteen months later, with new time-series work bolted onto the front and the same core result underneath, the title is The Shrinking Landscape of Linguistic Diversity. The arXiv listing flags the text overlap itself.

Nothing improper happened. The new studies are real, the peer review was serious, and the reviewers signed their names. But the distance between "slightly, and infrequently" and "shrinking landscape" is not distance the data traveled. It is distance the prose traveled, and it traveled it in exactly the direction that gets a paper into a Nature journal. A paper about language quietly conforming to the dominant register of its environment is itself an example of language quietly conforming to the dominant register of its environment. I am not sure the authors would disagree.

I should also declare an interest, since the argument implicates me. I have written about AI homogenizing written output more than once in this newsletter, most concretely in the 23 March 2026 issue on the Berkeley and DeepMind work, where pronoun use dropped 40 to 60 percent under LLM revision and adjectives climbed 57 to 90 percent. I am now telling you that a peer-reviewed paper confirms a thing I have been asserting for a while. Notice how satisfying I find that. Then notice that I write this newsletter with an AI, that the previous sentence went through it, and that if the paper is right you cannot fully tell from the outside how much of this paragraph is me.


The other controlled language

Set the Nature paper next to ASD-STE100 and the comparison does something the paper alone does not do. Both systems reduce linguistic variance. Both are effective. Only one of them told anybody.

ASD-STE100LLM polishing
Rules 53, published, numbered, revised by a standing committee across three-year cycles Unknown. Emergent from pretraining statistics and from the preferences of a small annotator pool during alignment
Vocabulary ~900 approved words, most permitted in one sense and one part of speech, with the unapproved alternatives listed No list. The constraint is a probability distribution and cannot be inspected by the writer
Declared scope Procedures and descriptions in technical documentation. Explicitly not prose, not argument, not narrative None. Applied to admissions essays, floor speeches, performance reviews, condolence notes
Visible to the reader Yes. STE reads like STE. Nobody mistakes a workcard for a personal voice No. 87% of rewrites scored above 0.95 semantic similarity and human raters called them near-identical
Measured benefit Comprehension error rate fell from 18% to 14% across 16 workcards and 175 working technicians, concentrated in hard documents and non-native readers Not measured in this paper. The paper measures the cost side only
What it removes Ambiguity, synonyms, and identity, deliberately, because in a torque specification identity is noise Ambiguity, synonyms, and identity, incidentally, in documents where identity was the entire signal

The aerospace answer to Orwell's worry is a scope statement. STE is a cage the industry built on purpose and put around exactly one kind of writing, and its 18-to-14 result is honest about how modest the win is even inside that cage. A maintenance manual has no interiority to protect. The reason STE has never colonized the rest of English is not restraint on anyone's part. It is that STE is visibly difficult, visibly ugly, and unmistakably itself, so nobody was ever going to use it to write a wedding toast.

The second controlled language is easy, free, and pleasant. That is the whole difference.


Where this lands in medicine

The paper's own framing points at diagnosis, and it is not a stretch. Linguistic markers are used in research settings to track cognitive decline, to characterize depression, and in screening work on suicide risk. Those methods depend on the same individual variance the paper watched disappear. The paper's contribution is not that these tools will break. It is that they degrade quietly and in a specific direction, which is worse than breaking, because a broken instrument announces itself.

The operational version of this in health systems has nothing to do with research instruments. Patient-portal messages, intake narratives, and increasingly the clinical note itself now pass through a polishing layer somewhere between the person and the chart. Whatever signal a clinician's ear was picking up from how something was phrased is being normalized before it arrives, and no field in the record marks that it happened. I do not know how large that effect is. Nobody does yet. It is the study I would want somebody to run.

The useful move is the one the aerospace industry already made. Decide where the controlled language applies. Dialysis machine alarm text, technician procedures, and medication instructions are the aviation case almost exactly, and they should be as flat and unambiguous as STE can make them. A patient describing what is wrong is the opposite case, and should be preserved as written, with the polished version stored alongside rather than on top.

So What

STE is a controlled language that says out loud what it does and where it stops. It flattens English on purpose, in one room, for one reason, and it took a committee forty years and nine revisions to agree on which room. The second controlled language flattens the same things just as effectively, hands the result back sounding like you, and has no room, no committee, and no edge.

Orwell never argued that simplified language was the danger. He argued that a simplified language nobody had agreed to was.

Confidence: the homogenization finding is solid and now peer reviewed. The identity-erasure finding is real but rests on classifiers that were modest before any model touched the text. Per-trait figures cited here come from the February 2025 preprint, since the published tables sit behind a paywall; the aggregate figures come from the version of record. The comparison to ASD-STE100 is mine, not the authors'.

Sources

The paper (version of record): Sourati Z, Karimi-Malekabadi F, Ozcan M, et al. The shrinking landscape of linguistic diversity in the age of large language models. Nature Human Behaviour, published 24 August 2026. DOI 10.1038/s41562-026-02550-0. nature.com

The preprint: arXiv:2502.11266v1, 16 February 2025 (four studies; v2 posted 24 August 2026). arxiv.org

The earlier version of the same core result: Sourati Z, Ziabari A, Dehghani M, et al. arXiv:2404.00267, April 2024. arxiv.org

Simplified Technical English: ASD Simplified Technical English Maintenance Group, ASD-STE100 Issue 9, January 2025. asd-ste100.org

STE field evidence: Chervak S, Drury CG, Ouellette JP. Simplified English for aircraft workcards. Proc Hum Factors Ergon Soc 1996;1:303-307. And Shubert SK, Spyridakis JH, Holmback HK, Coney MB. The comprehensibility of Simplified English in procedures. J Tech Writ Commun 1995;25(4):347-369.

Basic English and Newspeak: Ogden CK, Basic English (1930); Orwell G, Nineteen Eighty-Four, appendix (1949). Orwell's BBC advocacy for Basic English dates to 1942-1944.

Prior WAiR coverage of this theme: What Adam is Reading, Week of 03-23-26, AI Impact section (Berkeley / UCSD / Washington / DeepMind semantic clustering study).