A thread came across my feed claiming that AI is about to uncover an entire hidden world of animal communication, and that what it has already found is insane. Eight bullet points followed. Elephants with names. Monkeys with names. Bats arguing about sleeping arrangements. Whales with a phonetic alphabet. Birds talking to a language model. Dolphins being taught a shared vocabulary. Crows announcing themselves at the nest. A robot bee giving directions to a hive.
I went and pulled the papers using Claude, which helped me find them, summarize the research, and work out what was really going on in each one. Six of the eight are real and reported accurately. One contains a fabricated detail, and it happens to be the most impressive-sounding part of that bullet. One is genuinely wonderful and has nothing whatsoever to do with artificial intelligence.
But the interesting problem is not the error rate. It is that every single one of these systems does the same thing, and it is not the thing the thread says it is doing.
Accurate, and the underlying work is careful. Rumbles recorded at Samburu and Amboseli between 1986 and 2022. A random forest identified the intended recipient of 27.5 percent of calls, well above its own permuted-feature null. In playbacks to 17 elephants, animals approached the speaker roughly 128 seconds sooner, vocalized 87 seconds sooner, and produced about 2.3 times more calls when the rumble had originally been addressed to them.
Two things the thread leaves out. 27.5 percent is a real signal and a modest one, which is what you would expect if labels appear in only a minority of calls, and the authors say as much. And the claim that the labels are arbitrary rather than imitative is the genuinely novel part, more novel than the word "name."
The number is roughly right and the finding is genuinely disputed in the peer-reviewed literature, which the thread does not mention. The study covers ten captive marmosets from three family groups, not a wild population.
Jaakkola's objection is the good kind. If a caller subtly shifts its own phee toward its partner's call, a behavior already documented in other species, you would get exactly this pattern. Calls to a given receiver classify as addressed to that receiver. Receivers respond more to calls that resemble their own. Same-family calls cluster together, because phee calls already encode the caller's social group. No naming required. Omer's group answered in 2025, arguing that cross-caller models reveal family-level label conventions and that accommodation does not account for the effect.
What the data actually show is that marmoset phee calls carry receiver-specific information. Whether that information is a name or a side effect of vocal convergence is unresolved as of now.
The four categories are real. Food, sleeping position, unwanted mating advances, and perching too close. So is the finding that calls are directed at individuals rather than broadcast, which was the actual contribution.
The colony of thousands is invented. This was 22 captive bats in acoustically isolated chambers, monitored continuously for 75 days, yielding roughly 15,000 vocalizations for which both context and identities could be established. Wild colonies of thousands are where the authors said the method might someday go. The addressee identification the thread describes as working "inside a colony of thousands" was in fact the weakest result in the paper, closer to a coin flip on the addressee's sex than to picking a bat out of a crowd.
Also worth noting. This is 2016, and the tool was a modified speaker-verification algorithm of the kind that had been shipping in consumer products for years. It sits inside a thread about what AI is "about to" do.
Both numbers are exactly right, which is unusual for a thread like this. The paper states that at least 143 feature combinations are frequently realized, and describes a repertoire nearly an order of magnitude larger than previously believed. The four features are rhythm, tempo, rubato, and ornamentation. Two are context independent and two are modulated during exchanges, which is the actual news.
The scope is narrower than it sounds. 8,719 codas, one clan in the Eastern Caribbean, roughly 60 animals. Not sperm whales in general.
The word "alphabet" is doing promotional work rather than technical work, and linguists said so within a day. What was demonstrated is a combinatorial coding system with more distinguishable units than anyone had catalogued. What was not demonstrated is that any unit carries meaning. Sharma has been consistent and public about this. They do not know what the whales are saying.
The claim is accurate to the preprint. More than 1.5 million female zebra finch calls, more than 1,000 hours of interaction, and a generative audio model that exchanged calls with live birds in real time. When birds interacted with the model, their vocal production and flexibility resembled natural exchanges in ways that did not occur with fixed playback. Ablations showed that call timing and call structure contribute differently.
Two caveats the thread drops. This is a preprint, and the rapid retitling across three versions is the visible fingerprint of reviewers pushing on the framing. And the birds were all female, which matters because zebra finch song and calling are strongly sex-differentiated.
What was shown is that a model can produce timing and structure sufficient to sustain a natural-looking exchange. That is a claim about interaction dynamics. It is not a claim that anything was said.
Accurate, and notably the most honestly worded bullet in the thread. It says "the goal is," which is the correct tense. No wild dolphin has requested an object. The system is deployed and the experiment is ongoing.
Worth being precise about what DolphinGemma is. It is a next-token predictor over dolphin audio, which makes it useful for spotting recurring structure and for recognizing in real time when a dolphin has mimicked a synthetic whistle amid ocean noise. The synthetic whistles are invented by the researchers and are deliberately unlike natural dolphin sounds. If this works, humans and dolphins will share a small artificial vocabulary that humans wrote. That is a real and interesting achievement. It is not translation of dolphin.
Accurate, and the thread preserved the paper's hedge, which is more than most reposts manage. The paper says these calls may announce nest visits, and the thread says may.
The methodological point is the real story here and it is not about AI. These are low-amplitude calls that directional microphones simply never picked up. Putting a 12.5 gram recorder on the bird is what made them audible. The machine learning came second, to sort 127,000 detections into tagged adult, other adult, chick, and parasitic cuckoo nestling. New sensor first, new classifier second, new biology third. That order matters.
Preprint status on the repertoire mapping, so treat the specific call functions as provisional.
The experiment is real and it is my favorite thing in the whole list. It is also the cleanest counterexample to the thread's own thesis.
There is no AI in it. A robot performed a motor pattern and live bees decoded it. The reason that worked is that von Frisch and successors spent seventy years establishing what the waggle dance means, click by click, using nothing more computational than patient observation and, later, harmonic radar. The robot is a test of a theory, not a discovery engine.
The overstatement is small but real. "The bees changed their flight paths based on its instructions" reads as reliable command. What the radar showed was that bees which followed the robot adjusted their flight paths toward the indicated vector, and the lab's own later description concedes it worked at least sometimes.
Here is the part worth sitting with. This is the only entry on the list where a human unambiguously sent a specific, novel, semantic message to an animal and the animal acted on the content. It required no machine learning at all. It required knowing what the signal meant.
Laid side by side, the pattern is hard to miss. Every model in this thread takes acoustic input and assigns it to a category, predicts the next unit, or generates a plausible continuation. None of them recovers meaning. Meaning, where it exists at all in this list, was established the old way.
| System | Computational task | Output |
|---|---|---|
| Elephant rumbles | Random forest classification | Which elephant a call was for |
| Marmoset phee calls | Random forest classification | Which marmoset a call was for |
| Bat squabbles | Speaker-verification classification | Who called, and about what, from four bins |
| Sperm whale codas | Clustering and feature decomposition | A larger inventory of distinguishable units |
| Zebra finch calls | Generative audio model | Plausibly timed calls that sustain an exchange |
| DolphinGemma | Next-token prediction | Likely next sound, plus mimicry detection |
| Crow biologgers | Detection and attribution | 127,000 calls sorted by who made them |
| RoboBee | None | A destination, understood and acted upon |
The best piece on the waggle dance robot, and the one that puts it in its actual context. Landgraf's group also built a system that watches real dances and translates them into a map, which means the hive can be read as well as written to. The stated ambition is steering colonies away from pesticide-contaminated forage. Note the hedging in the prose, which is more careful than the version that went viral.
Linguists reacting in near real time to the phonetic alphabet paper, and not gently. The core objection is that an alphabet is a writing system, so the metaphor concedes the thing it is trying to prove. Useful as a specimen of what happens when a computational team publishes into a field with its own hundred-year vocabulary.
Kelly Jaakkola's response to the marmoset paper is the most instructive four pages in this whole file. A domain expert reads a machine learning result and identifies, without rerunning anything, the confound that would produce the identical pattern. This is what the field needs more of and gets less of every year.
Pardo writing about his own study for a general audience, with the effect sizes intact and the claims sized correctly. A model for how this should be done.
Worth reading for the hardware more than the models. Twelve and a half grams, three breeding seasons, tags that release themselves after about eighteen days. The quiet calls were always there. Nobody could hear them.
Machine learning has established that animal signals carry far more distinguishable structure than we could hear unaided. It has not established what any of it means, and the researchers themselves keep saying so.
The distance between "these two calls are reliably different" and "this call means come here" is the entire problem. It is not a data problem, and more parameters will not close it. Meaning has to be anchored to behavior, in context, over years, by someone watching.
Sources
Elephants: Pardo MA, et al. Nature Ecology & Evolution 8:1035 (2024). 10.1038/s41559-024-02420-w
Marmosets: Oren G, et al. Science 385:996 (2024). 10.1126/science.adp3757. Critique: Jaakkola K, Learning & Behavior (2025). Response: Oren G, et al., Learning & Behavior (2025).
Bats: Prat Y, Taub M, Yovel Y. Scientific Reports 6:39419 (2016). 10.1038/srep39419
Sperm whales: Sharma P, et al. Nature Communications 15:3617 (2024). 10.1038/s41467-024-47221-8. Coda vowels: Beguš G, et al. Proc R Soc B 293:20252994 (2026).
Zebra finches: ZF-AIM preprint, bioRxiv 2026.02.12.705387 v3 (2026). Not peer reviewed.
Dolphins: DolphinGemma and CHAT, Google with Georgia Tech and the Wild Dolphin Project (2025). blog.google
Crows: Animal Cognition (2025), 10.1007/s10071-025-02018-0. Repertoire mapping: bioRxiv 2026.04.02.715916 v2 (2026). Not peer reviewed.
Bees: Landgraf T, et al. Dancing honey bee robot elicits dance-following and recruits foragers (2018). arXiv:1803.07126
Note on provenance: The eight claims were transcribed from a social post supplied by the reader. The post itself is not cited as evidence for anything. Every claim was checked against the primary literature independently.