AI is doing genuinely impressive work in archaeology. I want to say that clearly before I say anything else. Researchers are reading scrolls that have been sealed shut for two thousand years. They are reassembling shattered hymns from fragments scattered across museums on different continents. They are reconstructing the rules of games that nobody has played since the fall of Rome. This is real.
It is also consistently overstated. I have been collecting claims about AI in archaeology for a few weeks. I asked Perplexity to verify the numbers and timelines. Then I pulled the primary sources myself (a dozen articles via Firecrawl, plus the original journal papers) to check what the actual researchers said versus what the press reported they said.
The pattern is the same one we see in clinical AI coverage. Each relay layer (journal to science press to general press to social) retains the conclusion, drops the caveats, and inflates the adjectives. By the time a finding reaches your feed, "a family of plausible rulesets" has become "AI cracked the rules." A partial reading of a damaged papyrus has become "full text of a book nobody knew existed." An incremental refinement of a known burial area has become "Plato's exact grave found."
Here are the six claims. I scored each one.
A Yamagata University and IBM team published in PNAS (September 2024) showing their AI identified 1,309 candidate sites on the Nazca Plateau. Human field teams with drones verified about a quarter of them. They confirmed 303 new figurative geoglyphs in six months, nearly doubling the 430 previously known. The AI increased the discovery rate 16 fold. One confirmed geoglyph per 36 AI flagged candidates.
What the headlines inflatedVery little, actually. This is one of the cleaner cases. The main distortion is telescoping years of iterative development (IBM and Yamagata had already found 143 geoglyphs with AI in 2019) into a narrative of sudden magic. The AI flagged candidates. Humans walked to them.
Enrique Jiménez at LMU Munich, collaborating with the University of Baghdad, used the Electronic Babylonian Library (a platform hosting 1,402 digitized cuneiform manuscripts) to match 30 clay tablet fragments scattered across multiple museum collections. The AI used n-gram matching of cuneiform sign sequences to flag possible joins between fragments that had been separately cataloged for decades. The result is a complete 250 line hymn to the city of Babylon, describing its temples, gardens, treatment of foreigners, and the role of priests.
What the headlines inflated"Rediscovered after 3,000 years" is slightly misleading. The fragments were excavated and sitting in museums. What is new is the recognition that 30 separately cataloged pieces belong to a single text. The "AI" here is closer to sophisticated text matching than the generative AI of popular imagination. Less glamorous than it sounds. More useful than it looks.
A team at the University of Groningen built an AI model called "Enoch," trained on 25 radiocarbon dated scrolls. Enoch learned correlations between letter shapes and absolute dates, then applied that model to 135 additional scrolls. Published in PLOS ONE (June 2025), the results generally push dates 50 to 150 years earlier than the prevailing paleographic consensus. The model correctly sequenced 24 of 25 training samples.
What the headlines inflatedA lot. Christopher Rollston at George Washington University told the Biblical Archaeology Society the results are "very interesting" but "not earth-shattering" because many conclusions align with what the great paleographers said 60 years ago. The study's own co-author (Popović) emphasizes the value is in dating individual manuscripts, not in a blanket "everything is older" claim. The press stripped this nuance. "Near-original" framing for Daniel fragments and theological spin about rewriting biblical history are overinterpretations. The single point of failure: the entire model depends on 25 calibration anchors being representative.
Archaeologist Walter Crist found a carved limestone slab in a Dutch museum (Het Romeins Museum, Heerlen) during COVID lockdowns. Microscopic use-wear analysis showed localized abrasion patterns consistent with game piece sliding. The Digital Ludeme Project's AI system (Ludii) then simulated thousands of rounds across many candidate rulesets and compared predicted wear patterns to the physical evidence. Published in Antiquity (2025), the best-fitting rules describe an asymmetric blocking game: four "dogs" try to trap two "hares." You can play it online.
What the headlines inflated"Cracked the rules" implies a single definitive solution. In reality the method identifies a family of plausible rulesets that best explain the wear data. The archaeological context is also weaker than ideal (the stone lacks excavation provenance). But the finding that blocking games existed in Roman Europe, previously thought to arrive only in the Middle Ages, is legitimate and interesting.
This is the big one. The Vesuvius Challenge, launched in 2023 by computer scientist Brent Seales and Silicon Valley backers, has awarded over $1.8 million in prizes for reading scrolls carbonized by the eruption of Mount Vesuvius in AD 79. In February 2024, three students won the $700,000 grand prize after reading 2,000+ Greek letters from a sealed scroll using AI ink detection on synchrotron CT data. In June 2026, the team fully unwrapped PHerc. 1667 for the first time, revealing 1.5 meters of text across 20 columns. Papyrologist Federica Nicolardi dates it to the 2nd or possibly 3rd century BC, potentially the oldest scroll in the collection. The text discusses Stoic philosophy (impulse, practical wisdom, the nature of good and evil). A separate scroll revealed "On Gods, Book 8" by Philodemus, proving that series was at least 8 books long when only Books 1 and 3 were previously known.
What the headlines inflated"Full text" and "a book nobody knew existed" are hype. Nicolardi herself notes that interpreting a text from which you can read only the bottom few lines of the final pages is "tricky." The process remains slow (each scroll requires individual AI parameter tuning). Seales' vision of automated 90% accuracy within 24 hours is aspirational. Nearly all progress depends on one lab (Seales at Kentucky), one synchrotron facility (Diamond Light Source), and one funding structure (the Challenge). Those are single points of failure dressed up as a broad scientific movement. The technology is transformative. The relay language inflates at every step.
Graziano Ranocchia's team used AI-assisted imaging to read about 1,000 new words (30% of the text) from a Herculaneum papyrus. The text appears to refine Plato's burial location to a specific garden within the Academy in Athens, near a small shrine called the Museion. This narrows a location previously known only as "somewhere in the Academy grounds." The text also yielded new information about the timing of Plato's enslavement.
What the headlines inflatedSubstantially. "Found Plato's exact burial spot" overstates what the text delivers. The Museion no longer stands. The claim refines a textual reference, not a GPS coordinate. The dramatic "final night critiquing a Thracian flute player" anecdote, which appeared in press coverage, does not appear in the available scholarly reports and is likely narrative embellishment at the Daily Mail relay layer (it may derive from a separate ancient biographical tradition that predates this papyrus). The distance between Ranocchia's careful language and the press headlines is the largest gap of any claim in this review.
The thread connecting all six claims is not the AI. It is the information supply chain.
At the primary research layer (PNAS, PLOS ONE, Antiquity, conference presentations), the language is appropriately cautious. Researchers describe families of plausible rulesets, provisional dates, partial readings. At the quality science press layer (Nature News, the Guardian, National Geographic), the caveats are mostly preserved but timelines get telescoped and complexity gets smoothed. At the general press layer (CNN, Daily Mail, BBC), the caveats disappear and the verbs become "cracked," "decoded," "revealed." By the social and AI aggregator layer, the finding is fully laundered into a single triumphant sentence.
This is the same mechanism that transforms "a phase 2 trial showed modest improvement on a surrogate endpoint" into "breakthrough drug shows promise." It is the same mechanism that transforms "our clinical decision support tool was trained on a curated dataset" into "AI-powered diagnostics." The relay strips the uncertainty. The adjectives do the rest.
I should also note: "AI" is doing a lot of rhetorical work across these six claims. The actual techniques range from n-gram text matching (Babylon) to synchrotron imaging with machine learning ink detection (Herculaneum) to game-theoretic simulation (Roman board game) to paleographic pattern recognition (Dead Sea Scrolls). Calling them all "AI" implies a coherent technological revolution. What actually exists is a set of specialized computational tools, each impressive on its own terms, each dependent on deep human domain expertise to mean anything.
The technology is genuinely impressive. The hype is genuinely unnecessary. Tell your dinner companions that AI is producing real results in archaeology and that every specific claim they have seen has been inflated at least one notch by the time it reaches their feed.
Take the field seriously. Take any individual headline with a calibrated grain of salt.
Nazca: Sakai et al., PNAS (Sep 2024). pnas.org · Gigazine · BBC Newsround · SCMP · Heise
Babylonian Hymn: Jiménez et al., LMU Munich/Univ. Baghdad. Techno-Science · AInvest · Electronic Babylonian Library
Dead Sea Scrolls: Popović et al., PLOS ONE (Jun 2025). plosone.org · Biblical Archaeology Society · El País · DW
Roman Board Game: Crist et al., Antiquity (2025). ZME Science · Earth.com · Scientific American
Herculaneum Scrolls: Vesuvius Challenge. scrollprize.org · Guardian · CNN · Nat Geo · Nature
Plato Burial: Ranocchia et al. (Apr 2024). Daily Mail