In 2022, “delve” was a verb used modestly. Two years later, it had become a linguistic flare. A Science Advances analysis of more than 15 million PubMed abstracts found that certain words associated with large language models (LLMs) rose at a speed unmatched by earlier shocks to scientific writing, with at least 13.5 percent of 2024 biomedical abstracts showing signs of LLM assistance.
Beyond the word itself, what really matters is the speed at which a statistical preference became a public trace.
By June 2026, the obvious signals of so-called artificial intelligence (AI) writing are already fading. Writers and editors know that “delve,” “intricate,” and “meticulous” can make a text sound polished by a machine, just as model providers adjust outputs to make the machine read less like one. A recent study on human-LLM coevolution in academic writing finds exactly this moving pattern: conspicuous AI markers decline once they become socially visible, while subtler terms continue to spread.
That is precisely the tipping point of AI-mediated language. Essentially, the issue has moved beyond awkward prose, and machine suggestions are becoming part of the infrastructure through which societies write, translate, summarize, govern and think.
Language as Civic Infrastructure
Policy makers usually regulate AI through familiar categories: privacy, safety, discrimination, intellectual property, labour and security. Language now belongs in that list. Wittgenstein’s warning that intelligence can be “bewitched” by language is especially relevant in an age when machines generate the words through which institutions explain, justify and decide. Language is the operating system of public life. Laws, school curricula, medical guidance, research summaries, grant applications, procurement notices, public service chatbots and diplomatic texts all depend on words that carry precision, context and trust. When language becomes automated at scale, errors in meaning become errors in governance.
When AI becomes a routine layer between thought and written expression, the machine does more than correct grammar and starts to propose what sounds professional, confident or persuasive. Each accepted suggestion shifts expectations for how we communicate. Over time, “good writing” may begin to resemble the statistical preferences of AI-mediated language itself. The strongest evidence so far comes from domains where text can be measured at scale. Scientific publishing already shows clear traces: a large-scale analysis of more than 15 million biomedical abstracts found an abrupt post-ChatGPT rise in certain style words associated with LLM-assisted writing. While this evidence should not be stretched into a sweeping claim about everyday speech, it might serve as a flag to note. Science, education, administration and workplace communication are early warning systems because they already rely heavily on AI-mediated writing.
AI writing assistance offers real gains, that much is true. For example, it can help non-native speakers cross language barriers, make bureaucratic language clearer, support people with disabilities and reduce the time spent turning knowledge into acceptable prose. For researchers, civil servants and students working outside dominant language norms, that can open doors to additional forms of expressing their thoughts, and lead to previously inaccessible professional opportunities.
However, the bargain becomes costly when fluency substitutes judgment. A Royal Society Open Science study on LLM-generated scientific summaries found that many models overgeneralized research findings, presenting narrower results as broader conclusions. This matters for policy, as a summary that makes a limited study sound universal can travel faster than the caveat it erased.
The same pressure appears in writing style. Cornell researchers found that AI writing suggestions made Indian and American participants’ writing more similar, with Indian participants shifting more toward Western styles. The change was subtle: fewer local markers, more generic phrasing, smoother sentences and less cultural texture. Nothing looked broken on the surface, but that is precisely why it matters.
This imbalance also sits inside the machinery. Many AI systems process language in tokens, and languages do not tokenize equally. Research on tokenizer fairness found that the same content can require vastly different token lengths across languages, affecting cost, latency and usable context. A 2025 study on the “token tax” reached a similar conclusion for African languages: inefficient tokenization can raise computational costs and depress accuracy.
This is an infrastructure bias. An English speaker may get cheaper, faster, better support than a speaker of a morphologically complex or underrepresented language, even before the model begins to “understand” the request. Policy makers who care about digital inclusion must therefore look below the visible interface. Fairness begins in data, tokenization, benchmarks, pricing and procurement.
That said, communities are now responding to this bias. Nigeria’s N-ATLAS is building open-source AI for Nigerian languages. India’s Bhashini and BharatGen pursue multilingual public infrastructure. Aya Expanse and Portuguese-focused models such as Tucano show that alternatives to English-centric defaults are possible. Malaysia’s ILMU is trained on Bahas and tailored in harmony with local culture.
There is also a cognitive dimension, although the evidence should be handled with care. A 2025 MIT Media Lab preprint, based on a small electroencephalogram (EEG) study, reported lower neural connectivity among participants who used ChatGPT for essay writing compared with those who used search or no tool. The study is preliminary and deserves cautious use, but it reinforces a practical insight every writer recognizes: drafting is more than output. A society that automates drafting too early may gain speed while weakening the habits that make judgment possible: hesitation, revision, doubt and the search for a word that has not already been offered by the machine.
The Policy Task
The next phase of AI governance should treat language as public infrastructure. The European Union’s General-Purpose AI Code of Practice has advanced transparency obligations for model providers. The United Nations Educational, Scientific and Cultural Organization’s (UNESCO’s) Global Roadmap for Multilingualism in the Digital Era goes closer to the linguistic core, calling for community involvement, multilingual standards and data sovereignty.
Four policy moves should follow. First, require linguistic impact assessments for AI systems used in education, health, justice, science and public administration. These assessments should test whether systems preserve meaning, cultural markers and dialectal variation.
Second, benchmark models by language, dialect and register, rather than reporting averaged “multilingual” performance. A model that works in formal French may still fail in Haitian Creole, Wolof, Nigerian Pidgin or rural speech patterns.
Third, make procurement contracts reward plural expression. Governments should ask vendors how their systems perform across local languages, how they handle tokenization disparities, whether communities contributed to data sets and how users can preserve their own voice.
Fourth, integrate double literacy from kindergarten, covering human and algorithmic literacy to sharpen awareness of the direct and indirect consequences of generative AI on our ability to think, communicate and understand in a hybrid context. Students, teachers and professionals across sectors need more than prompt engineering skills.
Ultimately, the aim is to govern the conditions under which language changes. Language has always evolved through contact, commerce and power. The task now is to prevent a handful of systems from narrowing the expressive range through which billions of people participate in public life.
The best AI writing tool should help a person sound more precise, more grounded and more fully themselves. The best AI language policy should make that possible for every language community, including those currently underrepresented in the machine.