Anagha Nair
Anagha Nair
Published: September 4, 2026

When the Footnote Points Nowhere: How Unverified AI-Generated Citations Are Dismantling Scholarly Evidence

When the Footnote Points Nowhere: How Unverified AI-Generated Citations Are Dismantling Scholarly Evidence

Science does not advance one paper at a time. It advances through a vast, interconnected web in which every study builds on, cites, and connects to the studies before it. A single citation is a thread; millions of citations together form the fabric that allows researchers to trace a claim back to its origin, verify it, and stand on it to reach the next discovery.

That fabric is enormous. The open scholarly graph indexed by OpenAlex now contains over 250 million works connected by more than 2 billion citation links. Each citation represents a small act of trust: an author asserting that a real source, somewhere in the record, supports what they have written.

Authors misusing generative AI are now quietly cutting those threads, and, more dangerously, weaving in threads that lead nowhere. The problem isn’t only that AI can produce fabricated references (a well-documented issue). It’s that phantom citations break the connectivity of scientific knowledge itself, and because of how citation actually works, a single fabricated citation can spread across the network long after its original source has been forgotten.

A Phantom Citation Is Not Just a Mistake; It’s a Broken Node

Large language models such as ChatGPT, Gemini, and Claude generate text by predicting plausible sequences of words. They may not query authoritative bibliographic databases when generating references. So, when a model produces a citation, it is producing something that looks like a legitimate citation—a credible author string, a real-sounding journal, a plausible year and DOI—without any guarantee that the underlying work exists.

The scale of the problem is now measurable. A 2025 evaluation of GPT-4o citations in mental health research found that of 176 generated references, 35 (19.9%) were fabricated outright, and of the 141 that did exist, 64 (45.4%) contained bibliographic errors — most often incorrect or invalid DOIs. That leaves roughly 44% fully accurate. Another recent study on AI-assisted literature retrieval in nephrology reported that 31% of references were fabricated outright and 68% of the links were incorrect.

It is tempting to file this under “quality control.” But a phantom citation is structurally different from an ordinary error, and the difference matters for the whole network:

  • A real but flawed citation is a node you can find and fix. Even if a paper contains errors, is weak, or is later retracted, it still exists. It has a DOI. It can be corrected, retracted, or otherwise flagged within the scholarly record.
  • A phantom citation is an edge that points to a void. There is no node to retract, no author to contact, no DOI to mark. The scientific record has no mechanism to “withdraw” a paper that was never written. The thread is anchored to nothing, yet it sits in the reference list looking exactly like every legitimate thread around it.

When these phantom edges enter the literature, they don’t just sit harmlessly. They begin to spread.

How One Fake Citation Becomes Many

Here is the mechanism that can turn an isolated error into a network-wide contamination event. Researchers do not always consult the original sources they cite. Instead, they may reuse references from the bibliographies of other papers.

In an analysis, Simkin and Roychowdhury studied the spread of identical misprints in citations to high-profile papers and concluded that roughly 70–90% of scientific citations are copied from other papers’ reference lists rather than read from the original. In their dataset, a single misprinted citation propagated 78 times—copied and copied again, by authors who never looked at the source. The frequency of repeated misprints followed a power law, the signature of a self-reinforcing network process.

Now apply that same mechanism to AI-generated phantom references. A fabricated citation only has to slip into one published paper—past a rushed co-author, an overstretched reviewer, an editorial pipeline that assumes accuracy was checked “somewhere.” Once it is in the record, the same citation-copying dynamic that propagates misprints can propagate fabricated references. The citation may get reused, re-listed, and indexed. “Looks cited” gradually hardens into “is cited.” The phantom acquires the one thing it never earned: apparent legitimacy.

This is why phantom citations are a connectivity problem, not merely an accuracy problem. They don’t just add noise; they create false paths through the knowledge graph—paths that future researchers, databases, and automated tools will follow as if they were real.

The Precedent We Should Already Be Worried About: Zombie Citations

We don’t have to speculate about what happens when broken nodes persist in the citation network. We have decades of evidence from retracted papers, and it is sobering.

Retraction is intended to remove a discredited paper from the active record. In practice, retracted papers continue to be cited as though nothing happened; bibliometricians call them “zombie papers.” A large biomedical citation-context study published in 2021 found that only about 5% of post-retraction citations acknowledged the retraction at all. A 2025 scientometric analysis tracked 35,514 retracted publications in Scopus and documented how their “lost citations” continue to ripple through literature reviews and meta-analyses years later.

If the network struggles this badly to suppress real papers that have been formally flagged, consider how hopeless it is to suppress phantom papers that were never flagged; because they never existed, and nobody knows they’re fake. Retracted papers are at least visible. Phantom citations are invisible by construction.

Where It Hurts Most

The damage becomes concrete at exactly the points where science checks itself.

Reproducibility

Reproducibility depends on being able to trace a method or claim back to its source and re-run it. A phantom citation severs that trace at the root: a reader who tries to follow the reference to verify a method finds nothing there. The chain of evidence simply ends abruptly.

Meta-analyses and systematic reviews

A meta-analysis works by traversing the citation graph, gathering the body of evidence on a question, and aggregating it into a pooled conclusion that then informs guidelines and clinical practice. They are only as trustworthy as the network they walk through.

The retracted-paper precedent shows precisely how this fails. In one Science-reported analysis, of 88 papers that cited a fraudster’s retracted work, 39 had conclusions that would have been substantially weaker if the discredited studies were removed.

A reviewer can, in principle, exclude a retracted study from a meta-analysis—it’s identifiable. A phantom study cannot be excluded, because it can’t be reliably identified: it has no retraction notice, no flag, no record of non-existence. It can only quietly inflate the apparent weight of evidence behind a conclusion that may rest, in part, on sources that were never real.

A Corpus That Trains the Next Model

There is a longer-term reason to act now. The same literature that AI is contaminating is the literature that future AI systems will be trained on. Undetected phantom citations and AI-distorted text that settle into the permanent record become training data for the next generation of models, which may then reproduce and amplify the same fabrications. Left unchecked, the network of scholarly evidence risks degrading into a system that increasingly cites itself citing things that were never there.

What Responsible Use Looks Like

None of this is an argument against AI in research. Used well, generative tools are genuinely useful for organizing, summarizing, and formatting. The argument is for verification-first habits that protect the network’s integrity — and those habits work best as two layers, not one.

The first line of defense is automated verification. Tools that cross-check a reference against real bibliographic databases — CrossRef, OpenAlex, PubMed, Semantic Scholar — can catch a large share of fabricated citations mechanically: a DOI that doesn’t resolve, an author-title-journal combination that doesn’t exist anywhere in the record. This is existence-checking, and it’s exactly the kind of task automation scales well. Increasingly, the fix belongs even earlier in the pipeline: AI writing tools built on retrieval-augmented generation, which query real databases at the point of drafting rather than generating references from memory, prevent phantom citations from being created in the first place.

The second line of defense is human judgment — and it can’t be automated away. A citation-checking tool can confirm a paper exists. It cannot confirm that the paper actually says what the author claims it says. That’s a question of fit between a claim and its evidence, not a lookup problem, and it’s precisely the failure mode that slips past every automated check: a real, resolvable, perfectly legitimate citation attached to the wrong conclusion.

Major publishers, research integrity organizations, and AI governance frameworks consistently emphasize that authors remain fully accountable for the accuracy of every reference, regardless of whether AI was involved. The citation network is shared infrastructure, and protecting it requires both layers of defense working together: automation to catch what shouldn’t exist, and human expertise to catch what shouldn’t be claimed.

Disclosure: The above image is generated using NotebookLM for illustrative purpose.

So, Start With Human Oversight

That is exactly the principle at the heart of Enago’s Responsible Use of AI initiative: AI should augment human scholarship, never replace the human judgment that keeps the scientific record trustworthy. Most major publishers now expect that a manuscript be reviewed by a human subject-matter expert before submission.

Phantom citations are difficult to detect. They typically do not trigger plagiarism detection software, they don’t fail statistical checks, and they look indistinguishable from legitimate references right up until someone tries to follow them. The single most effective safeguard remains a qualified human reviewer who verifies the evidence before it becomes a permanent part of the record.