CiteProvebeta

Why AI chatbots invent citations — and how to catch them

In short

AI chatbots hallucinate citations because they generate text that statistically looks right, and a reference — author, year, title, journal, volume, pages — is a highly regular pattern a model can fabricate without any real source behind it. The only reliable way to catch them is to verify each reference against authoritative databases like Crossref, OpenAlex and PubMed, rather than trusting how well-formed it looks.

Language models generate text that looks right. A citation is a highly regular pattern — authors, year, title, journal, volume, pages — and a model can produce a perfectly formatted one without any underlying source existing. These fabrications have already reached courtrooms, peer review, and student work.

The four common fabrication patterns

  • Pure inventions — plausible authors and a plausible journal that simply do not exist. Easiest to catch: no database has them.
  • Real identifier, wrong paper — a genuine DOI or PMID attached to a fabricated title. Looks authoritative; resolves to something completely different.
  • Mangled real papers — right paper, wrong year; swapped author lists; a near-miss title with one word changed. The source exists, the citation is still wrong.
  • Chimeras — real authors from one paper, the title of another, the journal of a third.

How CiteProve verifies each pattern

Every reference is checked against Crossref, OpenAlex, PubMed, and CourtListener. Identifiers are resolved and then validated — a DOI that resolves to a different work than cited is flagged, not trusted. Field-level scoring catches mangled years, swapped authors, and near-miss titles, and shows you exactly what differs with a corrected citation ready to copy.

Paste your reference list and check it →