Is ChatGPT Reliable for Thesis Writing in 2026? What the Research Shows

·

Is ChatGPT Reliable for Thesis Writing in 2026? What the Research Shows

Is ChatGPT reliable for thesis writing? For structure and language, largely yes. For facts and citations, the research says no — several peer-reviewed studies have measured ChatGPT fabricating a meaningful share of the academic references it generates, sometimes over a third of them. If you’re relying on ChatGPT for any part of your dissertation, the difference between those two answers is exactly where your risk sits.

This isn’t a hypothetical concern. As generative AI use among students has climbed sharply in the past two years, so has the volume of fabricated citations slipping into submitted academic work — and the data on why this happens, and how often, is now well documented enough to give you a genuinely evidence-based answer rather than a guess.

Quick answer: ChatGPT is reliable for structural tasks — outlining, rephrasing, explaining concepts — but unreliable for facts and citations. Peer-reviewed studies have found ChatGPT fabricates roughly 20-55% of citations depending on the model version and topic, and introduces errors into a further chunk of the real references it does generate. Treat every citation and factual claim from ChatGPT as unverified until you check it against the original source.

1. What the Studies Actually Found

The numbers are consistent enough to take seriously. A widely cited study published in Scientific Reports found that ChatGPT fabricated roughly 20% of the citations it generated outright and introduced errors into around 45% of the real references it produced. Separate analyses of literature-review generation found ChatGPT-3.5 fabricating between 39.6% and 55% of citations, while GPT-4 versions showed lower but still meaningful fabrication rates of roughly 18% to 28.6%.

More recent cross-model testing has found variation between AI systems too — one comparison of eight AI chatbots for bibliographic reference retrieval found some models like Grok and DeepSeek outperforming ChatGPT on accuracy, though none of the eight tested were fully accurate.

Reported ChatGPT Citation Fabrication Rates by Study
Study / Model Fabrication Rate
Scientific Reports (fabricated outright) ~20%
Scientific Reports (errors in real references) ~45%
ChatGPT-3.5, literature reviews 39.6%-55%
GPT-4 versions 18%-28.6%
Fabricated DOIs linking to unrelated real papers 64% of fabricated-DOI cases

Practical tip: The 64% figure above matters more than it looks — when a fabricated citation does include a DOI, most of the time that DOI actually resolves to a real, unrelated paper, which makes the error far harder to catch with a quick manual check than a citation that simply doesn’t exist at all.

2. Why ChatGPT Makes Up Citations

It’s a predictable consequence of how the model works, not a bug that gets fixed by asking nicely. ChatGPT generates text by predicting statistically likely word sequences based on patterns in its training data, rather than retrieving verified records from a live citation database. When you ask for a reference on a specific claim, it produces an author name, title, journal, and year that fit the statistical pattern of a real citation — without any built-in mechanism to confirm that combination actually exists.

Practical tip: This is also why simply asking ChatGPT “are you sure this citation is real?” doesn’t reliably fix the problem — the model can confidently reaffirm a fabricated reference just as fluently as it generated it the first time.

3. Reliability Depends Heavily on Your Topic

Close-up of a student cross-referencing a printed academic journal article with a laptop screen
Verifying a citation against its original source is the only reliable way to catch a fabricated reference.

Well-studied topics fare much better than niche ones. Research comparing fabrication rates across subject areas found sharp variation — citations related to well-documented conditions like depression were around 94% real, while citations for less-studied topics such as binge eating disorder and body dysmorphic disorder saw fabrication rates near 30%. If your thesis covers an emerging or niche area of your field, treat ChatGPT’s citation suggestions with extra scepticism rather than assuming the aggregate rates above apply evenly.

This pattern also shows up in publication data more broadly: the rate of fabricated references appearing in published papers has risen sharply, from roughly 1 in 2,828 papers in 2023 to about 1 in 458 in 2025, and reportedly around 1 in 277 in the first weeks of 2026 — evidence that this isn’t a problem confined to student drafts, but one that’s increasingly slipping past peer review too.

4. Where ChatGPT Is Genuinely Safe to Use

Structural and language tasks are where it earns its reputation. ChatGPT performs reliably well when you use it to brainstorm chapter structure, rephrase a dense paragraph for clarity, explain a concept back to you in plain language to test your own understanding, or draft an outline you’ll fill in yourself with verified content.

  • Outlining and structure. Low factual risk, since you’re organising your own ideas rather than sourcing new claims.
  • Clarity and flow editing. Rephrasing your own already-verified sentences carries minimal factual risk.
  • Explaining a concept in plain language. Useful as a comprehension check, provided you cross-reference anything unfamiliar against your course materials or a textbook.

5. Where It’s Genuinely Risky

Anything involving a specific fact, statistic, or citation needs independent verification. The riskiest uses are asking ChatGPT to generate a literature review’s reference list from scratch, asking it for specific statistics or study findings without a source you can check, and asking it to summarise a paper it hasn’t actually been given the text of.

  • Generating a reference list unsupervised. As the data above shows, a meaningful share of what comes back may not exist or may misattribute a real source.
  • Trusting specific statistics without a source. If ChatGPT states a percentage or figure without linking to a real, checkable source, treat it as unverified.
  • Summarising papers it hasn’t read. Unless you’ve pasted the actual text in, a “summary” of a paper by title alone risks being a plausible-sounding fabrication.

6. How to Verify Anything ChatGPT Gives You

Graduate student comparing an AI-generated draft with printed research papers at a library desk
Treat every AI-generated claim as unverified until you’ve checked it against a real source.
  1. Search the exact title in Google Scholar or your library database. If it doesn’t appear, treat it as fabricated until proven otherwise.
  2. Check the DOI resolves to the claimed paper, not just any paper — remember that a working DOI can still point to an unrelated real study.
  3. Cross-check specific statistics against the original source, not against ChatGPT’s own restatement of the source.
  4. Keep a running verification log as you write, noting which claims you’ve personally confirmed, so you’re not re-checking the same citation twice or missing one under deadline pressure.

7. Safer Alternatives for Citations Specifically

For citation generation specifically, purpose-built academic tools that pull from real reference databases and citation-management standards are inherently safer than a general-purpose chatbot predicting plausible text, because they format records that actually exist rather than generating statistically likely ones. Tesify’s automatic bibliography generator is built around this distinction — it formats citations from verified reference data rather than free-text generation, which removes the fabrication risk described throughout this article.

If you want a deeper comparison of what ChatGPT genuinely does and doesn’t handle well across an entire dissertation, our honest 2026 guide to ChatGPT for thesis writing covers the full picture beyond citations specifically. For checking your own writing before submission — a separate but related integrity concern — our ranked comparison of free plagiarism checkers tests the tools students actually use. And if fabricated or inconsistent references have already crept into a draft, our guide on how to reference correctly with AI walks through fixing the problem chapter by chapter. Tesify’s AI thesis writing platform is built specifically to keep citation and fact-checking workflows separate from free-text drafting, which is the core distinction this whole reliability question comes down to.

FAQ

Is ChatGPT reliable for thesis writing?

ChatGPT is reliable for structural and language tasks like outlining, clarifying prose, and reformatting, but it is not reliable for facts and citations. Multiple academic studies have found ChatGPT fabricates or introduces errors into a significant share of the references it generates, so any citation it produces must be independently verified before use.

How often does ChatGPT fabricate citations?

Reported fabrication rates vary by study and model version. A widely cited Scientific Reports study found ChatGPT fabricated roughly 20% of citations and introduced errors into around 45% of real references, while other analyses of ChatGPT-3.5 found fabrication rates as high as 39.6% to 55% in literature review contexts, with GPT-4 versions showing lower but still meaningful rates around 18% to 28.6%.

Why does ChatGPT make up fake references?

ChatGPT generates text by predicting statistically likely word sequences rather than retrieving verified records from a database, so when asked for a citation it can produce a plausible-looking author, title, and journal combination that simply does not exist, especially for less well-represented or niche topics.

Is it safe to use ChatGPT for parts of my thesis?

It can be safe for tasks like brainstorming structure, improving clarity, or explaining a concept back to you in plain language, provided you treat every factual claim and citation as unverified until you check it against the original source yourself.

What should I use instead of ChatGPT for citations?

For citations specifically, purpose-built academic tools that pull from real reference databases and citation-management systems are safer than a general-purpose chatbot, since they generate formatted references from records that actually exist rather than predicting plausible-sounding ones.

Do universities allow ChatGPT for thesis writing?

Policies vary significantly by institution and department. Many universities now permit AI assistance for tasks like editing and brainstorming while prohibiting it for generating original analysis or fabricating data, so check your specific institution’s academic integrity policy before relying on it.

Sources: StudyFinds — ChatGPT’s Hallucination Problem, STAT News — Fraudulent AI Citations in Academic Papers, arXiv — Assessing AI Chatbots in Bibliographic Reference Retrieval.

Write your thesis with AI

Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.

Leave a Reply

Your email address will not be published. Required fields are marked *