I have noticed an increase in reference questions that ask me, as an information specialist, to evaluate the quality and accuracy of artificial intelligence output. I do not think this is inherently a bad thing. However, I will admit that I do not find it the most fulfilling form of research.
Checking the output of a chatbot is, quite literally, why I changed my undergraduate major from computer science to history. Computer science largely consisted of debugging and, occasionally, angrily yelling [and praying] for my code to work. History, on the other hand, was about identifying patterns, uncovering hidden truths, and wrestling with difficult research questions.
Of course, I have friends who genuinely enjoy debugging code and find historical research to be the bane of their existence. To each their own.
For me, verifying AI output feels remarkably similar to writing proofs in high school calculus. First, you arrive at an answer. Then, you must prove that the answer is correct. Like many of my high school calculus assignments, I often discover that I was wrong, I switched a letter or hallucinated an integer and must go back to the chalkboard.
Sometimes the hallucinations are obvious, such as attributing the wrong author to the wrong article in the wrong journal. Other times they are subtle, such as mischaracterizing a single but important aspect of an otherwise legitimate source. Increasingly, I find that the time required to verify AI-generated information offsets much of the time saved during the initial search. The work has not disappeared, it has simply shifted from the original researcher to me, the human verification system.
When conducting traditional legal research, I approach sources with a critical eye, evaluate their authority, and piece together a narrative or legal argument as I move through the project. Verifying AI output requires all those same skills, but with an additional layer of scrutiny. I must confirm that the sources themselves are authentic, ensure they are not AI-generated, and then compare the original sources against the AI’s characterization of them. At times, the process feels like the most meticulous cite check imaginable…
That brings me to what I believe is one of the more significant challenges: hallucinated citations in academic writing.
I serve as the library liaison for two law journals, and both have encountered submissions containing suspicious citations. Over time, I have noticed several recurring patterns. Often, the citation includes a real scholar, a legitimate article title, and an actual journal, but the article was written by someone else. Other times, the journal published an article with that title, but under a different author. Occasionally, every element of the citation is individually correct, yet they have been combined into a source that never existed.
When we encounter these citations, I encourage the journal editors to return them to the author for verification. We need to know whether the cited article, author, journal, and title are actually the source supporting the proposition in the manuscript.
This process requires substantially more work than correcting a citation or locating a missing source. If two authors have similarly titled articles and the citation contains only partial errors, we often need to retrieve, read, and compare both articles before determining which one the author intended to cite. What might once have been a straightforward cite check becomes a lengthy investigation.
Put another way, AI often shifts work from the front end of research to the back end.
Ironically, I do use AI in my own research—but primarily at the end of a project. After completing my own research, I sometimes ask an AI system whether it identifies any additional sources I may have overlooked. Occasionally, it surfaces useful material that I missed. Used in this way, AI can save me time by supplementing a completed research process rather than replacing it.
More and more, however, I find myself feeling like a chatbot.
“Dear Librarian, is this ChatGPT or Claude response correct?”
Three meticulous hours later, my response is often:
“Dear X,
No—and here is why.”