Nick Evershed, Ariel Bogle and Mike Hohnen 

‘Scary’: how misinformation and AI hallucinations are infiltrating Australia’s parliament

Exclusive: Guardian analysis reveals that dozens of policy submissions from across the political spectrum incorrectly summarise real research and invent or wrongly cite sources
  
  

Illustration of Parliament House, with the shadows filled with binary code
In some cases, committee reports have cited submissions in which the majority of sources appear to be AI-generated – because they do not exist. Composite: Getty Images/Guardian design

Australia’s government inquiry process is supposed to help parliament make better decisions, hearing from experts and constituents alike. But Guardian Australia can reveal that the system is being flooded with AI-generated material, which is inventing studies and attributing nonexistent research to real academics and authors.

In some cases, committee reports have cited submissions in which the majority of sources appear to be AI-generated “hallucinations”, when large language models (LLMs) invent content that looks real but doesn’t actually exist.

An AI hallucination takes place when an LLM like ChatGPT or Claude outputs text that is plausible but incorrect, or confidently generates information that appears relevant but is actually unrelated.

This happens because these AI systems use complex statistics to figure out the text they generate and they have no innate ability to distinguish between correct and incorrect content. They simply predict the next most likely bit of text based on their training data, with a certain degree of creativity or randomness involved.

Hallucinations can become more common when the AI is asked to generate content on topics that are not well covered in its training data.

These hallucinations are not bugs and although the frequency can be reduced by increasing the amount of training data or adding web-search results and other information, they are an inherent feature of LLMs and can never be fully eliminated.

If the phenomenon continues, there is a significant risk of parliamentarians “making decisions based on evidence that doesn’t exist”, says Christian Downie, a professor in the Australian National University’s school of regulation and global governance.

The situation is made more confusing by AI services including Google search. Previously, looking up an invented reference might produce results that would indicate the reference did not exist. Now Google’s AI summary will sometimes create summaries of a fake reference as though it were genuine.

In some instances, Google’s AI summary and ChatGPT will also cite the inquiry submission with the fake reference as a source, creating an ongoing cycle of misinformation.

One document submitted to an inquiry into family violence and suicide included a hallucinated reference attributed to Divna Haslam, a University of Queensland associate professor and clinical psychologist, and misstated her team’s research findings.

Haslam, who researches child and family adversity and maltreatment, says the reference was “scary” because it looked so real that someone taking a cursory look could be convinced. Google’s AI summary summarised the reference as if it were a real paper.

Sign up for the Breaking News Australia email

“It’s very frustrating to … to have invested time, money, effort, expertise in rigorous research,” she says. “Then to see something that’s inaccurate and inappropriately attributed anyway is really concerning.”

The author of the submission, Drilldown Reports, told Guardian Australia it used AI in its research process and had identified the errors in its follow-up submission – but the deadline had passed to have them corrected.

“We do strongly believe in the human factor and that it should be a vital part of the full quality review process,” a spokesperson said. “Unfortunately, the human factor also lets us down when uploading the correct document.

“I get that you feel you have a juicy article to add the AI fear machine but this was a human error, not AI.”

Interactive

The chair of the standing committee on social policy and legal affairs, which oversaw the inquiry, said all committees received a range of materials of varying quality and positions: “It is the job of the committee to both accept the evidence offered on face value but then also interrogate it through the inquiry process.”

Haslam says not only does this phenomenon risk devaluing legitimate research but there are serious risks if misleading information makes its way into government policy. “Particularly in a [domestic violence] space,” she says. “We just can’t risk that.”

Dozens of submissions contained errors

The paper was just one of dozens the Guardian identified by building a computer program that extracted references from all inquiry submissions made to the current parliament and checked those references against several online academic databases.

Documents that were flagged as having a high proportion of references not matching anything were then manually checked.

All current commercial services for AI detection have one important limitation – they get it wrong sometimes. A “false positive” detection can lead people to be falsely accused of using AI services such as ChatGPT or Claude.

For this reason, we took a different approach. Analysis of reports produced with AI and studies on AI-generated journal articles showed that LLMs can often produce references that are incorrect in some way, from getting pages, titles and authors wrong to making up details entirely.

We built a custom program that extracts references from documents and then searches those references in CrossRef (a database of academic papers and books) and Google Scholar. We also programmatically checked digital object identifiers (DOIs) if they were present, to see if they resolved to a reachable URL.

Documents that had 20% or more references that were unable to be matched were then checked manually and a large number of submissions with incorrect references were checked with the original authors to verify our method.

It's important to note that this method can only be used to point towards AI usage in documents with references and so will always be an underestimate of actual AI usage.

All links within documents were also checked for ChatGPT metadata tags. These are tags added by ChatGPT when it provides links to users and can persist when copied into a document. ChatGPT tags do not necessarily indicate that someone used ChatGPT directly, as they could be copying the text from a third-party article or similar.

Using these methods, the Guardian found at least 39 submissions to political inquiries containing what appear to be hallucinated references.

These findings represent a conservative estimate of the amount of AI-generated material in political processes, as it only identifies submissions with incorrect citations. This approach would not find documents that used AI to generate text without any references, for example.

In another indication of how often LLMs are being used, more than 100 papers included ChatGPT url tags in reference links, which are automatically added by the platform in its results.

The submissions containing AI-generated material range from documents with a small number of incorrect citations to submissions in which every reference cited does not exist. They were authored by individuals and organisations across the political spectrum.

When Guardian Australia contacted the people and organisations responsible, many were unaware AI could get things wrong in this way or were aware of the potential for errors but had missed them before submitting the document.

How governments will grapple with AI-generated errors in material intended to influence policy, or whether they have the tools and resources to identify them, is an emerging question. Last year the consultancy giant Deloitte issued a partial refund to the federal government after it was revealed that AI tools had included fake references and even a fake court reference in a $440,000 report.

Downie says if public documents like government submissions or even court judgments are found to contain fake material or fake citations, the implications could go beyond bad decisions.

“We’re also starting to undermine the public’s trust and confidence in the types of institutions that underpin our democracy,” he says.

‘A first draft, not a final source’

While misleading claims or invented citations are not a new problem, greater access to LLMs has allowed this to grow at unprecedented levels.

One submission to a housing inequity inquiry this year included what appears to be references to work by Nicole Gurran, an academic at the University of Sydney, that does not exist.

Transparency and contestability are a vital part of the research and peer-review process, which is why citations are used as evidence to back up claims, says Gurran, a professor of urban and regional planning.

“Fake citations, even if an accidental ‘collage’ where the claim is correct but the chain of reference to evidence is broken, undermines that,” she says.

Another paper submitted last year to a Senate inquiry into climate change misinformation by the National Rational Energy Network said the Guardian was “left-leaning, activist-oriented” – and appeared to attribute that claim in part to an “article” by Margaret Simons, a journalist and academic, in a top journal.

But Simons – who is on the board of the Guardian’s owner, the Scott Trust – never wrote such a paper.

The endnote so accurately replicated other references that Simons briefly wondered whether it really existed. “Even though I know I didn’t write this, I did have that moment of self questioning,” she says.

NREN did not respond to a request for comment.

Senate advice for inquiry submissions warns that use of AI can present risks to the quality of information and that accuracy is the responsibility of the submitter.

Downie says the inquiry process needs to remain as open as possible but new guidelines may be needed to “encourage truthfulness”.

“Whether it’s climate change, immigration, health, we want our elected officials to be making decisions based on real information and real evidence not on fake citations and fake claims,” he says.

Guardian Australia asked OpenAI and Google how the companies were working to combat their role in the spread of misinformation.

OpenAI said addressing hallucinations was an ongoing area of research and that users should “use ChatGPT as a first draft, not a final source”, ensuring that they verified quotes, data or references to external documents.

A Google spokesperson said AI Overviews operated “like traditional Search” in that they aimed to match content from the web to the words in the query: “AI Overviews – like traditional blue link results – will find and surface the web pages that include those terms.”

 

Leave a Comment

Required fields are marked *

*

*