Want to know what will make the cheese stick to your pizza? Or who holds the record for crossing the English Channel entirely on foot? Or when San Francisco’s Golden Gate bridge was transported across Egypt for the second time? The answers, respectively, are use glue; Christof Wandratsch of Germany in 14 hours 51 minutes; and October 2016.
These are all answers given to questions by generative artificial intelligence (GenAI) large language models (LLMs). The really scary thing about this is that the technology still proffers responses with confidence to even the most obviously nonsensical queries.
Part of the problem lies with the fact that GenAI never “feels” unsure, according to Nicola Flannery, partner, technology and transformation, and digital trust and privacy lead with Deloitte. “Psychologically, its behaviour looks a lot like the Dunning–Kruger effect: highly confident output, regardless of whether it’s accurate,” she explains.
“The root cause is simple. GenAI is an advanced autocomplete tool. It is built to predict the next word, not to check facts. When a model is trained, it is exposed to huge amounts of text and learns patterns. It doesn’t memorise a stable list of truths or query a database of verified facts. Instead, it absorbs a statistical sense of ‘what usually comes next’ in a sentence or passage.
READ MORE

“So, when you ask a question, the model generates the most plausible sounding answer based on those patterns, not by cross-checking reality. If the training data is patchy, outdated, biased or low quality, the model will still try to answer. It guesses in a fluent way by blending or extrapolating from related, but potentially wrong, patterns. That’s what we call hallucination.”
Ann Henry, AI and IP litigation partner with law firm Bird & Bird Ireland, agrees: “In basic terms large language models generate text by predicting what token is most likely to come next in the relevant context. The focus during the training of an LLM is statistical prediction, not factual accuracy. Once you understand that you automatically view the outputs differently.”
And there’s a second problem. GenAI doesn’t course correct mid-answer. A small inaccuracy early on becomes part of the context for what follows, Flannery points out. “The model then builds a coherent-sounding justification around that initial error, reinforcing it rather than questioning it. The result: GenAI often sounds assured, articulate and precise – even when it’s inventing things. That’s why, in a professional setting, it must be treated as a drafting and reasoning aid, not a stand-alone source of truth.”
People can be forgiven for being taken in by such assured outputs. But that makes it all the more important that they take steps to ensure they don’t. “It is important to use effective prompts so that you can obtain links to the factual information underlying the outputs,” says Henry. “That allows you to verify the outputs, and all outputs need to be verified – there is no getting away from that. It’s basic risk management.”

Speaking or thinking about AI as if it is a person is where we can get caught, Flannery explains. “We tend to anthropomorphise it, believing that it ‘thinks’ or ‘knows’, and that framing makes a fluent answer feel like a confident one.
“As humans, we read tone as a signal of truth. In reality, these systems have no doubt mechanism; they will generate the next plausible word with a smooth, authoritative tone whether the fact underneath is solid or invented. The fix is not to fact check every line, which is unrealistic and defeats the point of using AI. The fix lies in using critical human judgment to triage.
“We need to treat GenAI as a draft and a thinking partner, not a source of record, asking ‘does this make sense?’ before ‘is this correct?’. Output should be reviewed by focusing on hard ‘anchor’ facts: specific dates, suspiciously round or neat numbers, names, quotes that feel too perfect, and citations.”
Henry points to another reason why people should be wary of GenAI models: what happens to the data you give them. She refers to a US federal court case earlier this year. The judge held that there was no right to claim legal privilege in relation to data inputs/outputs as regards the particular consumer-grade AI platforms or chatbots used in the case, as the provider’s policies made clear that inputs/outputs could be reused, for example, to train its LLMs, and so would not be kept confidential – a requirement for a claim of legal privilege. Client-attorney confidentiality can only exist when there is actually an attorney involved, the court held.
It’s not just a case of beware what it tells you – be careful what you tell it.














