Why does AI make things up?

·

Levitation Education

A language model makes things up because generating plausible text is the only thing it does. True statements are a subset of plausible statements, and nothing in the mechanism distinguishes the two. When the model tells you the capital of France and when it invents a journal article, it is running exactly the same process with exactly the same confidence.

This is the single most useful idea to give a teenager, because every other failure mode follows from it.

The mechanism, in one paragraph

The model has been trained on an enormous quantity of text and has learned, statistically, which words tend to follow which other words in which contexts. Given a prompt, it produces a likely continuation, one piece at a time. It is not looking anything up. There is no database of facts inside it that it consults and occasionally misreads. The facts it gets right are right because true statements were common in the training text — not because it knows they are true.

Why the invented parts look so good

Because plausibility is precisely what the system optimizes for. A fabricated citation has a real-sounding author, a real journal, a plausible year and a title that fits the subject, because that is what citations look like and the model has seen a million of them. The fake is well-formed for the same reason the real ones are.

This is why "it sounded convincing" is worthless as evidence, and why a learner's instinct — trained on humans, where fluency correlates with knowledge — actively misleads them here.

Where it happens most

Specifics it has seen rarely. Exact numbers, dates, page references, small towns, minor figures, recent events. The thinner the training data, the more the model has to improvise, and improvising is indistinguishable from recalling.

Anything that must exist to satisfy the request. Ask for five sources and you will get five sources. The request creates pressure toward a shaped answer, and the shape is easier to produce than the substance.

Questions with a false premise. Ask about a policy that does not exist and many models will describe it helpfully rather than challenge you.

Its own reasoning. When a model explains why it produced an answer, that explanation is itself generated text, not an introspective report. It can be wrong about itself.

What does not fix it

Telling the model not to make things up does not work, and neither does asking it whether it is sure. Both produce text about confidence, which is not the same as confidence. Search-connected models are genuinely better, because they can cite something that exists — but they still summarize, and a wrong summary of a real page is harder to catch than an invented page.

The honest position is that this is a property of the approach, not a defect awaiting a patch.

Three habits worth teaching

Check the specifics, ignore the prose. Names, numbers, dates and citations are where errors live. The connecting argument is usually fine and usually not the thing being assessed.

Notice when you cannot check. If a learner cannot verify an answer, the correct response is not to trust it — it is to treat the inability as the finding. This is the hardest habit and the most valuable.

Ask it something you know. Before trusting a model in a subject you do not know, test it in one you do. The error rate you find is the error rate you should assume elsewhere.

Why this is an ethics lesson, not a tech lesson

The interesting question is not that machines err. It is what happens when a confident error is handed to someone who cannot check it, and nobody is sure who is responsible for the result. A student who understands the mechanism can predict the failure; a student who understands the consequence can decide what to do about it. Both halves are the job, and the second is the one a course usually skips.

FROM READING TO DOING

World 1 of ZEROTH is free

An evening's work, no email address, no account with us. It is the quickest way to find out whether this is the kind of AI teaching you want in your house or your classroom.

Levitation Education

Published by Levitation Automation.

ZEROTH and the Operator turtle are marks of Levitation Automation. Independent curriculum. Not affiliated with or endorsed by Google or any other tool provider named in the course. Not accredited; awards no credit.

© 2026 Levitation Automation.