cyberivy
AI SummariesHuman MemoryMisinformationAI ResearchChatGPTGeminiAIES 2026Cognitive Bias

Misleading AI summaries nearly halve accurate recall

September 29, 2026

Eine leuchtende Gehirnform liegt auf einer dunklen elektronischen Leiterplatte

In a study of 331 people, correct recall fell from 83.6% to 44.8% after a misleading AI summary. Labeling the summary as AI-generated did not help.

What this is about

A study published on September 28, 2026, by Georgetown University and the University of Washington shows that inaccurate AI summaries do more than misinform. They can also change what people later believe they remember.

A total of 331 participants watched an animated video of a traffic accident. Between 24 and 48 hours later, they read either an accurate or deliberately misleading summary. With an accurate summary, 83.6% correctly remembered the traffic sign in the video. After the misleading version, only 44.8% did.

What the study actually does

The researchers combined two investigations. First, they asked ChatGPT and Gemini to summarize the accident video under different prompts. On average, the outputs omitted 51.6% of central details. In 95% of the summaries, they even left out the most important event: the collision between the car and the pedestrian.

For the memory experiment, participants were randomly assigned to groups. They saw either a stop sign or a yield sign. Later they received a summary in which that detail was presented accurately or inaccurately. Some people were told the text came from AI, while others believed it had a human author. The label did not measurably change the effect.

Why it matters

AI summaries now appear in search results, meeting notes, news apps, and document systems. In medicine, policing, or law, they may become the basis for later statements and decisions. The usual safeguard of having a human review the output may be insufficient if the inaccurate text has already influenced the reviewer's memory.

The study gives the risk a concrete scale. Correct responses fell by 38.8 percentage points. Neither general trust in AI nor an explicit machine-generated label protected participants. This suggests that summaries used in important processes should link directly to the original material and visibly flag disputed details.

The paper is scheduled for presentation at the AIES conference in October 2026. It connects research on generative AI with the decades-old finding that information received after an event can reshape human memory.

In plain language

Imagine packing a suitcase and photographing its contents. The next day, someone gives you a list that wrongly includes a red scarf. When you later recall what was in the suitcase, the list can overwrite the photograph in your memory. An AI summary can similarly come between the original event and later recall.

A practical example

An insurer processes 1,000 claim videos per month and automatically summarizes each one. In 100 cases, the color of a traffic light is decisive. If employees read an inaccurate summary first and the study's measured rates transferred to this setting, about 45 rather than 84 people would remember the detail correctly. The system should therefore show the relevant video segment, support summary claims with timestamps, and never base a decision on the short text alone.

Scope and limits

  • The study used animated videos and one traffic scenario. Real police footage, medical conversations, or workplace meetings may produce different results.
  • The experiment involved 331 people. That supports the specific result but does not establish a universal error rate for every population and situation.
  • Some summaries were assembled and deliberately altered to create controlled conditions. The study therefore does not measure the everyday error rate of commercial summarization services.

SEO & GEO keywords

AI summaries, human memory, misinformation, ChatGPT, Gemini, Georgetown University, University of Washington, AIES 2026, generative AI, memory distortion

πŸ’‘ In plain English

An inaccurate AI summary can influence what people later believe is their own memory. In the experiment, correct recall fell from 83.6% to 44.8%, and an AI warning label did not protect participants.

Key Takeaways

  • β†’The two-stage memory experiment included 331 participants.
  • β†’Misleading summaries reduced correct recall from 83.6% to 44.8%.
  • β†’The tested AI summaries omitted 51.6% of central details on average.
  • β†’Labeling the summary as machine-generated did not prevent memory distortion.
  • β†’High-stakes summaries should always link claims to verifiable original material.

FAQ

Which models were studied?

The researchers analyzed summaries from ChatGPT and Gemini. The controlled memory experiment used text from ChatGPT, with some passages assembled for consistency.

Did an AI-generated label help?

No. Labeling the text as AI-generated did not reduce the measured effect on memory.

Does this apply to real police footage?

That remains unknown. The study used animated accident scenes, and the researchers plan to examine real-world material later.

Sources & Context