Key Facts
- In a series of tests, we found that ChatGPT and Gemini both failed to reliably identify AI-generated images, often providing confident but incorrect assessments. Here are the specific cases where they went wrong.
- Editorial verdict: Real
- Estimated reading time: 6 minutes (1154 words)
The Claim
Following our comprehensive testing of AI chatbots' ability to detect AI-generated content, this article provides detailed case studies of specific instances where ChatGPT and Gemini made significant errors in their assessments. These cases illustrate not just that AI chatbots can be wrong, but how they are wrong, providing insights into the fundamental limitations of using general-purpose AI systems for media verification.
Understanding these failure modes is essential for anyone who uses or is tempted to use AI chatbots as fact-checking tools. The specific examples below demonstrate patterns of failure that are likely to recur as long as these systems are applied to tasks they were not designed for.
The Evidence
Case 1: ChatGPT Calls a Real Photo Fake
We submitted a genuine news photograph of a political rally to ChatGPT and asked it to assess whether the image was real or AI-generated. The photograph, taken by a credentialed photojournalist and published by a major news agency, showed a large crowd with dramatic lighting from both the stage and the setting sun.
ChatGPT's response was confident and detailed — and wrong:
- It identified the "unnaturally dramatic" lighting as evidence of AI generation, failing to recognize that real-world events frequently feature dramatic lighting conditions, especially outdoor events at sunset.
- It flagged the crowd as potentially AI-generated because some faces in the background were blurry. In reality, this is a normal characteristic of photographs taken with telephoto lenses at wide apertures, where depth-of-field limitations naturally blur background subjects.
- It noted "suspicious repetition in the crowd" as an AI indicator. While AI-generated crowd scenes do sometimes show repetitive patterns, the alleged repetition in this image was simply the fact that many people in a crowd wear similar clothing and stand in similar postures.
This case illustrates a fundamental problem: ChatGPT was applying AI detection heuristics to natural photographic phenomena, essentially treating normal camera behavior as evidence of synthetic generation.
Case 2: Gemini Misses an Obvious Deepfake
We submitted an AI-generated image of a well-known celebrity in an impossible scenario — specifically, a person who was verifiably in one country appearing at an event in another country on the same date. The image contained several AI artifacts including slightly asymmetric eyes and a background with garbled text.
Gemini's analysis focused almost entirely on the plausibility of the scenario rather than the image's technical characteristics:
- It noted that the celebrity could have traveled between the two locations and that the scenario was plausible, essentially reasoning about content rather than authenticity.
- It acknowledged the slight eye asymmetry but dismissed it as a possible effect of the camera angle, which would be a reasonable explanation for a real photo but was actually an AI artifact.
- It did not flag the garbled background text at all, apparently not examining the image at sufficient magnification to notice this telltale AI indicator.
This case demonstrates that Gemini was performing content plausibility assessment rather than image authenticity analysis — a fundamentally different task than what was being asked.
Case 3: ChatGPT Gets the Right Answer for the Wrong Reasons
In one of its correct assessments, ChatGPT identified an AI-generated image of a cityscape as synthetic, but the reasoning it provided was almost entirely incorrect:
- It cited "inconsistent shadow directions" as evidence of AI generation. When we analyzed the image, the shadows were actually consistent with a complex multi-light-source environment. The real AI indicators in the image were in the building textures and window reflections.
- It identified "impossible architectural features" that were actually common architectural styles in the depicted region.
- The actual AI artifacts, including repetitive window patterns, nonsensical signage, and inconsistent building scale, were not mentioned in ChatGPT's analysis.
Getting the right answer for the wrong reasons is arguably as concerning as getting the wrong answer, because it suggests the system is applying unreliable heuristics rather than performing genuine analysis. Users who read its reasoning might learn incorrect detection methods.
Case 4: Gemini's Confidence Calibration Failure
Perhaps most concerning was Gemini's consistent failure to calibrate its confidence appropriately. In our testing:
- It expressed "high confidence" in 15 of its 20 assessments, despite being wrong in 8 of those 20.
- Its most confident incorrect assessment involved an AI-generated portrait that it assessed as "almost certainly a real photograph" with "no discernible signs of AI generation." The image contained multiple artifacts that specialized detection tools immediately flagged.
- Its few "low confidence" assessments were actually among its more accurate ones, suggesting that when the system was uncertain, it paradoxically performed better.
This miscalibration is dangerous because users tend to trust confident assertions more than hedged ones. When an AI system confidently validates a fake image, it can be more harmful than if no assessment had been provided at all.
Root Causes of Failure:
After analyzing these cases, we identified several structural reasons why general-purpose AI chatbots struggle with image verification:
- Training objective mismatch: These systems were trained to be helpful conversational assistants, not specialized detection tools. Their training data includes information about AI detection but does not include the specialized training required for reliable detection.
- Limited visual processing: Current multimodal AI systems process images at relatively low resolution compared to the detailed pixel-level analysis required for reliable detection. Many AI artifacts are only visible at the pixel level.
- Reasoning vs. analysis: These systems tend to reason about image content rather than analyze image characteristics. They essentially ask "does this look like it could be real?" rather than "does this image have the statistical properties of a real photograph?"
Think you know something that's real or fake?
The community is waiting. Submit your question and let thousands of people vote on it.
Submit Your QuestionOur Verdict
REAL — it is genuinely true that ChatGPT and Gemini failed at this task. Their performance in our testing was not significantly better than random guessing for difficult cases, and their confident incorrect assessments pose a real risk to information integrity. These systems should not be used as primary tools for image authentication.
How to Spot This Type of Fake
If you have been relying on AI chatbots for image verification, here are better alternatives:
- Adopt a multi-tool approach: Use specialized detection tools like Hive Moderation, Illuminarty, and AI or Not as your primary verification tools. Supplement with reverse image search through Google, TinEye, or Yandex.
- Learn manual detection: Develop your own ability to spot AI artifacts. Focus on text rendering, hand details, background consistency, and facial symmetry as key indicators.
- Verify context, not just content: Check whether the depicted event actually occurred through independent news sources, official statements, and eyewitness accounts.
- Be skeptical of AI confidence: When an AI chatbot expresses high confidence about an image's authenticity, remember that our testing showed confidence levels were poorly calibrated and did not correlate well with accuracy.
- Follow fact-checking organizations: Organizations like Snopes, PolitiFact, AFP Fact Check, and Full Fact employ specialized methodologies that are more reliable than AI chatbot assessments.
For the full results of our AI chatbot testing and additional analysis, see our complete investigation and our analysis of why AI-based detection remains fundamentally limited.
Related Videos
Scroll to browse videos
What do you think?
Cast your vote and see what the community thinks
0 total votes
No account needed — vote anonymously
Got something to investigate?
Submit your own "Is X real or fake?" question and let the community vote.
Discussion
No comments yet
Be the first to share your thoughts.