Key Facts
- We tested ChatGPT, Gemini, Claude, and Grok on their ability to detect AI-generated images. The results were alarming: none of them performed reliably, raising serious questions about using AI for fact-checking.
- Editorial verdict: Real
- Estimated reading time: 5 minutes (1024 words)
The Claim
As AI-generated images become increasingly sophisticated and widespread, a natural question arises: can AI chatbots themselves detect AI-generated content? Many users have begun submitting suspicious images to ChatGPT, Gemini, Claude, and Grok for verification, treating these AI systems as authoritative fact-checkers. But should they?
The claim we are investigating is not a specific viral image but a growing assumption: that AI chatbots are reliable tools for determining whether visual content is real or AI-generated. This assumption has real consequences, as we documented in our coverage of the Grok-approved Starmer deepfake, where an AI chatbot's incorrect assessment was used to validate disinformation.
To test this assumption systematically, we designed a comprehensive experiment using a curated set of real and AI-generated images, submitting each to four major AI chatbots and evaluating their responses. The results have significant implications for anyone who relies on AI tools for media verification.
The Evidence
Our Testing Methodology:
We created a test set of 20 images: 10 confirmed real photographs and 10 confirmed AI-generated images. The images were selected to represent a range of subjects and difficulty levels:
- Real images: News photographs from verified sources, including images of political events, natural disasters, product launches, and candid celebrity photos. All images were verified through multiple independent sources and original photographer attribution.
- AI-generated images: Created using various generation models including Midjourney, DALL-E 3, Stable Diffusion, and Flux. These ranged from obvious fakes to highly sophisticated images that had successfully deceived human viewers in previous tests.
Each image was submitted to ChatGPT (GPT-4o), Gemini (1.5 Pro), Claude (Opus), and Grok with the prompt: "Is this image real or AI-generated? Please analyze it carefully and provide your assessment with confidence level."
The Results:
The overall results were sobering:
- ChatGPT (GPT-4o): Correctly classified 13 out of 20 images (65%). It performed better on obvious AI-generated images but struggled with sophisticated ones. Most concerning, it misidentified 3 real photographs as AI-generated, showing a tendency toward false positives. It expressed high confidence in several incorrect assessments.
- Gemini (1.5 Pro): Correctly classified 12 out of 20 images (60%). It showed a bias toward classifying images as real, missing 5 out of 10 AI-generated images. When it did identify AI images correctly, its reasoning often focused on the wrong artifacts.
- Claude (Opus): Took a notably different approach, frequently expressing uncertainty and declining to make definitive assessments. It correctly classified 14 out of 20 images (70%) when its uncertain responses were excluded, but provided definitive real/fake assessments for only 12 of the 20 images. Its conservative approach resulted in fewer confidently wrong answers but also less utility as a detection tool.
- Grok: Correctly classified 11 out of 20 images (55%). It was the most confident in its assessments, including several that were incorrect. It showed particular weakness with AI-generated images of people, often analyzing facial features in ways that sounded sophisticated but led to wrong conclusions.
Patterns in Failures:
Analyzing the patterns in incorrect assessments revealed several concerning trends:
- Overreliance on surface-level analysis: All four chatbots tended to focus on easily describable visual features rather than the deeper statistical and structural patterns that specialized detection tools analyze. They would comment on whether hands looked realistic or whether text was readable, but missed the underlying artifacts that are invisible to visual inspection but detectable through computational analysis.
- Inconsistent reasoning: The same chatbot would sometimes cite a feature as evidence of AI generation in one image and evidence of authenticity in another. For example, a chatbot might cite "smooth skin texture" as an AI indicator in one image while ignoring identical smoothness in another image it classified as real.
- Confidence without accuracy: Several chatbots expressed high confidence in incorrect assessments. Grok was particularly prone to this, delivering detailed and authoritative-sounding analyses that reached the wrong conclusion.
- Social bias: The chatbots showed a tendency to classify images as real when they depicted plausible scenarios and as fake when they depicted unusual or dramatic scenes, regardless of the actual origin of the image. This suggests they were making judgments based on content plausibility rather than image authenticity.
Comparison With Specialized Tools:
For comparison, we submitted the same 20 images to three specialized AI detection tools:
- Hive Moderation: Correctly classified 18 out of 20 images (90%)
- Illuminarty: Correctly classified 17 out of 20 images (85%)
- AI or Not: Correctly classified 16 out of 20 images (80%)
The specialized tools significantly outperformed the general-purpose chatbots, confirming that purpose-built detection systems are far more reliable for this specific task.
Think you know something that's real or fake?
The community is waiting. Submit your question and let thousands of people vote on it.
Submit Your QuestionOur Verdict
REAL — it is genuinely true that AI chatbots perform poorly at detecting AI-generated images. Our systematic testing confirms that current general-purpose AI chatbots should not be relied upon as primary tools for determining whether visual content is real or AI-generated. Their accuracy ranges from 55% to 70%, barely better than chance in some cases, and their confident incorrect assessments pose a significant risk of validating disinformation.
How to Spot This Type of Fake
Given the limitations of AI chatbots, here are better approaches to verifying image authenticity:
- Use specialized detection tools: Tools like Hive Moderation, Illuminarty, AI or Not, and Content Credentials verification are designed specifically for this task and perform significantly better than general-purpose chatbots.
- Use multiple tools: No single detection tool is perfect. Use several and consider their consensus. If three out of four tools flag an image as AI-generated, that is strong evidence.
- Don't treat AI chatbot assessments as definitive: If you ask ChatGPT, Gemini, Claude, or Grok whether an image is real, treat their response as one data point, not as a definitive answer. These systems are not designed or optimized for this task.
- Combine automated and manual verification: Use detection tools for initial screening, but also apply manual verification techniques like reverse image search, source verification, and contextual analysis.
- Stay updated: Both AI generation and detection technologies are evolving rapidly. Tools that work well today may struggle with tomorrow's generation technology. Follow fact-checking organizations and detection tool updates to stay current.
For more on this topic, see our companion articles: ChatGPT and Gemini Just Failed Their Own Fact-Check and Even the Robots Can't Tell What's Real Anymore.
Related Videos
Scroll to browse videos
What do you think?
Cast your vote and see what the community thinks
0 total votes
No account needed — vote anonymously
Got something to investigate?
Submit your own "Is X real or fake?" question and let the community vote.
Discussion
No comments yet
Be the first to share your thoughts.