What AI image tools get wrong, and why detectors don't fix it

What AI image tools get wrong, and why detectors don't fix it

AI-generated and AI-edited images can contain useful visual clues, but the clues change and detector results depend on the models, edits, and test data involved. Use source and provenance checks instead of treating a score as certification.

What you should be able to do

AI image generation has real, learnable failure patterns - and a real, unsolved detection problem. Learn the patterns so you notice them. Do not rely on a detector score to settle a specific case, because no detector can guarantee it.

Saved only in this browser.
In this article

Somebody shares an image and asks, “is this AI?” The instinct is to look for the classic tells - too many fingers, garbled text, waxy skin - or to run it through one of the “AI detector” tools that have sprung up. Both instincts are reasonable. Both also fail in ways that matter: the visible tells are shrinking with every model release, and the detector tools carry a reliability problem that does not go away just because the interface looks confident.

This article covers what is actually true: the specific patterns where AI image tools still struggle, and why no detector - human or automated - can give you a guaranteed verdict on a single image.

The failure patterns that still hold up

The following categories are heuristics, not measured guarantees. They may justify a closer source check, but tool capabilities change and the absence of a clue says nothing conclusive about origin:

Hands and fine anatomy. Complex hand poses, overlapping fingers, and unusual anatomy still produce visible errors more often than faces or simple poses do, though frontier tools have improved a lot here.

Embedded text. Signage, labels, and text within an image remain uneven across tools and versions - some generations look clean; others still show near-words, inconsistent lettering, or text that is crisp in one area and smeared in another. Treat “looks readable” as weak evidence either way.

Precise counting and exact repetition. Ask a model to depict a specific number of identical objects and the count is often wrong. This applies to both generation and to a model’s own description of an image.

Consistency across multiple images or frames. A single generated image can look flawless; the same character or object generated twice, or across video frames, frequently drifts in small but noticeable ways.

Reflections, shadows, and physically consistent lighting. Generators produce plausible-looking light and shadow more often than physically correct light and shadow.

None of these are reliable enough to build a firm verdict on their own - a skilled or lucky generation can avoid all of them, and a genuine photo can contain an odd shadow or an awkward hand. Treat them as things worth a second look, not proof. They also age quickly: verify against current tools rather than treating this list as permanent.

Why “AI detector” tools cannot give you certainty

NIST’s synthetic-content transparency report reviews detection, watermarking, provenance, and labelling as distinct approaches with research gaps and implementation limitations. NIST’s ongoing GenAI evaluation programme and image challenge evaluate generators and discriminators under defined test conditions. That is the correct frame for a detector result: performance for a specified system, dataset, transformation, and date — not certification of every image from every generator.

Do not use a detector tool’s score as the deciding factor in a consequential decision - a workplace dispute, a relationship conflict, a public accusation. A false positive can wrongly brand a genuine photo as fake; a false negative can wave through a convincing fabrication. Use detector output, if at all, as one weak input alongside everything else in this article - never as the verdict.

What actually helps, in order of usefulness

Provenance information, where it exists. A Content Credential under the C2PA standard can show who signed a file and what creation or editing history they asserted - genuinely stronger evidence than a detector’s guess, though it depends on the tool and platform supporting it and surviving the file’s journey to you. See Content Credentials and watermark basics for what that badge does and does not prove.

Finding the original source. For many images, tracing where the file first appeared, from whom, and when is more useful than inspecting pixels alone. A reverse-image search or direct search for the claimed event may surface original context, but a failed search is not proof of fabrication.

Consistency with what else is known. Does the claimed event show up anywhere else, from an independent source? Does the setting, clothing, or timestamp line up with what is otherwise verifiable? This is slower than a detector button but far more trustworthy.

The failure patterns above, as a supporting signal only. Worth a look, never worth a verdict on their own.

Where this applies beyond a single suspicious photo

This is the same reason synthetic media provenance and consent treats a Content Credential as useful-but-not-sufficient, and it is the direct answer to a common question in the adult deepfake first-response plan: do not delay reporting while trying to “prove” something is synthetic first. If you are generating images yourself and want a practical guide to what current tools do well and where their limits sit day to day, AI image generation 101 covers the tool-selection and prompting side of the same limits described here.

Common pitfalls

  • Treating “no visible tells” as proof of authenticity. Absence of the classic errors listed above means a model avoided them, not that the image is genuine.
  • Treating a detector’s percentage as a fact about the image. It is a fact about how that detector scored that image against its own training, on that date.
  • Spending time “proving” synthetic origin before acting on something urgent. If a decision needs to happen now - reporting harmful content, for instance - do not let detector-chasing delay it. Why AI sometimes gives confident wrong answers covers the same overconfidence problem in a different context; the lesson generalizes.
  • Assuming future tools make the workflow unnecessary. Detection and generation both change. Re-check current, independently evaluated evidence rather than preserving a permanent list of visual tells.

Try it this week

Next time an image makes you pause, resist the urge to run it through a detector app and trust the number. Instead: look for the original source first, check the specific failure patterns above as a secondary signal, and check for a Content Credential if the platform supports it. Use the AI image limits checklist to build the habit of checking the right things, in the right order.

Read next

Continue through the same learning path with the next practical articles.