Proving what is real: provenance, watermarking, and Content Credentials
Intermediate7 min readAI Safety & Data Privacy

Proving what is real: provenance, watermarking, and Content Credentials

What C2PA Content Credentials and watermarks can actually prove, what disappears after screenshots and re-uploads, and how an SME can publish media with an honest provenance policy.

What you should be able to do

Provenance can show who signed a file and what history they asserted. It cannot prove that the depicted event is true, and missing credentials do not prove that media is fake.

AI Expert TeamPublished: Jul 28, 2026
Saved only in this browser.
In this article

“Is this real?” sounds like one question. It is at least four:

  1. Who or what created the file?
  2. Has the file changed since that point?
  3. What tools and edits are declared in its history?
  4. Is the event or claim depicted actually true?

Provenance technology can help with the first three. It does not settle the fourth.

This distinction matters for SMEs choosing how to label AI-assisted marketing, preserve evidence, or respond to suspicious media. Content Credentials, watermarks, metadata, and detection tools are related, but they do different jobs.

A valid credential is evidence about a signed file and its declared history. It is not a fact-check. The absence of a credential is not a verdict either: ordinary processing can remove provenance data from genuine media.

The four layers

Ordinary metadata

Files may contain camera details, timestamps, location, author fields, or editing-software names. This context is useful but often easy to alter or remove. Messaging services, content platforms, and export workflows may strip it.

Metadata alone should not be treated as tamper-evident proof.

Watermarking

A watermark adds a detectable signal to content. It may be:

  • visible, such as a logo or “AI-generated” label;
  • invisible, embedded in pixels or audio;
  • metadata-based, carried in a file field;
  • model-specific, designed to indicate output from a particular generator.

Watermarks can support disclosure or trace a workflow. Their robustness varies. Cropping, compression, transcoding, editing, screenshots, recording playback, or deliberate removal may damage a signal. A detector may also produce false positives or negatives.

Ask a vendor which transformations its watermark survives and how independently the claim can be verified. “Watermarked” without a threat model is not a control.

Content Credentials and C2PA

C2PA — the Coalition for Content Provenance and Authenticity — publishes an open technical standard for Content Credentials. A credential is a cryptographically bound manifest containing signed assertions about an asset: its origin, tools, ingredients, or editing actions.

The current C2PA specifications index listed version 2.4 when rechecked on 28 July 2026. The C2PA technical explainer describes provenance as facts about an asset’s history and Content Credentials as the cryptographically bound structure that records it.

At a practical level, a validator can answer questions such as:

  • Is the manifest structurally valid?
  • Is the signature trusted under the validator’s trust model?
  • Is the credential still bound to this file?
  • What creation or editing actions did the signer assert?
  • Are earlier ingredients linked into the history?

That is materially stronger than an unsigned text field. It is still an assertion by a signer, not a guarantee that the depicted scene is truthful or unstaged, or that a caption is accurate.

Detection

Detection estimates whether content has characteristics associated with generation or manipulation. It may help triage suspicious media, especially when combined with context and forensic analysis.

It cannot establish universal authenticity. Models, edits, formats, and attacks change. Do not build a payment, disciplinary, or public-truth decision around a detector score alone.

What survives common transformations?

The exact answer depends on format, implementation, platform, and whether the workflow uses embedded or recoverable credentials. Test your own publishing chain.

TransformationLikely resultWhat to do
Edit in a credential-aware toolHistory may be extendedVerify the output and inspect declared actions
Export through an unaware toolEmbedded credential may disappearKeep the signed original; test export settings
Social-platform re-encodeMetadata or manifest may be strippedLink to an authoritative original where possible
Screenshot or screen recordingCreates a new file without the original bindingTreat it as a derivative; retain the source
Crop or heavy compressionMay break hard binding or watermark detectionRe-sign the derived asset if your workflow supports it
Copy caption or claim to another fileNo cryptographic continuityVerify against the publisher’s source

C2PA also defines soft-binding approaches intended to help recover provenance for derived assets, but availability and support differ. Do not promise that every screenshot can be traced back automatically.

Run a simple acceptance test before adopting a tool:

  1. create and sign one image;
  2. edit it in your standard application;
  3. export each format you publish;
  4. upload and download it through each platform;
  5. send it through your messaging workflow;
  6. take a screenshot;
  7. validate every result and record what remained.

The result is more useful than a vendor compatibility logo.

What provenance proves — and what it does not

A valid credential may support this statement:

“This exact asset is bound to a manifest signed by an identity trusted by this validator, and the manifest asserts this creation and edit history.”

It does not automatically support:

  • “the camera showed an unstaged event”;
  • “the person consented”;
  • “the caption is accurate”;
  • “every edit was disclosed perfectly”;
  • “the signer is honest”;
  • “the asset is lawful to publish”;
  • “another file without a credential is fake.”

A trusted publisher can sign a misleading image. A real photograph can lose its credential. Provenance adds accountable context; people and institutions still evaluate the claim.

A publishing policy for an SME

Use this as a starting policy and adapt it to your risk.

1. Preserve an authoritative original

Keep the highest-quality source, its credential if present, release approval, usage rights, consent evidence where required, and final published asset. Restrict write access and apply a documented retention period.

Do not rely on the social-platform copy as your archive.

2. Disclose material generation or manipulation

Label content when AI materially generated or changed a realistic person, voice, event, testimonial, or news-like claim. Put a human-readable disclosure near the content even if a machine-readable mark exists.

In the EU, Article 50 transparency rules apply from 2 August 2026. The European Commission’s current transparency overview explains the marking and disclosure duties, while the legally relevant details and exceptions remain in the Regulation and official guidance. Obtain legal advice for your particular use; this article is operational guidance, not a legal opinion.

3. Use provenance where the workflow supports it

Prefer credential-aware capture and editing for:

  • official statements;
  • product and safety evidence;
  • executive or spokesperson media;
  • documentary material;
  • campaign media where AI disclosure matters;
  • assets likely to be impersonated or disputed.

Do not delay a necessary disclosure merely because your tool cannot emit a credential.

4. Verify before republishing

For third-party media:

  • inspect any credential and its validation status;
  • identify the signer;
  • find the earliest authoritative source;
  • compare the claimed date and context;
  • contact the subject or publisher through known details for consequential use;
  • record the editorial decision.

A green credential indicator is one input, not permission to skip verification.

5. State what your label means

Define terms publicly:

  • “AI-assisted” might mean ideation or cleanup;
  • “AI-generated” might mean most pixels, speech, or text were synthesized;
  • “materially manipulated” might mean the realistic meaning or identity changed;
  • “Content Credentials available” should link to the credential or verification route.

Consistency is more valuable than an impressive label nobody understands.

6. Prepare for impersonation

Publish official contact channels. Tell customers how payment-detail changes, urgent requests, or executive instructions are verified. Keep a rapid correction and takedown process for fraudulent media using your identity.

For payment controls, use voice-cloning fraud controls for SMEs. For household scams, use the independent-callback procedure.

A one-page release record

For high-trust media, record:

Asset:
Business owner:
Source/capture:
Rights and consent:
AI generation or material edits:
Human-readable disclosure:
Credential created and validated:
Published destinations:
Authoritative original:
Retention date:
Correction/takedown owner:

This record is valuable even if your current tools do not support C2PA. Provenance starts with a disciplined publishing process.

The honest limit

There is no universal “real” badge. Standards and platform support are still developing. Credentials can be removed, trust lists can differ, signers can make false claims, and genuine media can arrive without provenance.

Use Content Credentials to make your own media more accountable and to gain context about media you receive. Use clear labels to communicate with people. Use independent verification for consequential claims.

Provenance can tell a stronger story about origin. It cannot turn origin into truth.

Read next

Continue through the same learning path with the next practical articles.

Take it further

Hand-picked external courses that go deeper on this topic.

AWS Skill Builder

AWS Security: Securing Generative AI on AWS

AWS Training and Certification

A cloud-vendor-specific complement to the Macquarie specialization: AWS's own Generative AI Security Scoping Matrix, OWASP Top 10 for LLMs, and MITRE ATLAS, walked through governance, legal, and compliance controls for five different AI deployment scopes — from consumer apps to self-trained models. Not GDPR-specific, but a genuinely practical advanced pick for teams whose AI workloads actually run on AWS and need concrete data-governance and compliance controls, not just theory.

Advanced~2 hours · self-paced (9 modules)
Coursera · Macquarie University

Cyber Security: Data, Privacy and AI Security

Macquarie University Cyber Security Hub faculty

The advanced, most explicitly on-target answer to our GDPR × AI gap: a three-course specialization from Macquarie University's Cyber Security Hub that goes from GDPR/CCPA fundamentals and privacy-by-design, through privacy impact assessments, to a dedicated third course on securing AI systems against adversarial attacks and model leakage. Genuinely bridges 'GDPR compliance' and 'AI security' rather than treating them as separate topics.

Advanced~47 hours · self-paced (3-course specialization)
EU Digital Skills & Jobs Platform · CyberSuite

Secure AI Adoption for SMEs: Cybersecurity and the EU AI Act

CyberSuite

The rare AI Act course written for the companies the Act actually reaches: SMEs adopting AI, not the labs building it. Hosted on the European Commission's own skills platform, it pairs the legal side — roles, obligations, risk classification — with the security side (prompt injection, data leakage, supplier due diligence) that most compliance courses skip. For an Estonian SME deploying AI, this is the practical starting point.

Advanced~15 hours · self-paced

See all courses for AI Safety & Data Privacy