A box of photographs, cassette tapes, or scanned documents is raw material rather than a described archive. AI tools can draft transcriptions, captions, or date hypotheses, but speed and accuracy vary by media and tool. The workflow must keep machine output separate from confirmed historical claims.
This is a candidate method in which an approved tool drafts transcription or organisation metadata and a person records consent, provenance, and confidence. The schema can make an unconfirmed date visible, but it cannot prevent someone or another system from later dropping that qualifier.
The three things every archive item needs
For each photo, recording, letter, or document you add to the archive, track three things separately, because conflating them is exactly how family archives drift from documented history into confident folklore:
- Consent: who is in the item, and has each identifiable living person agreed to be included, and under what terms (private family archive only, or shareable more widely)?
- Provenance: where did this item come from — whose collection, roughly when was it acquired or created, and who is the source of any information about it?
- Confidence: is each claim about the item (date, names, location, event) confirmed by a reliable source, or an estimate, and if an estimate, how was it arrived at?
Step 1: Transcribe and describe, with confidence tagged from the start
When you use a model to transcribe an audio recording or draft a description of a photo, ask it to separate what it can read or hear directly from anything it is inferring.
Transcribe this recording [attach or paste the audio/transcript tool's
raw output]. Where the audio is unclear, mark that section as [unclear]
rather than guessing at words. Do not fill in gaps with plausible-
sounding text — an incomplete accurate transcript is more useful to me
than a complete but partly invented one.
Here is a description of what's visible in this photo: [describe what
you can see, or the output of an image-description tool]. Help me
draft a caption using only what's stated. Separately, list anything
you would need external confirmation for — an approximate date based on
clothing style, a location guess, a name guess — clearly labeled as
"estimate, needs confirmation," not stated as fact.
A model asked to identify people, places, or dates in old photographs can produce a plausible-sounding guess with no visible signal that it is guessing, especially for clothing-based date estimates or generic-looking locations. Never enter a model-generated guess as confirmed fact. Record a living relative’s recollection as named, dated testimony with its basis; it is provenance, not independent confirmation by itself. Upgrade a claim only under a defined evidence rule—for example, corroboration by a dated original or independent sources—and preserve conflicting testimony rather than forcing agreement.
Face and voice processing create extra risk beyond captions. Auto-tagging faces in photo apps, uploading recordings for transcription, or “enhancing” old audio can create additional copies that may be reused. The FTC warns that scammers can clone a relative’s voice from audio posted online (FTC consumer guidance). Prefer local tools when practical; if you use a cloud transcription or vision service, check retention and training settings first, keep original media offline, and do not provide a person’s voice sample without a clear need and informed consent.
Step 2: Record consent explicitly, item by item
For any photo, recording, or story involving an identifiable living person, note their consent status before adding the item to any shared or shareable archive — not just for the whole project once, but per item, since comfort levels can differ by photo (a childhood photo they’re fine with, a specific event they’d rather not have archived).
Help me draft a short, plain-language consent question to ask a family
member about a specific set of photos or recordings they appear in —
covering whether they're comfortable being included in a private family
archive, and separately, whether they're comfortable with any of it
being shared more widely (e.g. with extended family, or publicly).
Keep it a genuine question, not a leading one.
Record their actual answer per item or per identifiable group of items, and revisit if someone’s comfort changes — a “yes” given once is not permanent consent for every future use of the archive, especially if the archive’s audience or purpose changes later. When consent is withdrawn, act on it: remove or restrict the item, redact the person from shared views, and note the change in the metadata. “The right not to be included” is operational only if removal has a clear owner and date.
Step 3: Children’s data gets a stricter default
Photos, recordings, and stories involving children need a stricter default than adults: assume restricted access (private family archive only, not shared publicly or with a wider circle) unless the responsible adult has a lawful, child-centred basis to share, and avoid uploading identifiable children’s media to any AI service whose retention and training policies you have not checked. Seek the child’s views in an age-appropriate way and revisit earlier decisions as they mature. UNICEF’s current guidance on AI and children prioritises safety, privacy, transparency, accountability, and children’s best interests. See what ChatGPT remembers, sees, and shares before uploading anything involving a child to a general-purpose tool.
Guardianship, consent, image rights, data protection, and children’s evolving capacity vary by jurisdiction and circumstances. Obtain qualified local review before creating a public, institutional, school, care, or disputed-family archive involving children.
Step 4: Build the metadata schema
Bring transcription, consent, and provenance together into one structured record per item, rather than scattered across captions, memory, and separate consent conversations.
Help me design a metadata table for a family archive with these
columns: Item ID, Type (photo/recording/document), Date (confirmed or
estimate — label clearly which), Source (whose collection, who provided
the information), People identified (with confidence: confirmed/
estimate/unknown), Consent status (per identifiable living person),
and Notes. Use the fields I've described; don't add new metadata
categories without asking.
A model can generate the table structure and help you fill in fields from information you already have; it should not generate the actual historical claims (names, dates, events) from nothing — those come from your family’s sources, confirmed or clearly marked as unconfirmed.
Do not let restoration or colorization quietly become invention
AI photo restoration and colorization tools may make some damage less distracting, and they can also add detail that never existed — an invented pattern on clothing, guessed eye colour, or a changed facial feature. Keep the original scan alongside every derivative and label the derivative clearly as AI-processed rather than more accurate. The US National Archives likewise recommends preserving a master digital copy and backup copies; adapt formats, storage, and retention to your archive. This is the same provenance discipline covered in provenance and watermarking.
Validation and fallback
Define a review cadence that fits archive use and family availability. Revisit items marked “estimate, needs confirmation” when new evidence appears or before publication. Keep the estimate label unless a named source supports the upgrade.
The escalation rule
If a family disagreement arises about what should be included, who should have access, or whether a specific story is even accurate, that is a family decision, not a metadata problem — the schema can hold “disputed” as a confidence status, but it cannot resolve the dispute. If the disagreement involves a deceased relative’s story that conflicts across different family members’ memories, record both versions with their sources rather than picking one to enter as the official account.
Start with one bounded, consented set
Pick one shoebox, one folder, or one recording to start with. Build the metadata table using the family archive metadata schema, get consent from any living people identifiable in it, and keep every date or name that came from a guess clearly marked as one until someone confirms it.



