AI can help you navigate a long document, but it should not replace reading material you are responsible for understanding. The safe goal is triage: map the sections, locate relevant passages, and prepare questions while keeping the original document open.
Providing the source narrows the task, but it does not eliminate omissions, OCR errors, invented quotations, context loss, or bad interpretation. Those failures matter most in contracts, leases, policies, regulations, medical records, and financial documents.
A three-prompt workflow covers most everyday long-document tasks.
How to feed a document to AI
Many assistants support file uploads, subject to plan, workspace, format, and size limits. Before uploading, confirm that the tool is approved for the document’s confidentiality and personal-data classification. Scans may require OCR and can lose characters, tables, footnotes, or handwriting.
A few practical notes:
- Long or complex PDFs may exceed extraction or context limits on any tier. Split only on clear section boundaries and keep a coverage ledger so no pages disappear between chunks.
- Scanned PDFs work but lose precision on small or stylised text. For a contract you are about to sign, do not rely on AI OCR alone for the fine print.
- Multiple-file comparison is available in some products and plans. Confirm the current limit, label every file, and require the answer to identify which file and page supports each comparison row.
- Email exports may be technically readable, but they can contain third-party personal data, confidential attachments, hidden quoted text, and incomplete threads. Use an approved tool, minimize the export, and verify that the full thread was captured before asking for a map.
The three-pass workflow
The mistake most people make is asking for “a summary” in one big prompt. You get a generic abstract that misses what you care about. Instead, do three short passes, each with a specific purpose.
Pass 1 — First impressions. What is this thing, and what are the three or four key takeaways?
Pass 2 — Risks and red flags. What in here could hurt me, surprise me, or cost me something?
Pass 3 — Decisions and actions. What do I actually need to do, decide, or ask about?
You may keep the three passes in one conversation for convenience, but do not let an earlier summary become the source for a later pass. Each prompt should direct the model back to the uploaded document, and you should check each result against the original before continuing. Starting a fresh conversation can also be useful when you want an independent pass rather than an answer anchored to the first output.
Pass 1: First impressions
A reliable prompt:
I just uploaded a document. Before going deep, give me:
- One sentence describing what this document is, who wrote it, and who it’s for.
- The three most important takeaways for someone reading it.
- The structure — what are the main sections, in order, and what does each cover?
- Anything that seems unusual or unexpected for this kind of document, given who I am.
About me: I am [your role, your situation, why you’re reading this].
The “about me” line is doing real work. A contract review prompt from a freelance designer should produce different highlights than the same prompt from a corporate procurement officer. The model can calibrate when you tell it.
Read the response carefully. This pass tells you where to begin your own review. It cannot establish that unread sections are safe to skip.
Pass 2: Risks and red flags
This is the highest-value pass for contracts, legal documents, terms of service, employment agreements, leases, and anything where someone is asking you to commit. Continue in the same conversation:
Now read it again with one question in mind: what in here could hurt me or surprise me later?
List, in priority order:
- Any clause that gives the other side a unilateral right — to terminate, change terms, change pricing, claim ownership, demand additional things from me, etc.
- Any obligation on me that is open-ended or hard to comply with.
- Anything that contradicts itself or is ambiguous in a way that would favour the drafter.
- Anything common in this kind of document that is missing (a notable absence).
- Numbers, dates, or terms that seem unusually high, low, or strict.
For each, quote the exact clause and explain in plain English why it matters. Mark anything you are unsure about with [unclear].
The output is a hypothesis list for checking against the original, not a risk assessment.
A model may miss an obvious clause or invent what is “common” for a document type. For a contract or legal document, use this pass only to prepare questions for qualified counsel and the other party.

Pass 3: Decisions and actions
Now the operational pass.
Based on this document, give me:
- The three decisions I need to make before responding or signing.
- The two or three questions I should ask the other side to clarify before committing.
- Any actions on me with specific deadlines or numerical thresholds (dates, amounts, response windows).
- Passages that may require qualified review, with page or section references and the reason each needs checking.
Quote only text present in the file. Give a page or section reference for every item. If you cannot locate a passage, say so rather than reconstructing it.
That last sentence makes the output easier to audit. It does not guarantee correct quotations or references, so open every cited location and compare it with the original before relying on it.
After pass 3, you have a navigation aid: a page-referenced brief and question list. It is not a complete reading, legal opinion, or permission to act.
Beyond contracts: other things this workflow handles
The three-pass structure generalises. Adjust the questions per pass but keep the shape.
Long research reports. Pass 1: takeaways. Pass 2: where is the data weakest, what are the unstated assumptions. Pass 3: what does this imply for my decisions.
Long email threads. Pass 1: who said what, who decided what, what is still open. Pass 2: where did anyone change their position, where is the disagreement hiding. Pass 3: who is waiting on what, who needs to do what next.
Annual reports. Pass 1: the three things that changed year over year. Pass 2: footnotes that contain the real information. Pass 3: what does this imply about the company’s trajectory.
Long meeting transcripts. Pass 1: the decisions and the open questions. Pass 2: anyone who said something noteworthy that did not get picked up by the group. Pass 3: action items by person, with [unclear] tags for ambiguous owners.
Policy documents and regulations. Pass 1: who is covered, what is required, what is prohibited. Pass 2: exceptions, edge cases, conditions where the rules differ. Pass 3: what do I specifically need to do to comply.
In each case, the model is performing a kind of structured reading that humans are bad at when tired and pressured, which is when we do most of our reading.
Three things that lift quality further
Once the workflow is comfortable, three small additions push the output from useful to excellent.
Tell the model what kind of audience the summary is for. “Write the summary as if you were briefing a busy executive who has five minutes.” Or “…briefing a careful lawyer who will check anything they’re not sure about.” Different framings produce different outputs.
Ask for coverage, not confidence scores. Require a list of pages or sections processed, unreadable regions, and every quotation’s location. A self-reported confidence number does not validate the answer.
Open external authorities yourself. If the document cites laws, standards, or precedents, locate the current primary source and check the relevant provision. Search output and model summaries are leads, not authority.
Try it on the next long document
A three-pass triage can make a long document easier to navigate. Its value depends on extraction quality, full coverage, exact source references, confidentiality controls, and your own review.
Try it first on a low-consequence, non-confidential document whose contents you can check. Record the page or section coverage, unreadable regions, and every mismatch you find. For legal or consequential documents, preserve the original text and page references and obtain qualified review before acting.
Sources and review basis
NIST AI 600-1: Generative Artificial Intelligence Profile documents confabulation and information-integrity risks. The NIST Privacy Framework supports identifying and minimizing privacy risk before a document is disclosed to a tool. The Mata v. Avianca sanctions order is a primary court record showing the consequences of relying on fabricated legal authorities.



