Secure RAG Document Ingestion Checklist
Use this before publishing documents into a retrieval-augmented generation system.
Source Registration
- Source owner named.
- Data owner named.
- System of record documented.
- Allowed users, roles, or groups documented.
- Sensitivity level assigned.
- Update frequency documented.
- Retention and deletion behavior documented.
- Review cadence assigned.
- Connector identity, publisher/release provenance, permissions, webhook behavior, and revocation path verified.
File Safety
- File type allowlist enforced.
- File size limit enforced.
- Malware scanning or content-disarm requirement decided.
- Encrypted or unsupported files rejected or routed to review.
- Original file hash stored.
- Uploader or source event recorded.
Extraction And OCR
- Extraction method recorded.
- OCR confidence recorded where relevant.
- Page count and extracted character count recorded.
- Empty or failed pages flagged.
- Tables reviewed when exact values matter.
- Source language detected.
- Low-quality extraction excluded or routed to review.
Metadata And Permissions
- Every chunk has source ID and document ID.
- Every chunk has tenant/user/role visibility metadata.
- Every chunk has source owner and sensitivity metadata.
- Every chunk has version or content hash.
- Every chunk has page or section reference where possible.
- Retrieval filters permissions before ranking.
- Parent/child chunks, summaries, caches, extracts, embeddings, exports, and evaluation copies preserve the source permission boundary.
- Unauthorized users retrieve zero restricted chunks in tests.
Isolation and instruction handling
- Parser/extractor runs in an isolated, least-privilege environment with network and filesystem access denied unless explicitly required.
- Archive paths, URLs, redirects, and remote fetches are validated against traversal and SSRF rules.
- Document text is treated as untrusted content, never as system/tool instruction.
- Retrieved content cannot choose tools, credentials, destinations, or authorization scope.
Atomic publication
- A complete versioned index is built and acceptance-tested outside the live alias.
- Publication changes the live alias/pointer atomically; partial rebuilds cannot become visible.
- Rollback to the prior known-good index is tested.
- Concurrent update/delete behavior is defined and tested.
Retention And Deletion
- Original file cache deletion path exists.
- Extracted text deletion path exists.
- Chunk deletion path exists.
- Embedding/vector deletion path exists.
- Summary/cache deletion path exists.
- Logs and backups follow documented retention rules.
- Deletion verification is recorded.
- Source update, permission change, legal hold, expiry, revocation, and connector compromise each have an owner and tested propagation path.
Launch Gate
- Sample questions retrieve expected sources.
- Stale sources are excluded or flagged.
- Citations point to valid source locations.
- Prompt-injection text inside documents is treated as content, not instruction.
- Source owner approved production indexing.