Document Retrieval Pipeline
A retrieval and question answering system on AWS, with an evaluation harness that measures whether it actually works instead of trusting how good the answers look.
The problem
Retrieval systems are easy to demo and hard to verify. An answer can sound right while the system quietly retrieves the wrong documents, and an LLM grading another LLM's answers can be confidently wrong.
What I built
A serverless pipeline on AWS that ingests documents, embeds them with Bedrock, stores them in S3 Vectors, and answers questions through an API. On top of it, an evaluation harness scores retrieval against human relevance labels and validates the LLM judge before trusting it. The whole stack deploys as code through GitHub Actions, with no stored AWS keys.
The storyAfter a bulk load, the error queue wasn't empty, yet every document was in the index. Throttled attempts had expired, then a second delivery succeeded. The lesson: confirm a load by counting what's actually stored, not by reading the error queue.
How it works, key decisions and results
How it works
- Documents land in S3 as batched files
- A Lambda splits them into overlapping chunks and embeds them with Bedrock Titan V2
- Vectors are stored in S3 Vectors
- A query endpoint embeds the question, retrieves matching chunks, and generates an answer
- The evaluation harness scores retrieval against SciFact's human labels and compares dense, keyword, and hybrid search
- A Claude model judges answer groundedness, validated against blind hand labels
Key decisions
Results
| System | nDCG@10 | MRR | Recall@10 |
|---|---|---|---|
| Keyword (BM25) | 0.651 | 0.619 | 0.771 |
| Dense (Titan V2) | 0.676 | 0.643 | 0.795 |
| Hybrid (RRF) | 0.706 | 0.659 | 0.878 |
Metric code was cross checked against pytrec_eval, the reference implementation behind published benchmarks.
Skills shown
RAG and retrieval systems, AWS serverless architecture, infrastructure as code, evaluation design, LLM as judge validation