0
All projects

MetARAG

A GPU-accelerated document-intelligence platform for CCC Intelligent Solutions that grounds every answer in the exact source paragraph.

Role
Led a team of 5+ engineers
When
Aug 2025 - Dec 2025
Status
Delivered
Context
Client project for CCC Intelligent Solutions
Stack
  • Python
  • LangChain
  • Embeddings
  • GPU inference
  • RAG

The problem

CCC Intelligent Solutions needed reliable answers out of hundreds of PDFs, and an answer nobody can check isn't much use. Every answer had to point back to the paragraph it came from.

I led a team of 5+ engineers on it and designed the pipeline end to end, from parsing through generation.

How it works

  1. 1

    Parsing

    Extracts text and structure from the client's PDFs.

  2. 2

    Chunking

    Splits documents into retrievable pieces sized for the questions people actually ask.

  3. 3

    Metadata enrichment

    Attaches metadata to each chunk so retrieval can filter and rank on more than raw text similarity.

  4. 4

    Embeddings and retrieval

    Embeds the chunks and retrieves the best candidates for each question.

  5. 5

    Source-grounded answers

    Generates an answer and cites the exact paragraph it came from, so every claim can be checked.

  6. 6

    Batched GPU inference

    Runs the model on batches spanning many documents at once, instead of calling it document by document.

Decisions and tradeoffs

  • Batch inference across documents instead of per document. This one change cut document-processing time by about two thirds and inference latency by about a third on a pipeline carrying 100+ GB.
  • Citations at paragraph level. We tuned chunking, retrieval and the citation logic together so each answer resolves to its exact source paragraph.

Results

  • 93%

    Retrieval precision across hundreds of client PDFs

    project evaluation, tuning retrieval and citation logic

  • about two thirds

    Less document-processing time on the 100+ GB pipeline, from batching GPU inference across documents

    project measurements

  • about one third

    Lower inference latency from the same batching change

    project measurements

A note on the code

This was client work, so the code is private and there's no repo to link. For public code that shows the same retrieval and evaluation skills, see ECI Pipeline and rag-redteam.

Next project

Chain-of-Thought on CLEVR

A controlled test of whether chain-of-thought supervision helps a small vision-language model reason about scenes, using BLIP-2 fine-tuned with LoRA.