IamBishalIamBishal
Featured

Document Intelligence Lab

Experiments with different RAG pipelines, chunking strategies, and vector databases to find the best setup for real-world use cases.

  • RAG
  • LlamaIndex
  • Qdrant
  • Updated Sep 5, 2026
  • Experiment
Type
Experiment
License
MIT
Language
Python
Updated
Sep 5, 2026

Overview

A test bench for RAG pipelines over messy real-world documents — scanned PDFs, tables and long policies. Each configuration is scored on the same question set so the trade-offs are visible.

Key Features

  • Pluggable chunkers (fixed, recursive, semantic, layout-aware)
  • Swap vector stores: Qdrant, pgvector, Chroma
  • Hybrid search with BM25 and re-ranking
  • Side-by-side answer comparison with scores

Tech Stack

  • LlamaIndex
  • Qdrant
  • OpenAI
  • PDF/Docs

What I Learned

  • Layout-aware chunking beats fixed-size chunks on tables and forms.
  • Re-ranking gives the biggest accuracy gain per unit of effort.
  • Smaller, well-chosen chunks reduce hallucinations.

Next Steps

  • Add multimodal (image + text) retrieval
  • Publish the benchmark results as an article