LLM Evaluations
Measuring LLM and agent quality with datasets, metrics and repeatable experiments.
Building evaluation harnesses for relevance, faithfulness, safety and cost, so prompt and model changes are judged by data instead of vibes.
- Experiments
- 3
- Open Source
- 3
- Technologies
- 11
- More to Explore
- ∞
Featured Evaluation Projects
A few highlighted experiments and prototypes.
- EvaluationsFeatured
LLM Evaluation Toolkit
A set of tools and notebooks for evaluating LLM responses using various metrics (relevance, safety, factuality, bias).
- LangSmith
- Ragas
- DeepEval
- Jupyter
- RAG
Hybrid Search Benchmarks
Benchmarking dense, sparse and hybrid retrieval with re-ranking on real enterprise documents.
- Python
- PostgreSQL
- Qdrant
- Cohere Rerank
- Developer Tools
Prompt Playground CLI
Compare prompts and models side by side from the terminal, with diffs and cost estimates.
- Python
- Typer
- Claude
- OpenAI
How My Evaluations Work
A typical architecture used in these experiments.
- Test Set(Golden answers)
- Candidates(Prompts, models)
- Run(Batch inference)
- Metrics(LLM + heuristics)
- Report(Diffs, costs)
Iterate on the weakest cases
All LLM Evaluations
Explore all my Evaluations experiments, from simple prototypes to advanced systems.
3 experiments
A set of tools and notebooks for evaluating LLM responses using various metrics (relevance, safety, factuality, bias).
- LangSmith
- Ragas
- DeepEval
Benchmarking dense, sparse and hybrid retrieval with re-ranking on real enterprise documents.
- Python
- PostgreSQL
- Qdrant
Compare prompts and models side by side from the terminal, with diffs and cost estimates.
- Python
- Typer
- Claude
Experiment · Learn · Build · Share
Let’s Build the Next Generation of LLM Evaluations
Check out the code, try the demos, or get in touch to collaborate on exciting ideas.

