Featured
LLM Evaluation Toolkit
A set of tools and notebooks for evaluating LLM responses using various metrics (relevance, safety, factuality, bias).
- Evaluations
- Developer Tools
- LangSmith
- Ragas
- Updated Aug 20, 2026
- Stable
- Type
- Toolkit
- License
- MIT
- Language
- Python
- Updated
- Aug 20, 2026
Overview
Reusable notebooks and scripts for evaluating prompts, models and RAG pipelines on your own datasets, with reports that make regressions obvious.
Key Features
- Relevance, faithfulness, safety and bias metrics
- Model and prompt A/B comparison
- Cost and latency tracking
- HTML reports for sharing results
Tech Stack
- LangSmith
- Ragas
- DeepEval
- Jupyter
Project Links
What I Learned
- A small, well-labelled golden set is worth more than a large noisy one.
- LLM-as-judge needs its own calibration.
- Track cost alongside quality — the best model isn't always worth it.
Next Steps
- CI integration to block regressions
- Agent trajectory evaluation