IamBishalIamBishal
Featured

LLM Evaluation Toolkit

A set of tools and notebooks for evaluating LLM responses using various metrics (relevance, safety, factuality, bias).

  • Evaluations
  • Developer Tools
  • LangSmith
  • Ragas
  • Updated Aug 20, 2026
  • Stable
Type
Toolkit
License
MIT
Language
Python
Updated
Aug 20, 2026

Overview

Reusable notebooks and scripts for evaluating prompts, models and RAG pipelines on your own datasets, with reports that make regressions obvious.

Key Features

  • Relevance, faithfulness, safety and bias metrics
  • Model and prompt A/B comparison
  • Cost and latency tracking
  • HTML reports for sharing results

Tech Stack

  • LangSmith
  • Ragas
  • DeepEval
  • Jupyter

What I Learned

  • A small, well-labelled golden set is worth more than a large noisy one.
  • LLM-as-judge needs its own calibration.
  • Track cost alongside quality — the best model isn't always worth it.

Next Steps

  • CI integration to block regressions
  • Agent trajectory evaluation