IamBishalIamBishal
AI Lab

LLM Evaluations

Measuring LLM and agent quality with datasets, metrics and repeatable experiments.

Building evaluation harnesses for relevance, faithfulness, safety and cost, so prompt and model changes are judged by data instead of vibes.

Experiments
3
Open Source
3
Technologies
11
More to Explore
∞

A few highlighted experiments and prototypes.

View All Evaluations

How My Evaluations Work

A typical architecture used in these experiments.

Modular & Extensible
  1. Test Set(Golden answers)
  2. Candidates(Prompts, models)
  3. Run(Batch inference)
  4. Metrics(LLM + heuristics)
  5. Report(Diffs, costs)

Iterate on the weakest cases

All LLM Evaluations

Explore all my Evaluations experiments, from simple prototypes to advanced systems.

3 experiments

Experiment · Learn · Build · Share

Let’s Build the Next Generation of LLM Evaluations

Check out the code, try the demos, or get in touch to collaborate on exciting ideas.