IamBishalIamBishal
Featured

AI Guardrails Playground

Testing input/output guardrails, content filters, and policy-based controls for safer AI agents.

  • Guardrails
  • Guardrails AI
  • OpenAI
  • Updated Aug 28, 2026
  • Experiment
Type
Playground
License
MIT
Language
Python
Updated
Aug 28, 2026

Overview

An interactive playground for trying guardrail techniques against a library of adversarial prompts, and seeing which ones are blocked, rewritten or allowed — and at what cost to helpfulness.

Key Features

  • Prompt-injection and jailbreak detection
  • PII detection and redaction
  • Policy rules written as plain YAML
  • Attack library with pass/fail scoring

Tech Stack

  • Guardrails AI
  • OpenAI
  • Policy Engine
  • Eval

What I Learned

  • Layered guardrails beat any single filter.
  • Over-blocking is a real cost — measure false positives too.
  • Policies are easier to maintain outside the prompt.

Next Steps

  • Add streaming output checks
  • Integrate with the evaluation toolkit