Featured
AI Guardrails Playground
Testing input/output guardrails, content filters, and policy-based controls for safer AI agents.
- Guardrails
- Guardrails AI
- OpenAI
- Updated Aug 28, 2026
- Experiment
- Type
- Playground
- License
- MIT
- Language
- Python
- Updated
- Aug 28, 2026
Overview
An interactive playground for trying guardrail techniques against a library of adversarial prompts, and seeing which ones are blocked, rewritten or allowed — and at what cost to helpfulness.
Key Features
- Prompt-injection and jailbreak detection
- PII detection and redaction
- Policy rules written as plain YAML
- Attack library with pass/fail scoring
Tech Stack
- Guardrails AI
OpenAI
- Policy Engine
- Eval
Project Links
What I Learned
- Layered guardrails beat any single filter.
- Over-blocking is a real cost — measure false positives too.
- Policies are easier to maintain outside the prompt.
Next Steps
- Add streaming output checks
- Integrate with the evaluation toolkit