Skip to content

Opening book details…

Can I read Detecting Undesirable AI Behaviors on EtoBox?

Detecting Undesirable AI Behaviors by Laura is a document available to read on EtoBox.

What is Detecting Undesirable AI Behaviors about?

This research proposal by Satvik Golechha focuses on detecting undesirable behaviors in neural networks under a near zero-knowledge setting, framing it as an adversarial game between a red team (introducing harmful behaviors) and a blue team (detecting them). The project aims to explore various strategies and challenges faced by both teams, including limited information access and model similarity, while proposing detection techniques such as anomaly detection and feature analysis. The ultimate goal is to e

Author
Laura
Language
EN