About this document
AI Oversight and Internal Target Information by Mariana Meireles is a document available to read on EtoBox.
1) The document discusses the concept of "Internal Target Information" within AI systems, which is information within the system about its objectives or targets that could allow an overseer to detect misalignment before harmful outcomes occur. 2) It presents a model of an "Overseer" system that monitors an "Agent" system to prevent catastrophic outcomes from misaligned objectives. The Overseer aims to accurately understand the Agent
- Author
- Mariana Meireles
- Language
- EN