About this document
cs188 Fa23 Note10 by candicekan125 is a document available to read on EtoBox.
This document discusses reinforcement learning, focusing on online planning where agents learn optimal policies through exploration and feedback. It outlines two types of reinforcement learning: model-based learning, which estimates transition and reward functions, and model-free learning, which directly estimates state values. The document further details methods such as direct evaluation, temporal difference learning, and Q-learning, emphasizing their roles in learning optimal policies without requiring p
- Author
- candicekan125
- Language
- EN