Skip to content

Opening book details…

Can I read MIA Vulnerabilities in LLM Alignment on EtoBox?

MIA Vulnerabilities in LLM Alignment by nareshbathala67 is a document available to read on EtoBox.

What is MIA Vulnerabilities in LLM Alignment about?

This paper investigates the vulnerabilities of Large Language Models (LLMs) aligned using Direct Preference Optimization (DPO) and Proximal Policy Optimization (PPO) to Membership Inference Attacks (MIAs). The authors introduce a novel attack framework called PREMIA, which highlights that DPO models are more susceptible to MIAs compared to PPO models due to their tendency to overfit on preference data. The study emphasizes the need for robust privacy-preserving techniques in LLM alignment to safeguard train

Author
nareshbathala67
Language
EN