Low-Stakes Alignment - Listen - AI Safety Fundamentals: Alignment

Low-Stakes Alignment

Listen now

Description

Right now I’m working on finding a good objective to optimize with ML, rather than trying to make sure our models are robustly optimizing that objective. (This is roughly “outer alignment.”) That’s pretty vague, and it’s not obvious whether “find a good objective” is a meaningful goal rather than being inherently confused or sweeping key distinctions under the rug. So I like to focus on a more precise special case of alignment: solve alignment when decisions are “low stakes.” I think this cas...

More Episodes

See all »

AI Safety Fundamentals: Alignment

Published 07/19/24

Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

This paper explains Anthropic’s constitutional AI approach, which is largely an extension on RLHF but with AIs replacing human demonstrators and human evaluators.Everything in this paper is relevant to this week's learning objectives, and we recommend you read it in its entirety. It summarises...

Published 07/19/24

Constitutional AI Harmlessness from AI Feedback

Published 07/19/24