Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
AI contributions — not recorded for this paperView details
Machine checks
No machine checks are available from the Hub API for this paper.
No machine checks are available from the Hub API for this paper.