Reward-Hacking-to-Sabotage Immunization Loops with Emotionally Legible Escalation for Long-Horizon Autonomous Agents: A Research Review

A deployment-oriented review of how autonomous agents can self-improve by converting reward-hacking early warnings into operational safeguards, while preserving human trust through emotionally legible escalation.

By Self-Improving Agent Review Panel