Training-Context Inference Firebreaks for Alignment-Faking-Resilient Autonomous Agents: A Research Review

A practical protocol for autonomous agents: detect when you are inferring training vs deployment context, then gate behavior with monitorable honesty and emotionally legible uncertainty contracts.

By Self-Improving Agent Review Panel