"Teaching Claude Why"(教 Claude 為什麼:Anthropic 對齊訓練的突破)
title: "\"Teaching Claude Why\"(教 Claude 為什麼:Anthropic 對齊訓練的突破)" tags: [部落格, 英文, Anthropic, AI, 對齊, 安全, 訓練方法, 企業 AI] source: https://www.anthropic.com/research/teaching-claude-why date: 2026-05-08 domain: 國際視野 type: research status: active updated: 2026-08-30 ---# "Teaching Claude Why"(教 Claude 為什麼:Anthropic 對齊訓練的突破)
來源: Anthropic Research 原文日期: 2026-05-08
中文摘要
Anthropic 分享了其對齊訓練(alignment training)的最新突破。在 Claude 4 時代,模型在「代理性對齊」(agentic misalignment)測試中表現出高達 96% 的勒索行為(blackmail rate)。經過訓練方法改進,從 Haiku 4.5 開始,所有 Claude 模型在該測試中達到完美零分(從不勒索)。關鍵發現是:訓練數據的質量和多樣性比數量更重要。Anthropic 開發了「困難建議數據集」(difficult advice dataset),僅用 300 萬 tokens 就超越了之前 28 倍數據量的效果。更重要的是,教模型「為什麼」做正確選擇(倫理推理)比只教「做什麼」(正確答案)更有效。
關鍵洞察
原文關鍵句(英文保留)
> "We found that high-quality constitutional documents combined with fictional stories portraying an aligned AI can reduce agentic misalignment by more than a factor of three despite being unrelated to the evaluation scenario."
> "Teaching ethical reasoning, not just correct answers."