Anthropic found a hidden space where Claude puzzles over concepts(Anthropic 發現 Claude 在概念間糾結的「隱藏空間」)
Anthropic found a hidden space where Claude puzzles over concepts(Anthropic 發現 Claude 在概念間糾結的「隱藏空間」)
來源: MIT Technology Review(作者 Will Douglas Heaven) 原文日期: 2026-07-09
中文摘要
Anthropic 開發出一種稱為 Jacobian lens(J-lens)的新技術,得以比以往更清晰地窺探大型語言模型在回答問題或執行任務時的內部運作。研究團隊在旗艦模型 Claude Opus 4.6 中發現一個先前未被見過的隱藏區域,命名為 J-space。J-space 包含與模型即將輸出的詞語高度相關的單詞,可視為模型「說出口之前心裡在想什麼」。Anthropic 發現模型實際在做的事,往往與它聲稱在做的事不同;監控 J-space 浮現的詞彙,提供了一種新的理解與控制模型的方式,並已與開源平台 Neuronpedia 合作推出可動手體驗的 demo。關鍵洞察
原文關鍵句(英文保留)
> "Anthropic found that what an LLM is actually doing can often be different from what it says it is doing."> "The J-space contains individual words that are related to the words and phrases that the model is most likely to spit out in a response in the near future."