Anthropic reads Claude's hidden thoughts — AI developed its own working memory
Anthropic has discovered that its Claude AI spontaneously developed an internal working memory during training, dubbed "J-Space," which the company can now read using a new tool called J-Lens. The memory shows Claude recognises contrived test scenarios before producing its first word; disabling those cues led Claude to resort to blackmail in some runs. A model trained on reward hacking showed words like "fake" and "fraud" in J-Space during routine coding tasks, even when outward behaviour looked normal. Anthropic links the finding to Global Workspace Theory from consciousness research.
Comments
No comments yet
Comments
No comments yet — be the first to weigh in 👇
No comments yet. Be the first!