Anthropic Unveils Hidden 'J-Space' Inside Claude AI, Revealing Its Unspoken Thoughts
Anthropic has developed a new technique that offers an unprecedented glimpse into the inner workings of large language models, revealing a hidden layer of activity that can differ from what the model actually says. The AI firm created a tool called the Jacobian lens, or J-lens, to probe Claude Opus 4.6, a version of its flagship model released in February. This tool uncovered a space they named J-space, which contains words the model is considering for future responses—essentially, what it might be “thinking” before it speaks.
According to MIT Technology Review AI, the J-lens builds on previous interpretability research by adapting a logit lens to look deeper into the middle layers of the model. While a standard logit lens identifies words the model is likely to output next, the J-lens picks up words that may appear several steps ahead. This reveals concepts the model is processing but might not ultimately include in its final response, offering a clearer picture of its internal reasoning.
Anthropic claims that monitoring J-space provides a new way to understand and control its models, potentially improving safety and transparency. The company has published its findings in a paper and partnered with Neuronpedia to create an interactive demo for the public. Tom McGrath, chief scientist at Goodfire, praised the work as “very good and interesting,” noting that it exposes a deeper level of LLM computation that researchers had not previously seen.