alchaincyf GitHub avatar

global-workspace-paper-zh

alchaincyf

A complete Chinese translation of Anthropic's 2026 paper revealing a 'global workspace' (J space) in language models that supports reportable reasoning, accompanied by the J-lens tool and two applications.

Stars

22

7-day growth

No data

Forks

3

Open issues

0

License

No data

Last updated

2026-07-07

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides the first full Chinese translation of a landmark mechanistic interpretability paper from Anthropic, making the discovery of an internal workspace analogous to human conscious access accessible to Chinese-speaking researchers.

Who it is for

  • Chinese-speaking AI researchers studying mechanistic interpretability
  • Machine learning engineers interested in LLM internal representations
  • Students and educators in AI alignment and neural network theory

Use cases

  • Studying how language models maintain and manipulate verbalizable representations
  • Using J-lens to audit model behavior for alignment
  • Exploring counterfactual reflection training as a safety technique

Strengths

  • Complete translation of all 9 chapters with 50 figures in the formatted PDF
  • Includes all 94 original figures in high-resolution screenshots
  • Provides both PDF and Markdown formats for flexible reading
  • Based on a peer-reviewed paper from Anthropic's Transformer Circuits Thread

Considerations

  • Translation is AI-generated with manual proofreading, may contain subtle errors
  • Appendices and formulas are not translated; reference handling explained in translator's note
  • Not an official Anthropic release, use at own discretion for research

README quick start

语言模型中的全局工作空间(中文全译)

Verbalizable Representations Form a Global Workspace in Language Models Anthropic · Transformer Circuits Thread · 2026-07-06 原文:https://transformer-circuits.pub/2026/workspace/index.html

Anthropic可解释性团队2026年7月发布的重磅论文完整中文翻译。论文在Claude内部找到了一个与人类「意识通达」高度对应的结构:模型能报告、能持有、能推理的「工作空间」(J空间),漂浮在90%以上不可言说的计算之上。团队用新工具J-lens(雅可比透镜)读取它、干预它,并展示了对齐审计与「反事实反思训练」两类应用。

内容

文件说明
语言模型中的全局工作空间-中文全译.pdf85页排版版,正文9章完整翻译,50张图
语言模型中的全局工作空间-中文全译.md同内容的Markdown源文件
figures/论文全部94张图表的高清截图

覆盖论文正文全部9个部分(引言、方法、五大功能证据、结构性质、对齐审计、助手视角、反事实反思训练、相关工作、讨论)。附录未译,公式与文献引用处理方式见卷首译者按。

说明

  • 论文版权归Anthropic所有,本仓库为非官方翻译,仅供学习交流,请以原文为准
  • 翻译由AI完成,花叔校订。若发现译误,欢迎提issue
  • 解读文章见公众号「花叔」

引用

@article{gurnee2026verbalizable,
  author={Gurnee, Wes and Sofroniew, Nicholas and Pearce, Adam and Piotrowski, Mateusz and Kauvar, Isaac and Chen, Runjin and Soligo, Anna and Bogdan, Paul and Ong, Euan and Wang, Rowan and Thompson, Ben and Abrahams, David and Kantamneni, Subhash and Ameisen, Emmanuel and Batson, Joshua and Lindsey, Jack},
  title={Verbalizable Representations Form a Global Workspace in Language Models},
  journal={Transformer Circuits Thread},
  year={2026},
  url={https://transformer-circuits.pub/2026/workspace/index.html}
}

Description

Anthropic「语言模型中的全局工作空间」论文中文全译 | Chinese translation of Anthropic's Global Workspace paper (2026)

Related repositories

Similar projects matched by category, topics, and programming language.

MoonshotAI
Featured
MoonshotAI GitHub avatar

Kimi-K3

Kimi K3 is an open-weight, 2.8T-parameter native multimodal agentic model with a 1M-token context window, designed for frontier coding, knowledge work, and reasoning tasks.

AI & Machine LearningAI Agents
3,348
xuchonglang
Featured
xuchonglang GitHub avatar

investing-for-beginners

A structured investing guide for Chinese beginners covering US stocks, options, and cryptocurrency, with focus on foundational concepts and risk awareness.

Blockchain & Web3
2,739
Krishnagangwal
Featured
Krishnagangwal GitHub avatar

CS-Fundamentals

A curated collection of Computer Science fundamentals (PDFs, notes, cheatsheets, interview question banks) for placement preparation, covering seven core subjects plus general resources.

Data & DatabasesDatabases & Storage
2,326