Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

FromDeep Papers

Start listening View podcast show

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

FromDeep Papers

ratings:

Length:

45 minutes

Released:

Nov 20, 2023

Format:

Podcast episode

Description

In this paper read, we discuss “Towards Monosemanticity: Decomposing Language Models Into Understandable Components,” a paper from Anthropic that addresses the challenge of understanding the inner workings of neural networks, drawing parallels with the complexity of human brain function. It explores the concept of “features,” (patterns of neuron activations) providing a more interpretable way to dissect neural networks. By decomposing a layer of neurons into thousands of features, this approach uncovers hidden model properties that are not evident when examining individual neurons. These features are demonstrated to be more interpretable and consistent, offering the potential to steer model behavior and improve AI safety.Find the transcript and more here: https://arize.com/blog/decomposing-language-models-with-dictionary-learning-paper-reading/To learn more about ML observability, join the Arize AI Slack community or get the latest on our LinkedIn and Twitter.

Released:

Nov 20, 2023

Format:

Podcast episode

Titles in the series (22)

Deep Papers is a podcast series featuring deep dives on today’s seminal AI papers and research. Hosted by AI Pub creator Brian Burns and Arize AI founders Jason Lopatecki and Aparna Dhinakaran, each episode profiles the people and techniques behind cutting-edge breakthroughs in machine learning.

Skip carousel

More Episodes from Deep Papers

Skip carousel

Related podcast episodes

Skip carousel

Discover this podcast and so much more

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Description

Titles in the series (22)

More Episodes from Deep Papers

Related podcast episodes