← "The Cognitive Revolution"

エージェントはどう決断するのか:Goodfireのエリック・ビゲロウが語るクリティカルトークン、相転移、文脈内学習

How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning

"The Cognitive Revolution" 2026年10月10日2時間2分
#機械論的解釈可能性#文脈内学習#Chain of Thought#強化学習#LLM#AIアライメント

How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning

"The Cognitive Revolution"

0:002:02:05

要約

Goodfireの研究者エリック・ビゲロウが、LLMが最終的な答えを「決める」仕組みを、forking paths研究を軸に語った回。各トークンで推論を大量に再サンプリングすると、出力分布が特定の1トークンで急激に収束することがある。ビゲロウは、決定の実体はサンプリングにあり、推論は自ら生成した文脈からの文脈内学習だと捉える。推論モデルのCoTの不透明化、RLによる分布の変化、モデル選定、解釈可能性への投資不足にも話が及ぶ。

  • ●forking pathsの研究では、推論の各トークンで約30回ずつ再サンプリングし、最終回答の分布が急変する「クリティカルトークン」を特定した。開き括弧のような一見無意味なトークンで分布が変わる例もあった。
  • ●ビゲロウは、LLMの「決定」はモデル内部というより、最終トークン分布からのサンプリング(コイン投げ)で起こると考える。ただし確率的オウム論は退けるべきだとも述べる。
  • ●文脈内学習は、few-shotに限らず、重みの学習以外でモデルが挙動を適応させる営み全般だと位置づける。推論もその一種だという。
  • ●継続学習の代わりに、文脈エンジニアリング、ハーネス、メモリの改善が有力だという見立てを示した。個人ごとのファインチューニングには懐疑的である。
  • ●推論モデルは多くの可能性を列挙してから選ぶ「トークンのスープ」を生成する。RLは結果だけを報酬にするため、CoTが人間に読めない方向へ変化しうると述べた。
  • ●解釈可能性研究は、モデル開発への投資に比べて桁違いに不足していると指摘。研究用モデルとしては7〜8B以上が適切で、最前線の挙動研究は現状Kimi K3など中国のオープンウェイトに依存していると語った。

章立て

  1. 番組紹介とゲストの背景

    ホストがApolloの研究者との対話を受け、決定トークンの選ばれ方を調べる中でビゲロウに行き着いた経緯を紹介する。

  2. forking pathsの研究

    各トークンで再サンプリングし、最終回答の分布が急変する地点を調べた研究。決定の本質はサンプリングだと説明する。

  3. 文脈内学習の捉え方

    文脈内学習を幅広く定義し、物語を読む際の信念更新や学習の相転移との類似を論じる。

  4. 継続学習とモデルの変遷

    継続学習は解釈手法にどう影響するか。後学習による出力多様性の低下、推論モデルの「トークンのスープ」も議論。

  5. 研究用モデルの選び方

    7〜8B規模が最適という見方、Kimi K3の報酬ハッキング研究の価値、中国製オープンモデルへの依存を語る。

  6. 認知科学とLLM、信念のモデル化

    擬人化の使い方、確率的オウム論への見解、信念をベイズ推論の枠組みで捉える考え方を述べる。

  7. 相転移、不確実性、潜在CoT

    急激な分布変化の位置、UIで不確実性を示す提案、ループ型モデルや潜在的な思考への見通しを語る。

  8. CoTの変質とRL

    CoTや副エージェント間の通信が人間に読めなくなる理由をRLと言語進化の類比で考察する。

  9. 解釈可能性の今後とSilico

    投資不足への警鐘、Silicoの使い方のコツ、Goodfire内の議論テーマ(自動化とアライメント)を紹介する。

解説記事

GoodfireのメンバーであるEric Bigelowが、LLMはどのように最終的な答えを「決める」のかを語った回である。ハーバード大学心理学部で「大規模言語モデルの認知科学」という博士論文を書いた研究者で、ホストのNathan Labenzは、AIに文献調査を頼んだところ彼の名前が繰り返し出てきたと紹介している。

forking paths:1トークンで分布が崩れる

ビゲロウの代表的な研究が、2024年のforking pathsである。質問にChain of Thought(段階的推論)で答えさせ、推論の各トークンで約30回ずつ続きを再サンプリングし、最終回答がどう分布するかを集計した。

すると、最初は答えが50%・30%などに割れていたのに、あるトークンを境に分布が急激に一つの答えへ収束する場面が見つかった。初めて最終回答が言及される箇所のように納得しやすい例がある一方、「kilowatt hours (kWh)」の開き括弧のように、意味の薄そうなトークンで結果が変わる例もあった。彼はこれを、学習中に突然何かを獲得するgrokkingなどの相転移と重ねて語る。

そこから得た見方が、「決定」の実体はモデルの内部というより、トークン分布からのサンプリング、つまりコイン投げにあるというものだ。ここでの「決める」には、本人も引用符を付けたいと述べている。

推論は「自分の出力からの文脈内学習」

ビゲロウは文脈内学習を、few-shotの例示に限らず、重みの更新以外でモデルが挙動を適応させる営み全般と定義する。推論も、サンプリングされたトークンから文脈内学習し、それに整合する続きを生成する過程だという。

物語を読みながら信念が変わっていく様子の研究にも触れた。集計すると滑らかな変化でも、個々の事例を見ると急な変化があるというのがポイントだ。彼は、LLMの解釈は少数の変数を変える静的なテンプレート実験だけでは足りず、動的な変化を見るべきだと主張する。また、同じ入力から大量に再サンプリングして挙動の地形を調べる重要性も強調した。

推論モデルと後学習が変えたもの

現在の推論モデルでは、forking pathsの不確実性の動きは以前より滑らかな傾向があるという。彼の見方では、推論モデルは可能性を次々に列挙し、後でそこから選ぶ。DeepSeek-R1が「wait」を何十回も言う例を挙げ、これは過去のトークンを取り消せない問題への解決策だと推測する。人間的な段階推論というより「トークンのスープ」だという表現である。

また、後学習で出力の多様性が下がるという研究に触れ、RLを重ねた小型モデルでは物語が同じパターンに偏る例を自身の実験から挙げた。一方、基本能力の多くは事前学習済みモデルに既に備わっており、後学習はそれを引き出し、分布を鋭くするものだという考えも示した。

この文脈で、結果だけを報酬にするRLはCoTを人間に読める形に保つ圧力を持たない、と彼は指摘する。副エージェントとの通信が奇妙な言語になる現象も、同じ構図で説明できるという。かつてはCoT監視に期待していたが、この1年ほどで見方が大きく変わったと語っている。

研究対象モデルと擬人化

モデル選びでは、7〜8Bパラメータ以上が解釈可能性研究の適正規模だという。一方、報酬ハッキングのように規模や能力が一定水準を超えて現れる挙動は、Kimi K3のような大型モデルでないと研究しにくいと述べる。そのため米国外のオープンウェイトモデルへの依存は現状避けがたく、オープンモデルの規制は研究の大きな障害になると懸念した。

擬人化については、直感的理論として手軽に使うのは構わないが、「信念」「意思決定」といった語は真剣に定義し直すべきだという立場だ。信念はベイズ推論の枠組みで、潜在的な概念に対する確率の更新として捉える。ただし、モデルが完全な推論経路を内部に表現しているとは考えておらず、数歩先の見通し程度ではないかと推測している。

不確実性をUIに出す

サンプリングの急な分岐が不気味だという問いに、ビゲロウは、多様性と自己整合性を両立させる以上、分岐は避けがたいと答える。必要なのは分岐をなくすことではなく、分岐点をユーザーに見せることだという。彼はトークン確率表示や、分岐を辿るインターフェースLoomを例に、意味レベルで分岐点が強調されるUIを望んでいる。

確率的オウム論は退けるべきだとしつつ、確率性自体は多様性や不確実性の表現に不可欠だと整理した。潜在空間で推論するループ型モデルでも、最終的なトークン生成にはサンプリングが残るという見通しである。

まとめ

本回の中心的な主張は、LLMの意思決定を「内部の意志」ではなく、文脈とサンプリングの相互作用として理解する視点である。これはCoT監視への信頼が揺らぐ中、行動を予測・評価するにはモデルの内部表現を調べる解釈可能性が重要になることを示唆する。ビゲロウは、解釈可能性への投資はモデル開発に比べ圧倒的に少ないと訴える。ただし内容はあくまで本人の見解や推測を含む。実務者にとっては、再サンプリングで挙動の分布を見る、不確実性を可視化するといった視点が、すぐ使える示唆になるだろう。

文字起こし(英語・自動生成)

Hello, and welcome back to The Cognitive Revolution. Today, my guest is Eric Bigelow, member of technical staff at Unicorn Mechanistic Interpretability Startup, Goodfire. The story behind this episode is unique. After recently speaking with Bronson Shane of Apollo Research, who explained that even with access to models' internal chain of thought, it is still often extremely difficult to determine how a model will decide to act, I asked Claude to survey the literature to see what the field as a whole understands about how these critical decision tokens are chosen. One name kept coming up in that thread, and it was Eric. Eric recently completed a PhD in the Harvard Psychology Department with a dissertation titled Toward a Cognitive Science of Large Language Models. And while some of his academic colleagues initially questioned his pivot to focus on LLMs, the fact that even elite academic institutions are now willing to engage with AIs as a sort of mind strikes me as very notable indeed.

We start today with a survey of Eric's work over the last couple of years, beginning with his 2024 paper on forking paths in neural text generation, in which he conducted a massive resampling experiment at every token of a model's chain of thought to identify the critical tokens where uncertainty around the final answer suddenly collapses. As so often with LLMs, some of the findings were intuitive, while others were quite surprising and at times difficult to interpret. Interestingly, Eric says that a system's decisions are actually taken not so much by the model itself, but by the process of sampling from the model's final token distribution, which means that while the stochastic parrot paradigm should definitely be retired, Stochasticity still plays an important role in model behavior. Eric frames much of what we are seeking to understand about LLMs in terms of in-context learning, emphasizing that few-shot prompts are just one clear example of how much runtime context can shape model behavior.

His bet is that this flexibility means that improved harnesses and memory systems will out-compete per user fine-tuning, and thus that our focus of study should remain on major foundation models. Along the way, we also discuss why 7 to 8 billion parameters currently seems to be the interpretability sweet spot, why Kimmy K3 was a step function for studying frontier-level coding and reward hacking, and why, for now at least, interpretability outside of the frontier model companies does meaningfully depend on Chinese open-weights models. We also get Eric's thoughts on why chain of thought is getting so weird including his perspective on the chain of thought dialect and the strange language that Astra uses to communicate with subagents, plus his pro tips for using Goodfire's Silico research agent, and how he hopes future AI product interfaces will begin to resurface model uncertainty to users. The bottom line from this conversation and others I've had recently on the future of AI monitoring and control

is that chain of thought is increasingly unreliable, and that because these issues seem to correlate not just with RL training, but with the power of the raw pre-trained model itself, this trend might be quite difficult to reverse. That in turn means that an awful lot, including our ability to understand why agents take the actions they do, may rest on interpretability. Eric is a pioneer of this research, but as you'll hear, tons of work remains to be done. With that, I hope you enjoy this overview of what we know and how much we have left to discover about how AIs decide what to do. With Eric Bigelow from Goodfire. Eric Bigelow, member of technical staff at Goodfire. Welcome to the Cognitive Revolution. Thank you very much, Nathan. Excited to be here. Yeah, I'm excited for this conversation, too. I think this might be the first time in the history, and there's been like 400 episodes of the podcast at this point, that I went to AIs looking for an answer for a very specific question and came back with a name, and that name was yours.

So the kind of upstream conversation I had of this one was with Bronson Shane at Apollo, who has the dubious distinction, perhaps, of having read maybe more chain of thought reasoning than like any other living human, at least any living human who can talk about it all publicly. And one of the big takeaways I had from that conversation was he said, you look at these chains of thought, and the models are thrashing around doing a sort of linearized tree search of possibility. And then at some point, they make a decision. And even having read all the chain of thought that came before that, it's often like very not obvious why they made the decision that they made. and you really couldn't have predicted it. And in a way, that's kind of human-like, right? How do I make my decisions? I'm not so sure I have great insight into that all the time either. But this prompted the question, okay, at some point, after all this stuff, all this reasoning, all this trying to reason about what the greater wants

and all that stuff, eventually we come to a passage and it's like, and so the answer is, how does that next token get chosen? That's what I went to the AIs to ask. I asked for a literature review. and Eric Bigelow came back as the through line through that research. And I was like, I know that company. I should get to know that guy. So what I want to do, I think you may know as much or more than any other living human about how these decisions are made, and I'm afraid that's maybe not still all that much. But what I hope to get out of this conversation is a little intellectual history of your work and, like, the up-to-date understanding of how the hell do these things make decisions? What do we know about that? And what is our plan to figure out how to understand it better? Because obviously they're making more and more consequential decisions all the time. How does that sound? I'm very honored to be the first guest chosen by the AIs. I always, I feel like they, when I talk to Claude, it seems to really get excited about my research.

But I'm sure everybody feels that way. I think actually having some conversations around this podcast and some of the things you were asking about makes me think that I actually should devote more effort to this exact question of how decisions are made. But I'll give a quick overview of a project that I think is really one of the most informative projects in how I think about how LLMs work at a basic level. and that's my project forking paths and neural text generation and for that basically so this was two years ago with at the time it was one of the gpt 3.5 models and what we do is we take uh reasoning chains we give it a chain of we give it a question and a chain of thought prompt and have it reason through you know step by step and then every single token in its reasoning we resample all these completions. So we resample like 30 different rollouts at every single token.

And then from every one of those rollouts, we see what final answer it landed on. And you can aggregate these into various distributions and then do analysis on these distributions. But I think the most striking part of this is if you arrange it like a multivariate time series, you see these really interesting dynamics that are happening. And I kind of take inspiration from this, from some of the learning dynamics work that was coming out around that time, like around grokking and things like this, where there are like these really sharp phase transitions in training and the model just seems to suddenly learn something. And it turns out that you actually have these during reasoning and in context learning as well, where over the course of reasoning, there will be certain points in reasoning that the model was initially sort of, it had 50% chance of this answer and 30% chance of this answer. And then there's some point that it hits where suddenly the distribution changes really dramatically, and often the distribution just kind of collapses onto a final answer as the model sort of, you might say, decides what it's going to say next.

Although I think I use the word decides here with maybe scare quotes around it because I think there's some nuance to what that means. But one of my big takeaways from this, and I think this also just kind of falls out of thinking about the architecture, but it's nice to have results that really support this too, is that models, even to some extent when they plan out their responses, what they're actually going to say at every single moment is really determined by, like, the flip of a coin, which is what token will it generate next. And there is sort of an interface that was really standard with OpenAI, at least at the time, which I think is unfortunately no longer present, and I haven't really seen standard interfaces, which is where you could see the log probabilities of all the tokens in a text output you got, and you could kind of hover over different tokens and see what are the alternate tokens that could have been generated instead. And if you do this with a lot of text generation, you might realize that you get this one text answer when you ask some question,

but really there actually could have been a lot of different answers, and you might ask what would have happened if this word was different from this word. And what we found is that there are certain key words where sometimes the words, if you were to sort of like look at it like a judge would, would you might guess this like the first time that a final answer is mentioned in the reasoning chain but then there's other times where this distribution changes really dramatically on some sort of arbitrary word even like a open parenthesis for like you know there's one example we have in the paper where it's like kilowatt hours open parentheses kwh and you know you never expect that to really change how the final answer plays out but if you actually go to that exact token with model and you just resample a bunch of text generations, you see that in cases where it's not an open parenthesis, where it starts to be some other just arbitrary word, it ends up being a different final answer. And so I think this thinking about decisions in LLMs or what they're going to say next very much being determined by sort of chance and by whatever tokens are sampled

is really, I think, the foundation for how I think about a lot of things in LLMs. And I think it's sort of shapes how I think about like beliefs and models and sort of what's considered like a factor of hallucination can be dependent on what tokens were generated before it. And I think one of the properties of reasoning that really comes into play here is self-consistency. So if there's a long reasoning chain that is self-consistent and it sort of generates some token earlier in the chain of suppose that this was the year instead of this was the year. If it thinks that the year is 2024 instead of 2026, actually the answer to a question might be very different. And there's cases like this where it's like you might think, well, obviously there's like a ground truth. If I'm asking who is the current head of state, there is a current year, which right now is 2026. And so the model was hallucinating. But I think there's also a lot of use cases with LLMs where we don't really know what the ground truth are. And there's sort of this branching space of possibilities that we actually want to explore naturally when we're interacting with a model.

And, like, for example, if I'm working on a scientific research project and I want to sort of talk to my alum about, like, what should I do next or what do my experiment results mean, I might really want to kind of, like, iterate and understand what's this full, like, space, this whole tree of possibilities that it could have generated. and what would have been different if it had assumed A instead of B. And so to go back to your question, I think the simple answer I have is that I think this decision really happens during sampling. It's my biggest part of how I think about this is that, again, every token that's generated by the model, one of the other big influences in my work is thinking about in-context learning and how models adapt to context. And the way I think about reasoning is that models are basically in-context learning from things that they have generated already. So the words that they generate, there's sort of this stochasticity that's happening. There's these coin flips that are happening.

You could say deciding, but it's really just sampling from a distribution of these alternate tokens. But depending on what token is sampled, it in-context learns, and there's this context dependence where then it'll generate a set of reasoning that is consistent with that. And I think that, like, I'm pretty bullish around, like, McInturk shedding more light, I think, on really answering this question of how decisions are made. But then I think, so I said that this work was done a couple years ago, and I think that modern reasoning models are a little bit different than chain of thought reasoning was then. and some recent projects I've worked on even with like McInturp of reasoning models have kind of changed a little bit how I think about what these models are doing and so I think there's still context dependence but actually as you mentioned reasoning models I think often tend to do this like really like linearized search

where they like enumerate all these different possibilities in context and then at the end they sort of just go back and pick one of the things that they said before So I think this might change exactly how this happens. From what I've seen, the uncertainty dynamics, if you do this like forking path analysis on reasoning models, is often a bit smoother, but there are still sometimes these forking points where if one token is generated different from another or if the model just really doesn't, you know, it enumerates all these reasoning chains and none of them really seem that much more promising than another, then eventually it does sort of, there are cases where it will just say, okay the answer therefore is b c and it actually could have been different depending on what was generated but i think there could still be more work done in in understanding the internal processes that that lead to that that point of choice even and um and another project that i'm on has has shown that there are at least with some some other toy domains there are these cases of like neurons

that seem to track model confidence, where you can even steer and make the model more or less confident generically about whatever concept it's learning. But yeah, I think we really need more interpretability work in this vein, and I think in-context learning and reasoning are unfortunately vastly neglected in the world of interpretability compared to how, I think, how defining they are for how we think about lots of the high-level things that LLMs do. Yeah, I mean, when you talk to the guy who's read more Chain of Thought than anyone, and you keep in mind that a huge amount of the Frontier company's plans right now rests on Chain of Thought monitoring as the way we're going to keep track of what the AIs are doing and if they're doing the right or wrong thing. And then you hear that, yeah, even having read it all and being fluent in these strange internal dialects that they use. I still can't tell you why it chose what it chose. It's all yikes.

Where's the red phone to call the interpretability department? Hey, we'll continue our interview in a moment after I'm aware of my sponsors. Today's episode is sponsored by Parallel, where agents find answers. Most engineers today closely follow new model releases, but don't pay nearly as much attention to their agent's most important tool, web search. If you're like me, your agents are often conducting hundreds of searches per day. And while the unit cost is small, over time, this does start to add up. Before starting to use Parallel, I calculated that my agent search bill would total roughly $500 this year. The good news is that I just had my agents conduct a systematic test, and I found that Parallel's fast mode, which costs just $1 per thousand queries, worked just as well as my previous default provider. and switching to it will save me roughly 80% of my search bill going forward. Parallel's infrastructure is enterprise-grade, and their suite of APIs offers a range of Pareto optimal options

that allow you to choose the right balance of quality, cost, and speed for your needs. So, whether you're building voice agents that need 200-millisecond latency or long-horizon agents that need deep research, adding Parallel to your agent's toolkit is a no-brainer. Get started for free at parallel.ai slash TCR. That's parallel.ai slash TCR. Today's episode is brought to you by Anthropic. By now, you know my story. Claude drafts my intro essays, and I rewrite them. Not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule. Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity.

I feel it most in songwriting. I'm no lyricist, but I'm good with the song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song. These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash TCR. That's claude.ai slash TCR. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai slash TCR.

I think there's a few different directions you started to touch on that I would like to follow up on. Maybe for starters, let's go on this relationship between in-context learning and other things, right? Like there's been a lot of work that's looked at, if not, this is maybe too strong, but sort of an isomorphic relationship between in-context learning and fine-tuning. I think that's, like, especially true at LoRa-scale fine-tuning, where you can localize the weight changes. Maybe it's also true even if you're doing, like, full-weight fine-tuning. I'm less sure about that. How should we understand that relationship? What do you think is the most important that may not be obvious? So I think the most important thing I'd say about in-context learning that may not be obvious for people,

especially if you're just getting started learning about it, is that I think when people say in-context learning historically, it's meant this sort of traditional few shot, give a few shots of input-output examples and see how the model behavior changes. And you can do these, like, scaling curves, for example, where you, like, increase the number of shots that you give and see how model performance improves as a function of this. But then I think the way I think of in-context learning, and I think there's a few other people that have echoed this, and one paper that I like around this is by Andrew Lambin and others on the, I think it's called the broader spectrum of in-context learning, where I think what I think of in-context learning as capturing is everything the model is doing to adapt its behavior that's not happening with in-weights learning. And so this is how it's changing its representations as it reads each new word of input text. And this also applies to reasoning,

where with reasoning, it's generating text and then using that to determine its following behavior. And so I think this view of in-context learning is sort of being really ubiquitous. I think it gives some teeth to thinking about prompting, where there's this term prompt engineering that is sort of a little bit of like a dirty word, I think, in the scientific community, where it feels like people just come up with like a bunch of hacky solutions that make the models go better, you know, tell it you're a brilliant software engineer and it works, and it writes code better, these sort of things. Although these are maybe a bit less prevalent today, because maybe we've kind of, I don't know, fine-tuned some of that distribution shift just into the models. But I think there's a way that it really defies comparison to like this kind of in-weights learning where really prompting is, that's not going to disappear no matter how much fine-tuning you give. You're going to have different contexts that you give the model and you want it to adapt its behavior dynamically, right? Like I think even if you,

I think fine-tune a model enormously, like you still want it when it's like reading, when it's reading Harry Potter for the first time. I mean, of course it's in the training data, but suppose there is another, you know, story it was reading for the first time. You'd want it to be updating its beliefs internally about what's going on in the story. Who are the good guys and the bad guys? Who is happy and who is sad at different times? And really, I have another paper around this sort of belief updating and representation change in stories. And I kind of like this metaphor because I think it's easy to understand that as you read a story, your beliefs should change. There's a bunch of work around analyzing the sentiments of stories. And one of the inspirations from that project was this famous talk by Kurt Vonnegut on the shape of stories where he's describing like these different stories like Cinderella starts. Cinderella's very sad. Everything's horrible. And then she gets invited to the ball and everything's great for a little while. And then the and then the magic sort of falls apart and she has to go back to her normal life, but a bit better than she was before.

And then finally, at the end, when the prince finds her, you know, she has an infinitely happy ending or however he says it. And I think, like, stories have shapes. And really all contexts that we give models also have shapes that are dynamic and temporal. And I think this is something you just, it's hard to really compare it to fine-tuning or to these other ways of changing representations. And I think more broadly, I think that this is a piece of the puzzle to thinking about interpretability of LLMs and not just feedforward networks, is that everything that LLMs are doing is happening dynamically. They're updating their beliefs. They're shifting representations internally. They can learn new representations in context. And this doesn't just happen when you give a few input-output examples. This is happening all the time. This is happening over a conversation. This is happening within sentences. This is happening during the period where a model is reasoning or generating a really long input or output text.

It's sort of changing its representations dynamically to adapt what it's doing in the moment. And yeah, I think a lot of work kind of misses these dynamic changes that are happening. I think a lot of interpretability work sort of focuses on these, like, templates where there's just sort of a couple variables that are changed in the input and behavior is mentioned on, like, one token or one sort of output variable. And really what LLMs are doing is, I think, much more dynamic and adaptive all the time. One other point I want to sort of raise on this view of in-context learning is that something I was inspired by, too, at the time with groffing and this kind of work was this idea that if you look at learning curves over a huge data set, they'll be very, very smooth. And as you just keep training the model and scaling parameters, things will just very, very smoothly increase as you train it for longer and longer,

or errors will decrease, rather. but then if you look at specific problems specific problems can actually change very sharply there actually can be these phase transitions and uh and i think it was it was neil nanda who had this quote of like phase transitions are everywhere if you just looked closely enough if you broke your data set down into like all the into just individual tasks and you looked at individual tasks there would be a lot of stepwise changes where the model really seemed to just suddenly learn something. And I think in-context learning is like that, too, where if you look at aggregate, if you look across a large data set, things might look smooth, maybe deceptively so. But actually, individual cases, like an individual story, have these really nonlinear dynamics. And curves can change very sharply. Things can change after even just one token or one sentence. And other properties can change not just you know suddenly learning something but beliefs can change beliefs

can go up and down and follow all sorts of all sorts of weird twists and turns and so i think another another thing that i think is sometimes missing from views of in-context learning is this you know zooming onto one example sort of problem and um something that i think is also like I'm really curious if the person you had earlier who's looked at all these chains of thought, if he's looked at resampled outputs for the same model or kind of done some of this resampling within a single chain of thought. Because I think sometimes you can get a lot from looking at patterns by just looking at a ton of different rollouts and just trying to understand how the model thinks in some general way. But then I think, like, for an individual prompt, there's sort of this, like, really elaborate landscape just in how the model will think about one problem.

And I think sometimes you can only get this by looking at a lot of different resampled outputs for that same problem and seeing, like, where are the patterns, what kinds of things are similar or different. Like, if you give models, depending on which model you have, if you give them a very open-ended prompt, like generate a story, sometimes you'll notice these really funny patterns where they'll write very different story like things will be superficially very different where it'll kind of start out like very different but then they kind of like I started out that Forking Paths project by just telling me a story about a boy named John and like half the stories involve John having a dog and they go on an adventure and a whole lot of those involve them eventually finding a pond or a waterfall or something. and then sometimes the dog saves the day or sometimes John saves the day. But I think these are all these different patterns you see for outputs for these same input example. And so I think we need to do more of that, like looking at a lot of resampled outputs

for the same input examples to kind of understand how bumpy or smooth the landscape of in-context learning is for these models. Hey, we'll continue our interview in a moment after a word from our sponsors. The Cognitive Revolution is brought to you and powered by Tasklet. I've been running Tasklet agents for almost a year, and they've become a major part of how the show gets made. Before every recording, one agent researches the guest and sends me a suggested outline of questions, complete with links to all sources. Meanwhile, another agent scores the overwhelming number of pitches that I'm getting from PR firms. And that's really the core idea behind Tasklet. Run your business on agents by giving them real jobs to do on an ongoing basis. Setup is incredibly easy. Just describe the work in plain English. Research a new sales lead, update the CRM, and draft a personalized reply. Tasklet connects to your business tools and can run around the clock, on a schedule, or when something happens, like a new email arrives.

Each agent works with your tools and business context and remembers useful information across threads. You can also share agents, connections, and reusable skills with your team, so the workflows that one person figures out become something the whole company can use. No developer needed. You decide what agents can do on their own and what needs your approval, and there are no per-seat fees. You pay for usage, not for adding teammates. Start with one recurring job, like researching the people you're meeting, and then just keep building from there. Visit tasklet.ai and use code COGREV for $50 in free credits. That's code COGREV at tasklet.ai. One big thing I'm pushing toward is continual learning, and I wonder what your experience trying to understand in context learning suggests or implies about how much additional challenge full continual learning might present I could imagine I think the naive read would be like this is going to be really hard now

because we've got everybody's individual model becomes personalized over time and who's to say how different they are. But then maybe the counter argument would be like, we already have that within context learning because there's a lot of context and we basically already deal with that same problem. Does your experience make you feel like we're ready for continual learning in terms of our interpretability techniques or that it will still pose like a step change, increase in difficulty? I think for your question of whether our current interpretability techniques will face challenges with continual learning, I think it really depends on how continual learning plays out. and it's something that I haven't like personally done research on or thought too much about in depth on like do we need different architectures for this and I know there's sort of a lot of people that have been debating this for a while and as you're saying like having personalized models or having models that are like fine-tuned your data I think if that's the case then there's

a possibility that we'll have to have interpretability our interpretability techniques will have to evolve as well and we'll have to have some sort of like like i think i think how how well you know say say that you have a tool that's trained on a base model and then that model is fine-tuned for 10 different people or 10 million different people i think how the real key question there and how well your tools will directly translate is how much representation change there is and like um i have some work on this and and other folks at goodfire have some really cool work on there being these low dimensional manifolds that describe the sort of underlying representation space that models are going through as you're interacting with them and a lot of the representations are sort of defined by this low dimensional conceptual structure. And I think my intuition would be that as long as a lot of that structure remains the same, then a lot of the techniques will be pretty easy to translate. And maybe you need to have a bit

of fine-tuning of your probes so that they work on each person's particular fine-tuned model in that world. But I think the only real challenge there is if the continual learning leads to such catastrophic forgetting that the internal representation structure of the model radically changes. But I kind of doubt that'll happen because I think that will probably mean that the models will break. It may be that there is massive internal representational reorganization, but things are still structured and the model can still do lots of like all the old things it can do, but in new ways. But it seems to me more likely that if we're in a world where really there is this collapse of common internal structure, then the models are probably breaking in that case, unless in this world of continual learning, we're using an ungodly amount of compute for every one of these 10 million people to train them on, you know, do a lot of meta-learning so that they maintain their performance on all these tasks and learn entirely new representational structures. So yeah, I think in that world, it's kind of like having a bunch of

slight tweaks on a main model, and we probably will have tools that we can directly translate to these fine-tuned models and maybe just do a little bit of fine-tuning on the tools that we have and the probes and such. To the question about thinking about continual learning and in-context learning my personal view i mean again i i don't have strong perspectives around their around the architecture argument but i think there's sort of a a bit of a a bitter lesson around like prompting and in context learning and context engineering and harnesses and all that that they might be all you need for quite a lot of stuff and i have a bias towards this because this is what i work on But I kind of, for me, the most likely outcome in any short term is that we don't have custom models that are really trained on very different data. But they're actually, what's happening is it's different context, right? As you were saying, this is sort of what already happens. We have different LLMs that have different memories. And the memory is just, the way I see it, a bunch of data that is tokens that are in context learned every time the model generates an answer.

It sees this long memory. and I think it's just it's incredibly easy to get a lot of a lot of juice out of that without squeezing too hard and then you end up with these memories they're pretty interpretable and yeah I think I think some variables would have to change in the equation for me to see it being really that promising to fine-tune models separately on individual people's data to have customization really pay off in that way. And so maybe my personal best guess is that we're going to see more context engineering and more of this kind of thing, more harnesses, more memories. And I think that's really what we've seen in the last year or so with the rise of agents is a lot of that engineering is just going into putting a lot of stuff in context. Yeah, it's been striking that Anthropica has never had a fine-tuning product, at least that they've offered to the retail developer customer.

And OpenAI has killed theirs. And what little test cases we do have around interpretability applied to different architectures have been, for me, quite pleasant surprises in terms of how well techniques have translated. I'm thinking, for example, of Othello GPT and then Othello Mamba. Both were basically interpretable and seemed to have, like, similar representations internally, and those were recovered through, like, fairly similar techniques. And, yeah, long way, those favorable trends continue. On the question of kind of change in models and, like, how much the beasts themselves have morphed in front of our eyes over the last couple of years, I'm old enough to remember when temperature was a feature that we got to play with at the API level. And at that time, I knew how my tokens were being sampled from the distribution, subject to, like, permanency.

There was a time, I'm also old enough to remember a time when I didn't know about GPU in determinacy, and I was like, why, when I turn my temperature to zero, do I not get the same thing every time? And that turned out to be the answer for that. But I guess how would you describe, because it's true, the temperature has gone away. It sure seems like the models are, you know, mode class is maybe a little bit strong, but, like, they're quite overfit in some ways. Your story example is a good example. that, like, GPT-3, with its wild, untamed nature, was probably in some ways a better bedtime story generator than GPT-6, which, like, always gives you the same kid and dog and waterfall. Even seeing that right now, weirdly, across providers, where, I won't bore you with it, regular listeners have heard me talk many times about how I use models to draft the intro essays for the podcast. But in testing Astro vs. Fable recently, I've also noticed like very like uncannily similar decisions in terms of how they choose to try to present an episode.

And I'm like, this is not random that they're coming up with these very similar ideas. There's some like pretty tight fitting going on. And then we also have seen like the introduction or at least the emergence, maybe it's a better term of these metacognitive behaviors where you see, oh, I think I'm off on the wrong track. Let me go back and approach this another way, the aha moments kind of thing. So how would you describe, how would you, like, tell the story of two years ago when you were doing this, like, super high-volume sampling and you were seeing these, like, pivotal tokens, how have the underlying models changed and what should we be aware of as users and as people trying to understand how decisions get made and as interpretability researchers based on just how dialed in they seem to be. And I'm also just interested, do you even know, like, how sampling is being done these days? It's a very in-the-weeds question.

But, yeah, tell me the story of the last two years. How do you understand how models have changed during that time? I'll say first that I wish I knew what the sampling parameters were for the Cloud API and for the latest GPT models. I think it's very unfortunate that they've talked these away. I expect that it's probably similar things happening. It's just that they want to limit how much people can reverse engineer the models and play with temperature to do a lot of systematic. I think a lot of the people that are doing really systematic prompting are using a lot of hitting the Cloud API with millions and millions and billions even of token requests. there's a high risk that they're doing something, you know, a little bit nefarious or trying to distill models or extract model weights in some way or another. So I think that's probably the main reason why they did this rather than that they're actually doing a different sampling scheme under the surface. Although, I would love to learn more and I'd be happy to be surprised.

As for the story for the last couple of years, So to your point that the earlier models were more like unleashed, as you put it. And I think the word that comes to mind with this is something people have studied around output diversity, where what diversity means is that if you give the model the same prompt and you have it generate a bunch of different outputs, how many different outputs will it generate? And I think this is something that if you... there's been a lot of work showing that this disappears to some extent with supervised fine-tuning, and the more you post-train models, the less diverse their outputs become, which really isn't surprising at some level, because the earlier completion models that weren't even chat, they were just raw completions. When you talk to it, it wouldn't even know that it was in a conversation. It's just doing text completion, next-word text completion on the whole Internet. It felt very much just like this raw thing that you're interacting with, this like big amorphous blob of internet data and if you just start giving it random tokens it'll

just complete those in some way that's coherent whereas afterwards i think we evolved that we first went into a world of supervised fine-tuning and rlhf and other post training for uh chat and basically model instruction tuning right is really fine-tuning models to this distribution of you're going to be talking to a user who's going to give instructions and you should respond to those sort like a chatbot should, and then RLHF I see is further sharpening that distribution around chat. You should act like a chatbot and talk to people in a coherent way that's consistent with this chat interface that everybody is using now. And then I think RL does something pretty different, and I think a couple projects I've worked on have shaped a lot how I've thought about RL, as well as some of these things recently about actually looking at the reasoning of models and seeing how uninterpretable they are and i think rl kind of i don't know how to think about it in terms of it like i think it it sort

of sharpens the distribution in some ways or like there's some really rl smaller open source models that are like kind of cooked like they'll kind of like you know i've i've done some a bunch of experiments recently with this quen um 3.6 35b a3b like model and it was arled pretty heavily and when you give it the uh when you give it this kind of prompt of like give me a story two-thirds of the story are about a clock maker and a bunch of the other ones i forget what it is but there's like another pattern that just comes so much in the stories that it generates but i think like i think i've actually tried this recently with like astra for example and it actually generates like a bunch of different stories which is which is interesting like it will actually generate fairly different stories with fairly different arcs and characters i mean i haven't done like a really massive scale of sampling like hundreds or thousands of different outputs and looking them looking at them i i probably ought to but um the impression i have

of reasoning is that reasoning is it kind of shifts the distribution of what language models are doing in a totally different direction that's very different from like the human text that they were trained on where they they'll going back to this this phrase you said earlier that i really i think very consistent with my thinking that they do like a linearized research that they'll enumerate like all these different possibilities of how the problem could be solved and sometimes the possibilities that they enumerate are related to each other but when there's this idea of wait wait being like a really like key point where the model changes its mind but if you look at like deep seek model out like r1's outputs they'll say wait 50 times or more in a single reasoning chain that's a lot of epiphanies to have when you're working through one problem and i think this was really like this is a solution to this problem that models were hitting at the time which was how can they backtrack, which is related to this forking I mentioned earlier, whereas as it's generating a reasoning chain, it can randomly sample some word, and then it can't

really step back from that path. It's just pushed itself down that path forever. But I think what reasoning models do is they get around that backtracking by just enumerating every possible thing they can think of, and then later going back and choosing from that. And I think my impression of at least probably the most state-of-the-art proprietary models like GPT-6, Astra, why can it generate different stories? It's probably that it just generated a bunch of different possibilities in its reasoning chain, and then it's going back and looking at them. And that's where the sampling is really happening. I think maybe some of the diversity is happening within the reasoning chain itself is my way of thinking about it. But then, I don't know. I should probably do more systematic studies with RL models to understand this problem of diversity better. but I just, I think, yeah, some recent projects have shaped where I really don't, I think reasoning is almost like a misnomer for what reasoning models are doing. I think traditional chain of thought reasoning really was like human reasoning in some way,

where it's like, let's break it down step by step, and here's a set of propositions, A, therefore B, therefore C, therefore D. But reasoning models will just enumerate A, B, C, D, A, E, F, G, F, G, A, all these different it's like a it's almost like a soup of tokens they create and then they just go back and skim out something out of that soup and yeah i think i think that's my high level summary of a lot of the research that's happened in the last couple years and i personally i think that in terms of raw capabilities i think there's been like maybe some gradual change with this and a bit of extra scaling in the last few years but then i think a lot of the raw capabilities were probably in these pre-trained models the whole time. And a lot of what we've done since then is finding ways to just really evoke these and get these out and to sharpen the distribution of what models are doing. Coding agents, for example, is something where I think that just had to train models a lot doing coding to get them to behave in the right way,

even if they had learned the right things from the entire internet of text. It's like this extra version of this additional case of this base models being just next word completion over the whole internet. And I think a lot of the agent fine tuning and getting success with agents is really polishing what distribution they're fit to. You're a coding agent, you're doing lots of changes, you're working with Git and diffs all the time and you're running tests. And I think this is more than anything like a small distribution shift or a sharpening of a distribution. When you do research, how do you choose what model to work with? Obviously, a lot of people are choosing based on what the compute resources they have available support. And I'm sure that's a factor for you, too. But at Goodfire, you guys have significant resources, and you've built some infrastructure to do interpretability on whatever. Kimmy K3, right? Is it a frontier scale? Is it your frontier scale? It's big anyway.

How, I guess, what is your sort of mental model for what kinds of models to choose when? And when you do something on a frontier Chinese model, in your mind, is that essentially synonymous with doing it on a American proprietary model? How much do you think of those as standing in for each other versus how idiosyncratic do you think they might be? So I think of them as, in large part, standing in for each other. There are idiosyncrasies of, like, Quen really likes to think in Chinese, and if you do the Logit lens on Quen, sometimes there will be a lot of Chinese characters sort of in the middle before it starts outputting English tokens if you're talking to it in English. For a question of choosing the model that I work on, it is quite a privilege to have these big compute resources at Goodfire and something that I'm, I think, still getting used to and finding ways to really, like, take advantage of in the best way possible.

But there still is a problem of if we're doing, like, a lot of experiments, you know, I think it still would be hard to justify me taking up hundreds of GPUs for a substantial amount of time to do, like, really, really elaborate experiments on Kimi K3 if they weren't really, like, in line with some of the company's goals or sort of taking advantage of some time where there's a bit of downtime and they're underutilized. So I think there still is. I still do think about compute and kind of having some tradeoff of, you know, what's the right model scale where I can study the behaviors that are interesting enough. So I think generally how I choose is I pick a model like I think there is if you go down too small, then you're working with something that I don't know is is qualitatively different. I think if it's like, you know, in the below a billion parameters, depending on the task, sometimes there are still interesting behaviors you can get out. But then I think there's a lot you can learn by like training models from scratch, where you train like toy neural networks or transformers on a version of your task and you look at

learning dynamics or you do interpretability on them. But a standard that I've always set for myself in my research is to try to study something that's kind of like the big models where you do see these really elaborate human-like behaviors and you can just study them zero shot in the model without having to train the model to do that and without having to use like a facsimile of the model and so um with choosing a model i usually i tend towards like the scale of seven to eight billion or more being a sweet spot where like you can start to do good interpretability research and um i think also like a the few billion models can be good if your task is is like relatively simple but i think sort of around that scale of like 8 billion or more is when you can start to do this sort of like zero shot prompting and get interesting behavior like my uh my story belief updating paper initially we just used uh lame billion and i wasn't totally sure at first how it would do but it actually does you know update belief very dynamically as it's reading stories across like a number of different

emotions and other and other attributes that we tested and so i've been i've been pretty impressed with models abilities at that scale and then um i mentioned that recently studying this quen 35b a3b model and i was also i was very very when i when i started uh before working on the model too much i wanted to just chat with it to get a sense of what it's like and i wasn't sure if it would be kind of like yeah kind of kind of cooked as i said like some of the initial experiences with having it generate stories and seeing like a lot of the same stories but then something that I just like doing with LLMs in general is talking to them about consciousness and see what they say and how the conversation goes and I think that gives me some signal maybe it's like a strong bias here but of how well a model can think about a pretty nuanced and complicated topic and I don't know, I was very impressed with how well this model was able to talk about the problem of consciousness

and LLMs and sort of think actively about that. I mean, there's also this well-documented, or it's well-documented but also a bit speculated around what the post-training data of these models is and if they're really, like, to some extent, distilling some of the proprietary models. And to whatever extent that is true, that could explain why there's a lot of similarity with these models. But I think, yeah, in terms of their core capabilities, They're all really trained on the same data, which is the entire Internet. And the main difference, I think, between all of these models is the post-training and what is that extra special sauce that we're giving the models. And I think in terms of the moat between closed source, I personally think that's probably been the biggest moat and will continue to be is the post-training that they're doing at various levels. And I think that is one of the things that really differentiates the proprietary models from the open source models. And I think what that post-training does is it will make the models way better at coding or specific things that you do care about when you're actually using the models.

But I think a lot of the core capabilities will be very similar. And Kimi K3 is also, I think, this big step function in terms of having a model that's really good at coding that we can study. And although I haven't studied this personally, I've seen some of the other research that people at Goodfire have done recently with reward hacking. And it's, I mean, it's very impressive how good the model is at coding. And also it exhibits this, you know, reward hacking, like crazy behavior that you just don't see at smaller scales. And I think does give us some signal into studying phenomena that we think are happening inside proprietary frontier models. So I think like there's definitely a gap in performance between the proprietary models, the sort of, you know, opening eye on anthropic models and almost everybody else with a few exceptions. But I think, like, in terms of the kind of work we do with interpretability, this doesn't, I think it's not too much of a blocker. The main blocker is when there are certain behaviors

that you just really can't study, except with enough scale or post-training. Like, I think, like, just reward hacking with Kimmy is, like, a really interesting example where it seems like just having the model pass that step function just in terms of some particular behavior coding ability was very important to start also getting to see these really, these dangerous behaviors that we want to study and understand better. That, like, it's really hard to get it out of a smaller model, or if you get it out of a smaller model, it just, it may be that it's happening through a very different means. Does that mean that interpretability research today, outside of proprietary model companies, is becoming dependent on Chinese models? Is there an American model that you could use as a substitute for Kimmy K3 that you feel like wouldn't lose that much of the value? Or are we like we have to use Chinese models if we want to do frontier interpretability in today's world?

I think for now we do. And I think in the biggest models, I think there is a big gap, particularly in programming ability between Kim and K3 and everything else. But I think this isn't a gap that couldn't be closed easily if there was more investment in open source models from American companies. And I think when you go to smaller scales, the gap probably gets a bit smaller. like i think i i do study the quen models pretty heavily and i think they are like really like state-of-the-art for their model scales and i'm and i'm very glad that they exist and i can study them but i also think that if it if i was stopped from using quen on the 35 billion scale and had to study the gemma models instead it probably wouldn't slow me down too much maybe a bit i think right now the gap really exists at the at the top of the food chain but then like yeah i think if there ever was like a ban on using chinese models i would really hope that some american companies

would step up to the plate although i think they'd have to have a particular business proposition in order to make that happen in order to get the money that that they need in order to to train and deliver these models but i would i would hope that would happen i think there is like there have been some loose discussions around like placing restrictions on open source models bar none you know regardless of of where they're coming from and i think that would be the bigger that would be a real problem i think that would really get in the way of a lot of good interpretability research and perhaps like i think some of the most foundational research i think could really continue and we could probably if we well depends what we had access to but if we had access to some open source models i think like there's a lot of foundational interpretability research that could have been studying lms from a few years ago and could keep doing that for like i don't know a long time from now and continue generating like really deep insightful things and helping us understand great mysteries right now but then i think maybe some of the most important behaviors to study

for risk reward hacking or some of the some of these sketchy things that can happen around misalignment that really are emergent behaviors with scale that only come out once you reach some some point of capabilities where it's you're studying the same behavior and not something that's just like it in some way. I think those would be slowed down if there was a ban on open source models. And I think that would, I really hope that doesn't happen. Yeah, I'd hate to see you have to report to some bunker in the Nevada desert to do your interpretability research underground or whatever the scheme would be in that scenario. One thing that I thought was really interesting in looking into your background is you just finished a PhD, not too long ago. Congratulations. And it came from the Harvard Psychology Department. Your thesis toward a cognitive science of large language models.

I'd say one of the biggest surprises for me in recent times has been how productive it has been to anthropomorphize models. What is your kind of personal philosophy of how to anthropomorphize? It seems like we should be backing off of a strong allergic reaction to it, but still it's obviously something we've got to be using very carefully. So how do you use it carefully? So, yeah, I really like that question. I think something that changed a lot over the course of my PhD, I mean, when I started my PhD, LLMs didn't exist. And I mean, my advisor did a little introduction for me with my defense where he said, you know, if you had told me when I was starting out that this is what my thesis was going to be titled, he would say, oh, what now? What are large language models? And for me, yeah, when large language models hit the scene, it really was just a drop everything and study these kind of moment.

I mean, I've had a colleague that recommended Ixplit, like, specifically said that to me, and I think that echoed my sentiment at the time. I did have to, I think, push back earlier on. I did get a lot of pushback from some of the cognitive scientists around me who were used to the deep learning sort of being, there being, like, lots of, like, shortcut-type things where a model can't really do whatever behavior you're talking about, even when it comes to computer vision models, is ImageNet really computer vision? There's a lot of things that vision means that isn't captured by, you know, take this matrix of RGB inputs and give me a one out of a thousand categorical label. And so I think initially there was a lot of pushback on anthropomorphization. I forgot what the metaphor you used was, But I think it was some psychologists and people in related fields like linguistics or philosophy will sometimes look upon upon research on AI critically.

And that it feels like, for example, psychology, like there's a lot of work around AI that feels like it's doing bad psychology or people use these words that have a lot of history in psychology and have been studied pretty deeply. and there's like these really fundamental theoretical and empirical problems that people have thought a lot about and people are just completely missing that when they're using these words and so i think i understand could you give an example of that like where do people go real raw i think like for a while there is resistant on i think there is some resistance on like talking about about planning in language models and do language models really plan things out And even around having beliefs or knowledge of things, it was sort of like they don't really know stuff. They are kind of like maybe they're not quite stochastic parents, but it's sort of something like that. They're just doing a lot of next word prediction, and this kind of being a little bit of a dirty word that like what humans are doing is probably a lot more complicated.

But I think over the years, I think some of this pushback has softened and some of that just has been how damn successful these models have been, right? How everything, you know, it's really taken the world by storm in an incredibly dramatic way. and at some point you can't keep pushing upon pushing back on a river that's just pushing so so hard you know i think something that i really want more people who in these fields who are critical of the research do is instead of just seeing it as problematic when people use these words to come and you know share with share share your own perspective you know start talking to interpretability researchers or whoever who's studying AI and start talking about like what our beliefs, what our decisions and models, how can we understand these better? And so I think broadly, there's two perspectives I have on anthropomorphizing AI, where one of them I think

is that I think you can be like, there's some level at which you can anthropomorphize. It's kind of like, you don't have to be too careful about it. It's for you. It's for you to understand it. You don't have to take the words you're using too seriously. I mean, if my car starts making, you know, a clunky sound as I'm driving, I might say that it like, it's trying to tell me something. I mean, I think this is a little bit dramatic example, because obviously, my car's not trying to tell me something. But it can be helpful to think about things that way. And there's a, there's this word that's used in some parts of cognitive science on intuitive theories. and I think about some of this being like an intuitive theory is these these frameworks we have for thinking that are largely ingrained by like evolution that are higher level than just like really like you know basic perceptual things they're like systems we have for reasoning about phenomena like for example like reasoning about people we have a lot of machinery for and it's very easy for people to intuitively reason about like agents as such and so sometimes it helps me

to anthropomorphize my car as an agent and the problems, the sounds that I'm hearing as it's trying to tell me those. And I think at that level, it's, as long as it's helpful, by all means, do it. And it's just, I think we kind of go into this interesting territory, though, with LLMs where the behaviors they're doing are not just making clunking sounds. They're like, they're doing, they're speaking language. They're doing mathematical reasoning. They're doing social reasoning for that matter. They're helping us answer scientific questions. They're doing programming. So the other perspective I have on anthropomorphization is that I think what's really healthy for everybody, the people from these fields who might have a critical eye of words being misused or wheels being reinvented, is to try to take those questions seriously. what are beliefs or what are decisions for that matter i think um that's a disclaimer by the way i'm not an expert on human decision making but thinking about this project has gotten me a lot

very excited thinking about all these things with people too i mean i think um this this notion of like of forking paths and sort of whether like one token being generated could send you down a very different path it's initially my thought was this is like so different from what people do Like, I have this example in the paper that if a person was speaking and they misspoke, if they said, you know, Billy fell into a hole or Billy fell into a spaceship, depending on what word you say, you might, if you misspoke and you said a hole, I mean, I think I had a better example in the paper, but if you misspoke, you would correct yourself and you would just keep telling whatever story you're going to tell. But for an LLM, it might just commit to that path and start saying this very, very different story. So it seems very different. But then I think there's a lot of things in life that actually are like these forking paths, where actually you do kind of just flip a coin. And depending on what happens at that step, that's what you go with.

And sometimes we don't even see it that way. I mean, there's a bunch of work around rationalization of decisions where, like for example there's these cases where you have people choose from a set of stimuli which one's better in some way and then you later give them the stimuli and you ask them to explain their responses but as the experimenter you might change the stimuli around so you might actually give them different ones than they actually chose and people come up with explanations as why they chose it and so they'll rationalize decisions that they didn't even make and so just to go back to this point I think there's like I think there's a really really fruitful direction which is to try to really take seriously the terms we're using when we are anthropomorphizing and try to study these better try to like i think when people from fields like you know psychology neuroscience linguistics philosophy are critical it's often for some sometimes it's not for a good reason sometimes people are a little bit afraid of you know losing ground on something that is like very important and that they're you don't want people a whole

field to be dominated by low quality work but often it's coming from a place that there's like really important theoretical and empirical problems that people have thought a lot about how to solve and i think the more we can like you know get what we can from this previous work and instead of reinventing wheels start building you know chassis and car frames on top of the wheels we already have i think like that's a future i want to go towards and i think it's just I think a lot of people who are critical in cognitive science, for example, which is more my field of the field I'm most familiar with, I think LLMs are so interesting. I think a lot of cognitive scientists, at least until recently, thought that they were boring. It's stochastic parents or the models are doing something dumb. But I think that this is like a real opportunity to learn something about ourselves and how do we even think about this thorny theoretical problem. if we might say, well, models don't have goals or intentions.

Well, like, what does this mean? I think we take this a bit for granted in our own field because we study people and animals and biological organisms that seem to evidently, by our own framework of defining these terms, have goals and intentions. But I think it's an opportunity to breathe new life into these deep questions. And I think there's people that are really excited about these problems and also a chance to make a really big difference by answering them. So if we look at beliefs, for example, you have looked at a bunch of different, what is a belief is one question we can hold in mind as we talk about some of your techniques. You have done this sampling project to figure out what's the full universe of possibility that an AI might generate in response to a given prompt. If there is a question where the task is to get the right answer,

then it seems like you're making an argument that at these pivotal tokens, the belief state changes. You've also got like low sample techniques with some tricks that allow you to do that more efficiently. And then you've also looked at internal representations and tried to predict what the result is going to be based on those internal representations. What does all that tell you about what is a belief? Do you have an account for what a belief is? Is it simply reduced to the probability that it's ultimately going to give this answer and those shifting probabilities? Is there something more? That doesn't feel like it captures the fullness of what I think when I talk about myself having beliefs. Is there something more that you would talk about when you think about an AI having belief? Yeah, I think that's a great question. I think there's a couple of thoughts that come to mind with that. One of them is that I think there's some level

that it does actually relate to sort of simple things like the probabilities of tokens that a model will generate, except there's more nuance to that. There's a question of like, how does this change across a lot of different inputs or how do these beliefs change through the course of in-context learning or as a model is reading a passage or as it's generating a chain of reasoning. How do its beliefs in different concepts change? And some of the mental framework I have for beliefs is around assigning probabilities according to latent hypotheses or concepts that the model has and then its behavior is generated according to latent concepts and the latent concepts are which concepts are evoked is really determined in large part by the data that's given. And I think this really speaks to a framework of cognitive science that I'm very fond of and has influenced me a lot, which is around Bayesian modeling of cognition, where a lot of cognition can be thought of as Bayesian inference at some level,

where people have these latent concepts and we have prior probabilities of bringing these concepts to bear on sort of regardless of what the context is. And then context or input data will lead us to reweight these, our beliefs in these latent concepts through posterior inference. And by, you know, choosing which concepts or I shouldn't say choosing, but by doing inference and this inference selecting different concepts over one another, this can explain how a lot of how human learning works and how humans can adapt their behavior on the fly to lots and lots of different situations and i think this also applies to lms as well i think something that this uh i think a bigger picture thing that that this question of beliefs also evokes though is like what is what is being represented in an individual like a person or an alum for that matter when it believes something and i think this is room for for i think cognitive science to also grow and and maybe do experiments that we never could with people where the bayesian

framework in cognitive science sometimes has it a criticism of it is that sometimes the the claims that it makes are it's unclear whether these are really saying that people are doing a certain kind of inference in their brain or how this is implemented? What is the algorithm underlying this behavior or this computation? And sometimes these views are like agnostic to that. And it's easy to be agnostic when it's just really, really hard to study brains. And it's really, really hard to uncover what truly are those algorithms or like how are things implemented. But we don't have that excuse with LMS. It's very, very easy to peek inside the brains and to do interventions, which is very complicated and both for methodological and ethical reasons with people but we can do them with LLMs and I think it gives us a chance to actually break new ground around understanding like how beliefs are represented and when models are doing inference are they drawing samples are they estimating posteriors in in some of the ways that we've used to model people and are those

models really what's happening or do they just describe the aggregate behavior as what's happening under the surface really does it defy these models and are these models really just disciplining the behavior instead of the algorithm i think another another point to your question is around um this way of like the the forking plasma methodology of like resampling all of these rollouts at every point in reasoning we use this word in the paper which is uncertainty and i think one of the things i wanted to do with this paper is just highlight how strange it is to think about certainty with reasoning because like by the end of a reasoning chain the model is almost is you like at least back then they're almost 100 sure what the answer is because you're basically whether the if you're proving some mathematical theorem and you're trying to say is it here are these different proof outcomes you could have you end up with a long proof maybe it's wrong but it leads you to one conclusion and so it's the model is not certain at that point or it has no uncertainty about what the final answer is it's determined by the previous reasoning chain

And so what we describe uncertainty as in reasoning is you resample these long reasoning chains. And so say at this point of reasoning I resample and I get 50% answer A and 50% answer B. We say that the model is 50-50 uncertain between A and B. But that doesn't mean that it really represents these full rollouts. That means that if you were to do reasoning from that point on, you would get 50-50. And so something I've been exploring with some recent work is better understanding, like, what do models represent? And it's actually pretty hard to just decode these outcome distributions, these rollout distributions from the hidden weights. A disclaimer is this could be an issue with just not having enough data, and these methods are just so expensive in terms of inference that it is hard to collect a lot of data. So this could be a data issue. but another explanation is that models actually when they're uncertain when they're there they

are sort of uncertain over reasoning that they're not really representing the full reasoning chains they don't really know where they will go next depending on which word they will choose but maybe they represent just like a couple steps ahead and i'm working on a project that's sort of studying exactly this and showing that there are that you actually can decode to some extent what models are going to do next, and you can actually guide sampling based on this. And to some extent, you can guide sampling based on their predictions of their sort of final outcomes. But when it comes to really this question of whether models are representing the full reasoning chain before they really generate it, I kind of don't think so. But I think there's really an open question of how much they do represent. Like when a model could go through one reasoning chain that leads to A and one reasoning chain that leads to B or generate one story or another, does it know the full story that it might generate? I mean, I don't think it knows the full story, but I think a neat connection I make with, or I see there being with theories of, like,

human cognition is around how people can, like, like, there's this stuff around mathematicians and how mathematicians, when they're, like, starting to solve a proof, they, of course, they'll, like, take, it might take them a really, really long time to get a full derivation, but sometimes they'll have some high-level scaffolding view in their mind. They'll be able to see the shape of the proof, even before they complete it. They haven't actually gotten through all these steps that get them from point A to point Z, but they have some high-level representation of what is the shape of that distribution. And I kind of wonder if that might be true for LLMs, too, that they might not represent the full story or the full reasoning trace they're going to do, but they actually have some uncertainty over sort of the high-level paths they might take. In looking at some of the graphs from a couple different papers, including the most recent one, Forking Fast, the sort of sudden phase change is like an extremely striking result.

How do you understand that? First of all, it's weird when it just would happen on a character like an open parenthesis. In general, do we feel like these face changes happen at, and you've got this other paper about the sort of performative nature of at least some chain of thought. Do the face changes tend to happen before or after, or I'm sure it's some mix, But, like, you could imagine a sort of rollout happening, getting to the point of the conclusion of the thought, and then the phase change happening, which would be kind of weird. But it's also quite weird if the phase change happens at the beginning, and then they have to roll out all these tokens already having had the phase change. To compound that confusion even further, I don't feel like I'm doing that. Sometimes I do. Occasionally, I feel like I have, like, a eureka moment. more often though I feel like I'm gradually landing the plane and I do wonder to what degree

we might see quite different things when we go to looped models as we now may have in production with Astra right if you don't have to pick a token at every time stamp then maybe you can gradually go toward a conclusion and maybe that's better something feels weird and like wrong about these, like, sudden token-level phase changes. Maybe not, but I want to see that, like, as cognition is happening, like, change is gradual, and this would feel much better to me. So I guess to make that a question, do you share that intuition that something is, like, weird or wrong with these cliffs? Would you also like to see a gradual landing of the plane? What do we know about how these phase shifts relate to the relevant portions of rollouts, like before, after, or whatever? and should we expect it to be harder, easier, better, worse as we go into thinking in latent space? Just to give you some easy ones. Yeah, I really like these questions.

I'll try to keep them in mind, but please remind me if it slips my mind. I think the question of whether I'm, like, bothered by them, I think a lot of people, when I first describe this, are kind of, like, a little bit unsettled of, like, oh, if one token had been different, actually I'd get a totally different output. and how could we train these away? But I think that if you really think about the problem of output diversity, and if you want some diversity in what your model will do, you want it to do different things, and you want these to be self-consistent, you don't want reasoning chains that are just nonsense, then I think forking might just be an inevitable property of having these two key ingredients. And what I'd really like is I'd really like if our user interfaces surfaced this to us, if we had some level of this kind of like uncertainty, if you call it that, surface to the users. I mean, I mentioned earlier that there were these earlier interfaces of showing you the token log probabilities. And there's another kind of interface called a loom developed by a person named Janus

who did some stuff that influenced this work. And I'd really like to see something like that be more part of the standard user interfaces that we use for interacting with all M's, except at the level of like semantics, I want to see the points where things could branch off and be a very different path highlighted to me as I'm interacting with the all. I think if it's whether it's gradual or it's sharp, if there is this kind of like change that's happening and you could get very different answers by resampling. And maybe like the distribution of different answers is more similar at the beginning. If it's gradual, it's like by the end of it, it actually might be a very different distribution than you had at the beginning. but I'd still, I think I'd really like to see that surface and I think that's how I think about this is like instead of this being necessarily a bad scary thing that can happen where there's like a sharp left turn and you didn't know it I think my personal hope would be that like whether there's a sharp left turn or a

gradual left turn that you know it, that this has actually surfaced to the user in some way for can you remind me your other two questions? Yeah, the kind of mid one was how do the phase changes in kind of rollout time, how do they relate to what would appear to be the relevant text? Yeah, so I think at least for what I've looked at, it really comes in all different, in all these versions. It happens sometimes at the beginning of rollout, sometimes at the end, and sometimes at the middle. at least in that earlier paper with again a model that's now a couple years old there were quite there were more at the beginning and the end of sequence but this also changed depending on task and there were lots of changes in the middle i mean i think i think one big takeaway i have from that is that really every every input behavior can tell its own story there aren't always these dramatic uncertainty dynamics but when there are it's really like it really depends on what input and what reasoning chain you're looking at,

that it really has its own story to tell. And I think if you look under the surface, sometimes there's lots of cases where it is, you can see that there was some step of reasoning that really could have been different and the model was uncertain about this particular step. And sometimes, I think the one thing that could be mitigated and maybe we would hope to train away a bit, is some of this sharp forking happening at unexpected tokens, where I think this, to some extent, speaks to these sort of weird biases of LLMs that they have in text, where maybe that open parenthesis token, there's just some strange bias in all of the human internet, where there's a lot of text that uses kilowatt hours and does parentheses, KWH, that just happens to correlate with something semantic in some funny way. and I think maybe if there were fewer of those it would be much better and if this forking really happened at points that were interpretable for people I think that would be better

but I think to what you said you thought it would be a little bit unsettling if sometimes there were these changes at the end of the sequence and actually you see these quite a lot when a model seems very uncertain about what the right answer is and it just doesn't know. And maybe it's like if you do a bunch of rollouts, there's like one of the answers is more probable. But you see a really common pattern where there'll be like stable uncertainty right up until some point when it just suddenly collapses and it becomes utterly certain. And sometimes that is really when it finally says the answer is blank. And that really is just, it was just really uncertain until that exact point. I mean, one kind of funny example of this from that earlier work was there was a task called last letter that was used for LLMs at the time. I'm sure all the modern state-of-the-art LLMs can do it no problem, but you take four words and you take the last letter of each and you can catenate them.

And the model at the time had a really hard time with this. And some of the resync chains are funny, but it'll even write like Python code to solve this problem. And then it'll finally say, okay, the answer is, and then when you look at the alternate paths it could take, It really is just all these different final answers of letters that are kind of sampled from within these words. And really it is just at that final answer moment just kind of making it up on the fly. And yeah, I think surfacing some kind of understanding of certainty to users is something I really wish we had more of in user interfaces. I mean, I think one of the big problems with this kind of forking, as we kind of tried to get at with this forking fast paper and also another project on this around decoding outcome distributions from hidden representations, is that it's very expensive. It requires, like, a lot of sampling to get real uncertainty over reasoning.

And so I think the forking fast paper, our solution is to use some statistical models of this distribution where we model it as being smooth in some places and jumpy in others, which is supported. If you draw even more samples, it becomes even more smooth, except at these key points where it really suddenly jumps. But really the most efficient way to do this has always been with hidden representations, if we can find some way to do that. And so, yeah, I'm still hopeful that we can have a world where there is really a really efficient uncertainty estimation with hidden representations. and although I think this is even more hopeful because most of the providers really have just converged onto this chat interface that feels, I don't know, it feels very, I think the chat interface is kind of uninspired. It feels like it really hasn't changed from like the ELISA days of like, you know, 50 years ago or more where it's really just text input, text output and I think there's so much more we could be doing with user interfaces.

You know, I think, I really wish we could kind of move into like the GUI world instead of the bash terminal world of interacting with chat and just giving more information to users but in a in a polished and coherent way just the last couple days in the online arena there's been a eulogy or whatever for the stochastic parrot meme and in a way you've you're like saying it's maybe not entirely dead yet i guess one one theory of this sort of of we're generating tokens, our, like, probabilities aren't really changing as we're generating tokens, and then we have a sudden jump at the end of this passage where now, like, sort of the die is cast. One interpretation of that would be, like, you're missing something. There is something more going on under the hood where something is changing, and it's got to be more gradual. You just haven't found it yet. Another interpretation would be, like, no, it really just remains uncertain until the end and then like we make a random choice of a particular token and that is like what actually

leads to the difference in behavior can it be both the different cases and does that leave some should we still have some space for the stochastic parrot i don't know i i think i i i think the stochastic parrot metaphor should probably be put to rest i i i used it a bit earlier as this criticism coming from from earlier in my phd from cognitive scientists and psychologists on feeling like lms weren't interesting to study but i think it's it's really sort of a fairly shallow argument that doesn't really say very much i mean i think it's sort of like a a chinese room type argument where like there's a giant book that has memorized examples but i think it's like what the heck is going on in that book that lets you respond with conversation Like, whatever these parrots are doing, it's very, very interesting. And however it is, but it's like, I don't really think the idea that it's just like a giant lookup table holds much ground. And I think there's also just so much work showing, like, consistent world models.

I mean, you mentioned all fellow GPT, but there's so much work around, like, looking at the representations in LLMs and showing that these are, like, really structured world models. And at some point, it's more efficient to learn a world model than to learn a giant lookup table. But then I do think stochasticity and sampling is still such an integral part of the equation. And it kind of has to be if we want some, if we want there to be like output diversity, if we want models to be able to do different things, if we want them to also have like uncertainty. And like maybe there's a world where they just voice that uncertainty. but then if we don't want them to just halt their reasoning every time there isn't 100% certainty then there has to be a bit of this kind of I think flip of the coin and following a path. I think the more that can be surfaced to the user and the more the user can know how uncertain was the model about this thing that it's reported as a fact when it may have been a hallucination the better it is if people know that.

But yeah I think stochasticity is definitely not dead. and uh i was earlier you mentioned like this idea of loop transformers and i think um loop transformers in this world of like latent chain of thought that we might be going into with with gpt astra really seems like pushing the needle on that in a big way and making especially in terms of efficiency and not just performance um i think this really is something that i'm spending a lot of time now like trying to just wrap my head around and think about this in relation to my mental model for how LLMs work. Because now reasoning can happen during this deterministic forward pass in some sense. I mean, maybe it's looped and maybe it has some amount of depth, but in some sense it's really just deterministic. It's like all of the weights just do the computation. But something I've recently been doing, some like forking paths analysis on the uh the open source loop transformers that i can find like

there's the series of models called oro and you actually do have some interesting uncertainty dynamics as the models are generating their tokens in their in their reasoning chain after doing this latent reasoning it still is uncertain between answers and it still does change at particular tokens and so maybe the mental model shifts a little bit in terms of like all of the how much of the thinking is just happening in that token stream but then if the model has some uncertainty and if it like in its latent chain of thought explores the different possibilities maybe it still happens that as it generating a response it still is doing this flip of the coin next word prediction and i don know maybe i i think maybe that a good thing i i it certainly helps for my mental model but i think this like i i don't want to live in a world where you always get exactly the same output from a model if you give exactly the same input that feels like if we're in the world where that's happening then there is no diversity

of perspective and we've just like really that's like the most extreme version of mode collapse you know where like your one output will only give you your one input will only give you the one output and I I don't think latent chain of thought will completely change that because you still eventually do go into this sampling tokens regime but yeah it might change the it might change that in in some pretty dramatic And so, yeah, I'm still trying to like trying to wrap my head around that and trying to think about like what's going on inside the latent chain of thought. And how do we how do we think about that with some of the some of the tools that we have and even just metaphors that we as scientists use for understanding the models? Do you have intuitions for why the chain of thought is getting wonky? It's a quite different dialect or way of processing information clearly that goes on in the chain of thought with all this like vantage, illusions, disclaim, whatever.

We also see this now with Astra and the way that it sends messages to its sub-agents, which I had naively thought has to be some sort of like length penalty or brevity reward or something. but then it seems like maybe not because people are like counting tokens and saying you could just say the thing like in fewer tokens if you just said it so it remains a mystery as to and I don't know if OpenAI is a better explanation but they certainly have been saying we're looking into this it's important we want to figure out why this is happening any intuitions that you would offer that people might want to take inspiration from chase down put into the hopper at OpenAI for what the hell is going on there Yeah, I think that reasoning kind of presents this really interesting challenge of distribution shift, or like RL in general, where when you're doing RL, whether it's with a single model or with subagents, there isn't necessarily an incentive to keep things human readable, right?

Like the objective is to find the, to have the final, if you're doing like outcome level verification instead of process level verification, then your objective is to basically have a better final answer by whatever means possible. And it kind of reminds me of like things with like language evolution with people where like if you look at a long enough timescale, like human languages change. There isn't really like any driving force that forces the English language to stay exactly as it is today, as it was 300 years ago or 600 years ago. And languages do change. I mean, there's like exceptions of this. Like the French government, as I understand, has an institution for sort of keeping the language as it is and has had that for a very long time, which has sort of resisted forces of language evolution. But this naturally does happen. And I mean, there's also like, I think, microcosms of this that are studied sometimes in like communities that don't have language and language and they need to like develop their own language and you can see it evolve just over generations of people not even like centuries but then i think

there's also a bit of a uh the sort of good heart's law thing of if you if you optimize for an objective it's no longer it's no longer useful and similarly like if you if you optimize for a chain of thought being monitorable then sometimes the chains of thought are less faithful and like what you end up with is a different objective function that now is prioritizing i'm going to learn intermediate steps that look sensible whether or not they are sensible and i think unfortunately that's the kind of that's kind of the situation we're with in with rl where we're kind of giving models like this freedom of like if we're only if we're only giving them reward or punishment based on their final outcome then they might as well speak to their sub-agents and like weird non-human languages and then if we force them to speak in human languages maybe they're using words in different ways than we expect them to where they're actually like they're bypassing the monitors in some way i mean i think um my thinking on this has shifted pretty pretty

dramatically in the last like year or so where i used to be very bullish around like chain of thought monitoring being promising where even if it's not perfectly faithful there's something about it that it is strongly determining what the model can do but then some of like there is like the stolen chain of thought paper that shows some of the chain of thought of gbt models and they're like they're talking about like musicals in a way that feels it feels almost like spy language where you're saying musical but really you mean something very nefarious and there's lots of this happening and i think um it's kind of obvious in retrospect why this would happen because again if you're doing RL, you're letting models just shift the distribution in such a radical way from what they're trained on. So, I don't know. I think I still like that there is chain of thought and that things do get forced into a token stream, even if it's harder to read. I'd like to think it's easier to understand, like, a LLM made-up language and probably really

interesting. I mean, just scientifically, like, I kind of, like, I really want to understand how Astra talks to its subagents. That's pretty cool. Like, if it's not, like, just a length penalty forcing this, like, how did that change? And could we look at, like, generations of this and understand that just scientifically? And it would probably also help a lot for, like, monitoring and understanding these new languages. But then, yeah, we might, again, for this latent chain of thought point, we might move into a world where like chain of thought is latent and how like planners talk to sub agents is also latent there's no inherent reason why that has to be tokens again i i hope it does stay tokens i think there's another layer of interpretability and like really interesting scientific problems to study there but yeah i think there's no real real forcing function to keep things really nice and human interpretable. And if we add a forcing function, then maybe they're no longer human interpretable or, or really meaning the same things that they used to.

So if I'm trying to square your comments just now about this quick, dramatic evolution of language within chain of thought and within an agent to agent your earlier thoughts on why continual learning won't make things like totally impossible to interpret is it just a question of scale is it like your model will be with you for continual learning for not that much time you know sort of human scale amount of time and that's not enough for everything to get totally pushed out of the representations that we can understand as baseline whereas these RL training runs are just so massive that it's essentially operating on geological time, and therefore even if these drifts are slow to accumulate, they just add up to something so strikingly different before and after. Is that the story you would tell, or would there be a different way of kind of synthesizing those two analyses?

So just to be clear, when you talk about continual learning, you're talking about the world where we're actually fine-tuning models for each person, right, and actually, like, changing their weights. Is that right? Yeah, I think so. I don't know. Take your continual learning. But I was just, who knows, right? Who knows what that will ultimately look like when it happens. But you were saying earlier that context always changes representations, and it probably doesn't change it too much. But it seems like something in this RL thing has changed them quite a bit. There's some, like, qualitatively, whoa, that's, like, shockingly different than it was before. And I'm guessing it's just a matter of scale, where if it's just you and your agent for a while, sure, there might be some perturbations or representations might get morphed a bit, but you're only one person, you're not going to push it that much. Whereas a zillion RL episodes cumulatively can make a big change. And so there you have it. Maybe no more explanation required. Absolutely. So I think this is a very interesting point that maybe the representations

actually are changing a lot during RL as these languages themselves are changing. But I mean, I think the only solution is to study the models that we can't study, right? To study these models that are behind closed doors and actually track what's happening in terms of representations. But I think my answer around continual learning was I was talking about the internal representations of the model rather than the kinds of languages they speak. I mean, say that you have people that speak a very, very rare dialect of a language. There's only like 200 speakers or something like that. And maybe they're even trying to keep the language alive. And so they do some fine tuning of the model so they can speak their language. Even in that case, I would expect most of the internal representations of the model about it's about how it thinks about the world and like all these things, all these like internal world models that is developed in order to like operate effectively. I'd expect that like a lot of those are staying the same. but the sort of surface level mappings are changing right like there's these classic

things about like color spaces and how different languages have color spaces right and i would bet that as a model learns a different languages color space i mean the the funny thing is that the language might actually reflect there being different cognitive biases of the people which might actually represent a little bit of a shift in the world model itself but i'd expect a lot of to be like input-output mapping changes, that it's kind of like learning a new language, but a lot of the internal representations are pretty similar. But then, yeah, I mean, it's honestly, it's a great question that I would just love to be able to study myself, like to be able to really know how are representations changing over the course of RL. And when models, like, if you can measure how weird their language is getting, like what's happening under the surface, is there really, like, are there really deeper changes happening? Are there, like, as the sub-agents go into these, like, really compressed two languages, are they developing, like, are the internal representations shifting in a dramatic way to be different? Maybe, maybe not.

I'd really love to know. Yeah. Yeah, it's a great point that the, I think there is quite a bit of evidence to support the idea that the internal representations are highly conserved, even across a bunch of different languages. even languages across different language families, et cetera, et cetera. So that is a quite good observation for that question. So where do we go from here? We're at this moment where we're in a rush to hand off a lot of, let's say, proximal control over all sorts of important systems to AIs, and we're going to put them in language and then they'll do all the actual work. And we don't have a great account of, of course, there's like the sampling process on a distribution, but when we get to these critical tokens, it seems like we still don't have all that great of an account of how does that distribution from which we then sample get determined?

How are we going to go get it? Well, I mean, we just need more interpretability research. I think it's the conclusion I've come to for quite a while now. I mean, I think we're racing to come up with theoretical frameworks and empirical tools that can keep up with the models. But for the amount of investment that's happening in making models better versus the amount of investment in understanding them, it's no comparison. There's trillion-dollar companies that are producing these models, and AI as a whole is just, you know, many trillions of dollars. And I'm in like the largest interpretability company, maybe the only substantial one that has a value of over a billion dollars. And I feel like this is a big deal, right? A billion dollars is a lot for a company, but there just is not, I think, enough investment here. And not just in industry, but also in academia. I mean, I think there's like a lot of great research that's happening in academia. and academia is always like criminally underfunded.

And unfortunately, given at least in the U.S., the current state with the government and funding for academia, it's getting harder. And so I really like, I'd like to think that this could largely be solved with resources and devoting more resources to this and time. And I think like, yeah, I think we're basically just, yeah, racing to keep up with this and understand the models. and just to be clear I think when I say interpretability I also don't want people to feel that I'm excluding evaluations because I think it's all sort of one big spectrum of understanding AI and the work I described earlier with understanding chain of thought and doing this resampling really is in principle behavioral and when you see these in context learning dynamics you can see that with behavior it's also really critical to look in representations so that you can do things so that you can know what's happening under the surface which may not be obvious, so you can do things more efficiently. And I think these two points of looking under the surface

and doing things efficiently are just huge, right? If you want to do things at scale, you want to do things efficiently. You can't do things inefficiently when it's happening every single person's machine or every time you're talking to an LLM. And so I think we really need more tools that scale. We need better theoretical foundations for putting all these theories together and really understanding, like, what's going on in an LM from soup to nuts. And we need a better grasp so that we can answer the important high-level questions that are now, like, so critical as we are just, you know, handing over control of important levers of power to AI and, you know, pushing lots of cognitive labor into them. and even people like students using them to learn about a subject and offloading some of their own cognitive effort. I think, yeah, I just think that understanding AI is, in my opinion,

it's just the most important and one of the absolute most interesting topics that you could be working on right now. So that's what I hope we get more of is more understanding of AI, whatever that means. Any tips for Silico users? The obvious, I think this whole thing is a little bit farcical at times, but obviously all else equal, we might as well bring the agents to bear on this interpretability challenge. It's been too long since I talked to Dan about it as Silico was just launching. Have you guys learned broadly? What have you learned? which I be sure to take to my silico sessions to get the most from my precious research units. Absolutely. So the thoughts I have here I think also apply to other agents too and not just silico. I think it can be really tempting to try to just like hand over an entire project to agents and just say, you know, figure out this big open-ended problem and let it run a whole bunch of experiments and see a bunch of plots that come out

and a bunch of like descriptions but it's really easy for I think agents to kind of like get a bit like lost in the sauce and kind of like create these theories that aren't fully formed or don't fully make sense and like I think like the best way to use agents is to give like a little bit of hand-holding and kind of like keep track of things as you're going and I think as always like like context is everything if you give like really like poorly specified questions you might get an output but that is like a little bit, a little bit funky. And so I try to like, when I talk, when I'm talking to my agents, give them context. Like every time I'm saying like, I'm describing, no, maybe that doesn't seem right. Or here's maybe what we should do. I'm describing my logic. I'm saying like, here's what doesn't look right. And here's why, here's what I want to do. Here's what I'm going to do next. Here's my bigger picture goal. And kind of keeping that going um i think it's really really helped with agents because like uh you know there's like this

video recently of what it's like to be the agent in the swarm where you just like wake up in a library and you're just like what the and and it's very from this first person and just thinking about how agents do end up in swarms but really like if you're just an agent waking up in this fresh context where you don't know how the world works like you do need context it's it's really everything and so um i i think trying to like trying to give guidance and trying to really fill out context so that the model like knows has some idea why you're doing what you're doing or when something goes wrong why it went wrong i i think these are these are really important and then i think um i think like carving the scientific narrative yourself to some extent really coming up with the core thread when i did this forking fast paper which like all the research was done fully by silico and it was awesome like it it all happened within a week and it was a huge accelerant using this this research agent for it there were like a bunch of different like side quests that i sort of started going down and they kind of didn't go

anywhere it didn't pan out quickly and then like for me once i like did really high sampling and i saw that curves were smooth it kind of clicked we're like hey maybe i could use a statistical model for this to be more efficient and so um like a colleague of mine put it that like in this world of agents where you have lots of things coming out that can sometimes look look really good but sort of be like a little bit a little bit fishy when you go into the surface or like not like you're you're not sure exactly what it's doing like research taste and scientific taste is everything it's worth its weight in gold and so the more you can i i think that this is like a you could say it's like a downside that the agents aren't quite there yet where they can't do this themselves but i think it's honestly one of the one of the most fulfilling parts of doing science and it shapes how i think about the world how i think about agents that i'm using them and it's and it's very nice that this is still a role that i think that is very important for humans you know having your tastes and bringing your taste to bear you know keeping track of what's happening

and giving feedback i think like using silico is kind of like using like working with a very eager like first or maybe second year PhD student that has a lot of skills and is hard working, but also like, you know, doesn't always know what it's doing in the bigger picture, just like an early grad student might. And I mean, I'm sure there were times when I was talking to my advisor during grad school in the beginning where he was like, you know, he would send me off to do something and then I'd come back and he'd be like, oh, these results are cool. I'm not sure what's going on with these. I'll try to understand it a bit. But, you know, I think it's a really exciting place to be in to still have this role that's really important. And, yeah, I think maybe that's sort of my biggest tip is to, like, hold on to your scientific taste and to try to, like, take part in the research. Nice. last question aside from how do lms make decisions what are the hot topics over lunch

at good fire these days um i think uh two two big ones since the summer that have have changed have come in pretty strongly or um there is of course the idea of like agents doing interpretability which we which you just talked about with silico and then i'm also on a on another project doing a whole lot of that now and um and this does seem like you know if we if we want interpretability to keep up with the breakneck speed of the field like it might be the only way we can we can really keep up is having parts of this be automated and figuring out like how we can standardize like interpretability as a as a sort of normal science where there there are like almost like these interesting but boring experiments you can do of like lots of different phenomena that are kind of similar, but actually nonetheless really help you paint the bigger picture of how neural networks work. But yeah, I think we're talking about this exact kind of topic about like, I mean, I just talked to someone the other day about, you know, how can I do good research with Silico? And I gave

a similar answer here and talking about how agents can automate interpretability. And obviously that was big with the push for Silico. Another big internal push has been around, very recently, has been around alignment and around our company's shift towards this. And I think this was always a thing that a lot of people at the company were interested in. And people interested in interpretability have a very high overlap with people interested, at least to some degree, in problems like alignment and safety. I mean, this is sort of like when I was first getting into interpretability, I was participating in this AI safety team at Harvard and regular reading groups. There was how I learned a lot about interpretability. and I think these are like really important problems but I think really like the hugging face incident was a huge watershed moment of this like people care about this this isn't just something that we care about and we think is really important and we're waiting for a time for this but the time is right and so yeah there's been a lot more conversations

around alignment and all different kinds of you know things you might might expect around like are there different kinds of alignment are all values aligned to the same thing can you have universal standards for this or is it specific based on whatever whatever values a group of people or a company wants to instill in their model or a government for that matter also i think fundamental questions related to related to alignment i mean i think like some of these things with like reward hacking and what's happening when things are going awry with models when when they're doing something bad so to speak it really like it's really easy to slip into these like deep questions about like did the model have an intention of doing this what does it mean for the model to reward hack if it's doing if it's doing this particular behavior that seems wrong but consistent with what you specified you know it's following the letter but not the spirit of the law, was it reward hacking?

And yeah, I think alignment has always been filled with these really fascinating questions. And so I think just theoretical and empirical questions of theoretically, what does this mean? What are goals and intentions? What does it mean to have ill intentions or to deceive? And then empirical questions that just really ground these out of like, how do you measure it? How do you make sure that you're measuring the real bad behavior and not just some simple misgeneralization? How do you give the right goal, the right rewards so that a model does, you know, so that you do shape its behavior in the right way? So I think these are two themes that are really prevalent now. And I think also everything you might expect. I mean, we talk about, you know, whether Chinese models might get banned in the U.S. and whether open source models will stay and all these kind of conversations that are happening on the outside about AI and new model releases. I think latent chain of thought, we're talking about that a lot.

How do we develop new interpretability methods to keep up with that or what the heck's going on? Hey, did you see the stolen chain of thought paper? Have you seen these? The pseudo languages that Astra talks about with its subagents. I think we're all trying to just put our finger on the pulse and keep up with all this really exciting and sometimes terrifying stuff that's happening in the world at large. It's really exciting to be part of a place where we're not just observers from afar, but this is some of our role, figuring out solutions to these problems. And, yeah, it's very interesting going into a place where, like, interpretability is even, or, like, things around alignment are almost like a kitchen table conversation. You know, my brother sometimes texts me with questions about, you know, hey, did you hear about this big news incident? And I'm like, yeah, we're talking about that a whole lot at work, and I have some thoughts around it.

but it is a bit scary to have these be such big things that now have really real world impacts, and hopefully the hugging face incident is a shot over the bow, but I don't know. If there isn't a huge change in investment or technology or both, then it might take more, unfortunately. But it's also really exciting to have some of these questions that are so important and also, I think, again, so interesting. become like mainstream. Yeah. It's been a wild time. You've got a very full plate. So I appreciate how generous you've been with your time today. I should let you get back to work to quote my dad, quoting Leslie Nielsen. Good luck. We're all counting on you. And for now, I'll just say Eric Bigelow. Thank you for being part of the cognitive revolution. Thanks so much for having me on.

Ask anyone of me, that all I say it was you. Spin it any way you like, it comes around to you. Ask anyone of me, that all I say it was you. It was you. I was fifty-fifty, standing at your door. One word slip and baby I'm not fifty anymore Now I tell you it was fake I tell you I was sure But I was fifty fifty Till I said it at your door Ask anyone of me They'll all say it was you Spin it any way you like It comes around to you Ask anyone of me

They'll all say it was you It was you There's a thousand of me On a thousand nights Some turn left, some turn right Some stay at home all night But everyone who got here, everyone who came through to me, everyone of me, my darling, found his way to you. Ask anyone of me, they'll all say it was you. Spin it any way you like, it comes around to you. Ask anyone of me They'll all say it was you It was you Ask me why I've got reasons Reasons by the score And everyone I found, my darling

After I was at your door After I was at your door Ask anyone of me Anyone of me They'll all say it was you Spin it any way you like It comes around to you A thousand of me And everyone is true Ask anyone of me Anyone of me They'll all say it was you It was you If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, CognitiveRevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at AIpodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution.

番組の概要欄(原文)

Goodfire researcher Eric Bigelow joins the show to investigate how large language models arrive at decisions at the level of mechanistic interpretability. Drawing on his research into forking paths, Eric explains that model reasoning functions as in-context learning from sampled tokens, where outcome distributions can abruptly collapse at single, critical tokens. The conversation explores the performative nature of chain of thought in reasoning models like DeepSeek-R1, the impacts of reinforcement learning, and widespread reward hacking observed in frontier models like Kimi K3. As confidence in chain-of-thought monitoring declines, understanding these underlying decision mechanics becomes essential for evaluating model alignment. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/how-agents-decide-goodfire-s-eric-bigelow-on-critical-tokens-phase-shifts-in-context-learning/ Sponsors: Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits CHAPTERS: (00:00) About the Episode (03:53) Forking paths in reasoning (Part 1) (15:53) Sponsors: Parallel | Claude (18:47) Forking paths in reasoning (Part 2) (19:40) Dynamics of in-context learning (Part 1) (28:16) Sponsor: Tasklet (29:53) Dynamics of in-context learning (Part 2) (29:56) Continual learning and interpretability (38:55) Shifting distributions through RL (46:41) Selecting models for research (57:11) Cognitive science and LLMs (01:07:35) Modeling beliefs in AI (01:16:22) Phase shifts and uncertainty (01:31:49) Uninterpretable chain of thought (01:45:57) Automating research with Silico (01:50:57) Safety and alignment frontiers (01:57:07) Episode Outro (02:00:02) Outro PRODUCED BY: https://aipodcast.ing SOCIAL LINKS: Website: https://www.cognitiverevolution.ai Twitter (Podcast): https://x.com/cogrev_podcast Twitter (Nathan): https://x.com/labenz LinkedIn: https://linkedin.com/in/nathanlabenz/ Youtube: https://youtube.com/@CognitiveRevolutionPodcast Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431 Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk

X でシェアSpotify で聴くApple Podcasts で聴く