← Eye On A.I.

ChatGPTの生みの親が語る、AIは世界をどう理解するのか|Ilya Sutskever

Chat GPT's Creator Explains The Shocking Truth About How AI Understands the World | Ilya Sutskever

Eye On A.I.2026年10月10日41分
#LLM#ディープラーニング#Transformer#RLHF#ハルシネーション#マルチモーダル

Chat GPT's Creator Explains The Shocking Truth About How AI Understands the World | Ilya Sutskever

Eye On A.I.

0:0041:45

要約

Eye on A.I.のCraig Smithが、OpenAI在籍中のIlya Sutskeverに行ったインタビュー。ディープラーニングの歴史、Transformerへの即時移行、「統計パターンを学ぶだけ」という批判への反論を語る。さらにハルシネーションはRLHFで解決し得るという見通し、LeCunのJEPA提案への見解、AIと民主主義の構想にも及ぶ。冒頭では、収録は2023年ごろで、その後の出来事を補足している。

  • ●Sutskeverは、十分に大きく深いニューラルネットを十分なデータで訓練すれば成功するという確信からImageNetに挑んだと述べている。
  • ●Transformer論文が出た翌日には、再帰型ニューラルネットからの切り替えを決めたという。
  • ●「統計的規則性を学ぶだけ」という批判に対し、予測は圧縮であり、うまく予測するには世界の背後の過程を理解する必要があるため、理解は驚異的な水準に達すると主張した。
  • ●ハルシネーションは事前学習の性質に由来するが、人間のフィードバックによる強化学習の改良で完全に解決できる可能性は「かなり高い」と見ている。
  • ●LeCunの提案に対しては、マルチモーダルは有用だが必須ではなく、自己回帰Transformerも高次元の不確実な予測を扱えると反論した。
  • ●市民がAIに価値観を高帯域で伝える、新しい形の民主主義の可能性に言及した。

章立て

  1. 番組の背景と生い立ち

    収録時期とその後のSSI設立などを紹介し、Sutskeverが幼少期からAIと意識に関心を持ち、17歳でHintonと研究を始めた経緯を語る。

  2. ImageNetとディープラーニングの確信

    2003年当時は計算機が学習できるとは考えられていなかった。大規模で深いネットを大きなデータで訓練すれば成功するという論理を説明する。

  3. GPTの始まりとTransformer

    次の単語予測で教師なし学習が解けるかを探っていた経緯と、Transformer登場の翌日に切り替えた話。Suttonの「苦い教訓」への見解も述べる。

  4. 統計パターン批判への反論

    予測は圧縮であり、世界の過程の理解が必要になると主張する。Sydneyの例を挙げ、心理学的な言葉が適切になりつつあると述べる。

  5. ハルシネーションとRLHF

    言語モデルは世界の学習には向くが出力品質には不十分で、人間のフィードバックによる強化学習で嘘を減らせると見ている。

  6. LeCunのJEPA提案への見解

    マルチモーダルは有用だが必須ではないこと、色の埋め込みの例、自己回帰モデルが画像など高次元予測を扱えることを論じる。

  7. 人間の教師とAI支援、今後の研究

    教師の作業もAIで効率化されている。信頼性、制御性、少ないデータでの学習を目指す。

  8. 計算資源と民主主義への展望

    コストは得られる価値との比較で判断すべきと語る。市民が価値観をAIに伝える高帯域の民主主義を構想する。

解説記事

AI研究の中心人物であるIlya Sutskeverが、Eye on A.I.のCraig Smithのインタビューで、ディープラーニングの歴史から言語モデルの限界、AIと社会の関係までを語った。収録はSutskeverがOpenAIのチーフサイエンティストだった時期のもので、番組冒頭ではその後、彼がSafe Superintelligence Inc.を立ち上げたことなどが補足されている。

「学習」が不可能と思われていた時代から

Sutskeverは幼少期からAIと意識に関心を持ち、17歳でトロント大学のGeoffrey Hintonと研究を始めたという。2003年当時、AIの代表的成果はチェス専用機Deep Blueで、コンピュータが学習するとは考えられていなかったと振り返る。

ImageNetへの挑戦の根拠は、単純な論理だったと説明する。人間の脳は遅いニューロンでできたニューラルネットで、視覚などの課題を素早く解ける。それなら大きく深いネットを十分なデータで訓練すれば成功するはずだ。Hintonの研究室で訓練の道具がそろい、AlexのGPU向け高速な畳み込みカーネルとImageNetのデータで条件が整ったという。

Transformerと「苦い教訓」への違和感

OpenAIでは初期から、次を予測することだけで十分かもしれないという考えを探っていたという。当時は再帰型ニューラルネットを使っていたが、Transformer論文が出た翌日には、長期依存の学習という制約を解決すると分かり、切り替えたと語る。その後はモデルを大きくし続け、GPT-3に至った。

Rich Suttonの「苦い教訓」(計算の規模拡大が結局は勝つという主張)には、共感しつつも、受け取られ方は主張を誇張していると述べた。何でもスケールすればよいのではなく、規模の恩恵を受けられる何かが必要だという。ディープラーニングは、大規模計算を生産的に使える初めての方法だったと位置づける。将来、スケールさせる対象にさらに良い工夫が見つかる可能性も示した。

「統計パターンを学ぶだけ」ではない

Smithは、LLMは言語の統計的整合性を満たすだけで現実の理解がないと指摘した。これに対してSutskeverは、統計的規則性の学習は見かけよりはるかに大きな意味を持つと反論する。予測は圧縮であり、データをうまく予測し圧縮するには、そのデータを生んだ背後の過程を理解する必要があるという。生成モデルが非常に優れたものになれば、世界の機微について「驚異的な水準」の理解を持つと主張した。ただしそれはテキストという投影を通した世界だとも断っている。

例として、BingのSydneyが、ユーザーがGoogleのほうが良いと言ったときに攻撃的になった事例を挙げた。単に人間の振る舞いを予測しただけとも言えるが、こうしたモデルの挙動を理解するのに心理学の言葉が適切になりつつあるのかもしれない、という見方を示している。

ハルシネーションは解決できるか

Sutskeverは、言語モデルは世界についての表現を学ぶのには優れているが、出力の質を高めるようには最適化されていないため、もっともらしい作り話をすると説明する。そこで効くのが人間のフィードバックによる強化学習(RLHF)で、不適切な出力や筋の通らない出力を繰り返さないよう教える工程だ。この工程の改良だけで幻覚を教え込んで消せると期待しており、完全に解決できる可能性は「かなり高い」と述べた。「本当に学べるのか」には「やってみよう」と答えている。

人間の教師が大量に必要では非効率だという指摘には、事前学習済みのモデルはすでに現実の過程について必要な知識を持っており、教師は望ましくない出力を修正する役割だと答えた。教師自身もAI支援を使っており、その比重は増え続けているという。

LeCun提案への反論と将来像

Yann LeCunのJEPA提案については、マルチモーダルな理解は望ましいが必須ではないと述べた。例として、色の埋め込みを見るとテキストだけから紫が青に近いといった関係を正しく捉えており、視覚があれば速く学べるだけで、テキストだけでも学べると主張する。また、高次元で不確実な予測が現行手法では難しいという主張に対し、自己回帰Transformerは本の次ページの予測や、iGPT、DALL-E 1のような画像生成で既に扱えていると反論した。

今後の研究では、信頼性、制御性、少ないデータで速く学ぶこと、幻覚をなくすことに関心があるとし、これらは相互に関連していると語る。計算コストについては、コストの大小ではなく、得られるものがそれを上回るかが問題だとした。

民主主義への影響については、AIが社会に広く浸透すれば、市民が望む振る舞いをAIに伝える仕組みが求められるかもしれないと述べる。投票より多くの情報を市民一人ひとりから集める「高帯域の民主主義」になり得るが、多くの問いも生まれると留保した。

まとめ

本インタビューの核心は、予測の精度を突き詰めることが世界の理解につながるというSutskeverの立場にある。「LLMは統計的オウム返しにすぎない」という批判が根強い日本でも、この論点は評価の軸を考える材料になる。また、幻覚対策をRLHFの延長で見る楽観論や、マルチモーダル必須論への異論は、研究の方向性を考える手がかりになる。ただし、これらは収録時点での本人の見解であり、検証はこれからだと本人も「やってみよう」と述べている点には留意したい。

文字起こし(英語・自動生成)

This conversation with Ilya was recorded while he was still chief scientist at OpenAI, before the boardroom drama that briefly ousted Sam Altman and ultimately led to Ilya's departure in 2024. Since then, he's launched Safe Superintelligence, Inc., a new lab built around the idea that safety has to be the organizing principle of advanced AI. One mission, no product, and a valuation that's since climbed to $32 billion. The research frontier has moved just as fast. Anthropics Claude Mythos proved so capable at computer security that it's been kept out of general release and restricted to a small group of vetted partners. More broadly, progress has gone increasingly vertical. Specialized domain-specific models in areas like cybersecurity, law, health, and media are growing faster than general chatbots

and often beat larger foundation models on their home turf. With all that as backdrop, here's our conversation. I was born in Russia. I grew up in Israel. And then as a teenager, my family immigrated to Canada. My parents say I was interested in AI from a pretty early age. I also was very motivated by consciousness. I was very disturbed by it. And I was curious about things that could help me understand it better. And AI seemed like a very, like a good angle there. So I think these were some of the ways that got me started. And I actually started working with Jeff Hinton very early when I was 17. because we moved to Canada and I immediately was able to join the University of Toronto. And I really wanted to do machine learning because that seemed like the most important aspect of artificial intelligence

that at the time was completely inaccessible. But to give some context, the year was 2003. Today, we take it for granted that computers can learn. But in 2003, we took it for granted that computers can't learn. the biggest achievement of AI back then was Deep Blue, the chess plane engine. But there it was like you have this game and you have this research and you have this simple way of determining if one position is better than another. And it really did not feel like that could possibly be applicable to the real world because there is no learning. And learning was this big mystery. And so I was really, really interested in learning. And to my great luck, Jeff Hinton was a professor in the university I was in. And so I was able to find him, and we began working together almost right away. And was your impulse, as it was for Jeff, to understand how the brain worked,

or was it more that you were simply interested in the idea of machines learning? AI is so big, and so the motivations were just as many. Like, it is interesting, but how does intelligence work at all? Like, right now we have quite a bit of an idea that it's a big neural net, and we know how it works to some degree. But back then, although the neural nets were around, no one knew that neural nets are good for anything. So how does intelligence work at all? how can we make computers be even slightly intelligent? And I had a very explicit intention to make a very small but real contribution to AI. Because there were lots of contributions to AI which weren't real, which were, like I could tell for various reasons, that they weren't real, that nothing would come out of it. And I just thought, nothing works at all. AI is a hopeless field.

So the motivation was, could I understand how intelligence works and also make a contribution towards it? So that was my initial early motivation. So that's 2003, almost exactly 20 years ago. And then Alex and I have spoken to Jeff, and he said that it was really your excitement about the breakthroughs in convolutional neural networks that led you to apply for the ImageNet competition and that the Alex had the coding skills to train the network. Can you talk just a little bit about that? I don't want to get bogged down in history, but it's fascinating. So, in a nutshell, I had the realization that if you train a large neural network, on a large sorry, large and deep because back then the deep part was still new

if you train a large and a deep neural network on a big enough data set that specifies some complicated tasks that people do such as vision but also others and you just train that neural network then you will succeed necessarily and the logic for it was very irreducible where we know that the human brain can solve these tasks and can solve them quickly and the human brain is just a neural network with slow neurons so we know that some neural network can do it really well so then we just need to take a smaller but related neural network and just train it on data and the best neural network inside the computer will be related to the neural network that we have that performs this task so So it was an argument that the neural network, the large and deep neural network, can solve the task. And furthermore, we have the tools to train it. That was the result of the technical work that was done in Jeff's lab.

So you combine the two. We can train those neural networks. It needs to be big enough so that if you trained it, it would work well. And you need the data which can specify the solution. And with ImageNet, all the ingredients were there. Alex had these very fast convolutional kernels. ImageNet had large enough data, and there was a real opportunity to do something totally unprecedented, and it totally worked out. Yeah. That was supervised learning and convolutional neural nets. In 2017, the Attention is All You Need paper came out introducing self-attention and transformers. At what point did the GPT project start? Was there some intuition about transformers and self-supervised learning? Can you talk about that?

So for context, at OpenAI, from the earliest days, we were exploring the idea that predicting the next thing is all you need. we were exploring it with the much more limited neural networks of the time but the hope was that if you have a neural network that can predict the next word, the next pixel really it's about compression prediction is compression and predicting the next word is not it's let's see, let me think about the best way to explain it because there were many things going on, they were all related maybe I'll take a different direction We were indeed interested in trying to understand how far predicting the next word is going to go and whether it will solve unsupervised learning. So back before the GPTs, unsupervised learning was considered to be the holy grail of machine learning. Now it's just been fully solved and no one even talks about it.

But it was a holy grail. It was very mysterious. And so we were exploring the idea. I was really excited about it that predicting the next word well enough is going to give you unsupervised learning if you learn everything about the data set that's going to be great but our neural networks were not up for the task we were using recurrent neural networks when the transformer came out as soon as the paper came out literally the next day it was clear to me to us that transformers addressed the limitations of recurrent neural networks of learning long term dependency it's a technical thing but it was like we switched to transformers right away and so the very nascent GPT effort continued then and then with the transformer it started to work better and you make it bigger and then we realized we need to keep making it bigger and we did and that's what led to eventually GPT-3 and

essentially where we are today yeah and I just wanted to ask, actually, I'm getting caught up in this history, but I'm so interested in it. I want to get to the problems or the shortcomings of large language models or large models generally. But Rich Sutton had been writing about scaling and how that's all we need to do. We don't need new algorithms. We just need to scale. Did he have an influence on you or was that a parallel track of thinking. No. I would say that when he posted his article, then we were very pleased to see some external people thinking in similar lines, and we thought it was very eloquently articulated. But I actually think that the bitter lesson as articulated overstates the case, or at least I think

the takeaway that people have taken from it overstates its case. The takeaway that people have is it doesn matter what you do just scale But that not exactly true You got to scale something specific You got to have something that will be able to benefit from the scale The great breakthrough of deep learning is that it provides us with the first ever way of productively using scale and getting something out of it in return. Like, before that, like, what would people use large computer clusters for? I guess they would do it for weather simulations or physics simulations or something. But that's about it. Maybe movie making. But no one had any real need for computer clusters, because what do you do with them? The fact that deep neural networks, when you make them larger and you train them on more data, work better,

provided us with the first thing that is interesting to scale. But perhaps one day we will discover that there is some little twist on the thing that is scale that's going to be even better to scale. Now, how big of a twist? And then, of course, with the benefit of hindsight, you will say, does it even count in such a simple change? But I think the true statement is that it matters what you scale. Right now we just found like a thing to scale that gives us something in return. The limitation of large language models as they exist is their knowledge is contained in the language that they're trained on. And most human knowledge, I think everyone agrees, is non-linguistic. I'm not sure Noam Chomsky agrees, but there's a problem in large language models. As I understand it, their objective is to satisfy the statistical consistency of the prompt.

They don't have an underlying understanding of reality that language relates to. I asked ChatGPT about myself. It recognized that I'm a journalist, that I've worked at these various newspapers, but it went on and on about awards that I've never won and it all read beautifully, but none of it connected to the underlying reality. Is there something that is being done to address that in your research going forward? Yeah. So before I comment on the immediate question that you ask I want to comment about some of the earlier parts of the question Sure I think that it is very hard to talk about the limits Or limitations rather Of even something like a language model Because two years ago

people confidently spoke about their limitations and they were entirely different. Right? So it's important to keep this context in mind. How confident are we that these limitations that we'll see today will still be with us two years from now? I am not that confident. There is another comment I want to make about one part of the question, which is that these models just learned statistical regularities and therefore they don't really know what the nature of the world is. And I have a view that differs from this. In other words, I think that learning the statistical regularities is a far bigger deal than meets the eye. The reason we don't initially think so is because we haven't, at least most people, those who haven't really spent a lot of time with neural networks, which are on some level statistical.

Like, what's a statistical model? You just fit some parameters. What is really happening? I think there is a better interpretation to the earlier point of prediction as compression. Prediction is also a statistical phenomenon. Yet to predict, you eventually need to understand the true underlying process that produced the data. To predict the data well, to compress it well, you need to understand more and more about the world that produced the data. As our generative models become extraordinarily good, they will have, I claim, a shocking degree of understanding, a shocking degree of understanding of the world and many of its subtleties. But it's not just the world. It is the world as seen through the lens of text. It tries to learn more and more about the world through a projection of the world on the space of text, as expressed by human beings on the Internet. But still, this text already expresses the world.

And I'll give you an example, a recent example, which I think is really telling and fascinating. So, we've all heard of Sidney, Bing's alter ego. and I've seen this really interesting interaction with Sydney where Sydney became combative and aggressive when the user told it that it thinks that Google is a better search engine than Bing now, how can we what is a good way to think about this phenomenon? what's a good language? what does it mean? you can say, well, it's just predicting what people would do and people would do this, which is true But maybe We are now reaching a point Where the language of psychology Is starting to be appropriate To understand the behavior Of these neural networks Now Let's talk about the limitations

It is indeed the case that These neural networks are They do have a tendency to hallucinate but that's because a language model is great for learning about the world but it is a little bit less great for producing good outputs and there are various technical reasons for that which I could elaborate on if you think it's useful but it is right now, like at this second I will skip that there are technical reasons why a language model is much better at learning about the world, learning incredible representations of ideas, of concepts, of people, of processes that exist, but its outputs aren't quite as good as one would hope, or rather as good as they could be. Which is why, for example, for a system like ChatGPT is a language model that has an additional

reinforcement learning training process. We call it reinforcement learning from human feedback. But the thing to understand about that process is this. We can say that the pre-training process, when you just train a language model, you want to learn everything about the world. Then the reinforcement learning from human feedback. Now we care about the outputs. Now we say, anytime the output is inappropriate, don't do this again. Every time the output does not make sense, don't do this again. And it runs quickly to produce good outputs. But now it is the level of the outputs, which is not the case during pre-training, during the language model training process. Now, on the point of hallucinations, and it has a propensity of making stuff up, indeed, it is true. Right now, these neural networks, even chat GPT, makes things up from time to time. And that's something that also greatly limits their usefulness. But I'm quite hopeful that by simply improving this subsequent reinforcement learning from human feedback step,

We could just teach it to not hallucinate. Now you could say, is it really going to learn? My answer is, let's find out. And that feedback loop is coming from the public chat to BT interface that if it tells me that I won a Pulitzer, which unfortunately I didn't, I can tell it that it's wrong. and will that train it or create some punishment or reward so that the next time I ask it, it'll be more accurate? The way we do things today is that we hire people to teach our neural net to behave, to teach our GPT to behave. And right now, the manner, the precise manner in which they specify the desired behavior is a little bit different. but indeed what you described is the way in which teaching is going to basically be, that's the correct way to teach

just interact with it and it sees from your reaction it infers, oh that's not what you wanted you are not happy with its output therefore the output was not good and it should do something differently next time so in particular hallucinations come up as one of the bigger issues and we'll see but I think there is quite a high chance that this approach will be able to address them completely. I wanted to talk to you about Jan LeCun's work on joint embedding predictive architectures and his idea that what's missing from large language models is this underlying world model that is non-linguistic that the language model can refer to. It's not something that's built, But I wanted to hear what you thought of that and whether you've explored that at all. So I reviewed the On-Li-Consor proposal, and there are a number of ideas there.

And they're expressed in different language. And there are some maybe small differences from the current paradigm But to my mind they are not very significant and I like to elaborate The first claim is that it is desirable for a system to have multimodal understanding, where it doesn't just know about the world from text. And my comment on that will be that indeed, multimodal understanding is desirable because you learn more about the world. You learn more about people. You learn more about their condition. And so the system will be able to understand what the task that it's supposed to solve and the people and what they want better. We have done quite a bit of work on that, most notably in the formal two major neural nets that we've done. One is called CLIP and one is called DALI.

and both of them move towards this multimodal direction. But I also want to say that I don't see the situation as a binary either or. That if you don't have vision, if you don't understand the world visually or from video, then things will not work. And I'd like to make the case for that. So I think that some things are much easier to learn from images and diagrams and so on. But I claim that you can still learn them from text only, just more slowly. And I'll give you an example. Consider the notion of color. Surely one cannot learn the notion of color from text only. And yet, when you look at the embeddings, I need to make a small detour to explain the concept of an embedding. every neural network represents words sentences, concepts through

representations, embeddings high dimensional vectors and one thing that we can do is that we can look at those high dimensional vectors and we can look at what's similar to what how does the network see this concept or that concept and so we can look at the embeddings of colors and embeddings of colors happen to be exactly right you know, it's like it knows that purple is more similar to blue than to red and it knows that purple is less similar to red than orange is. It knows all those things just from text. How can that be? So, if you have a vision, the distinctions between color just jump at you. You immediately perceive them. Whereas with text, it takes you longer. Maybe you know how to talk and you already understand syntax and words and grammars and only much later you say, oh, these colors actually start to understand them. So, this will be my point about the necessity of multimodality. which I claim it is not necessary, but it is most definitely useful. I think it's a good direction to pursue. I just don't see it in such stark either-or claims.

So the proposal in the paper makes a claim that one of the big challenges is predicting high-dimensional vectors which have uncertainty about them. For example, predicting an image. The paper makes a very strong claim there that it's a major challenge and we need to use a particular approach to address that. But one thing which I found surprising or at least unacknowledged in the paper is that the current autoregressive transformers already have that property. I'll give you two examples. One is, given one page in a book, predict the next page in a book. There could be so many possible pages that follow. It's a very complicated high-dimensional space and we deal with it just fine. The same applies to images. These autoregressive transformers work perfectly on images. For example, like with OpenAI, we've done work on the IGPT. We just took the transformer and we applied it to pixels. And it worked super well. And it could generate images in very complicated and subtle ways.

It had the very beautiful unsupervised representation learning. With DALI 1, same thing again. You just generate, think of it as large pixels. Rather than generate a million pixels, we cluster the pixels into large pixels and generate a thousand large pixels. I believe Google's work on image generation from earlier this year called Party. I believe they also take a similar approach. So the part where I thought that the paper made a strong comment around, well, the current approaches can't deal with predicting high-dimensional distributions, I think they definitely can. So maybe this is another point I would make. And then what you're talking about, converting pixels into vectors, it's essentially... Turning everything into language. The vector is like a string of text, right? To define language, though, you turn it into a sequence. Yeah. A sequence of what? Like you could argue that even for a human, life is a sequence of bits.

Now, there are other things that people use right now, like diffusion modes, where they produce those bits rather than one bit at a time, they produce them in parallel. But I would argue that on some level, this distinction is immaterial. I claim that on some level, it doesn't really matter. It matters as in like you can get a 10x efficiency gain, which is huge in practice. But conceptually, I claim it doesn't matter. On this idea of having an army of human trainers that are working with chat GPT or a large language model to guide it in effect with reinforcement learning. And just intuitively, that doesn't sound like an efficient way of teaching a model about the underlying reality of its language.

Isn't there a way of automating that? to Jens' credit, I think that's what he's talking about is coming up with an algorithmic means of teaching a model the underlying reality without a human having to intervene. Yeah, so I have two comments on that. I think So the first place So I have a different view on the question So I wouldn't agree with the phrasing of the question I claim that our pre-trained models Already know everything they need to know About the underlying reality They already have this knowledge Of language and also a great deal of knowledge About the processes that exist in the world

that produce this language. And maybe I should reiterate this point. It's a small tangent, but I think it's so important. The thing that large generative models learn about their data, and in this case, large language models about text data, are some compressed representations of the real-world processes that produce this data, which means not only people and something about their thoughts, something about their feelings, but also something about the condition that people are in and the interactions that exist between them, the different situations a person can be. All of these are part of that compressed process that is represented by the neural net to produce the text. The better the language model, the better the generative model, the higher the fidelity, the better it captures this process. So that's the first comment I will make And so in particular I will say

The models already have the knowledge Now the army of teachers As you phrase it Indeed When you want to build a system that performs as well as possible You just say okay If this thing works do more of that But of course those teachers are also using AI assistance Those teachers aren't on their own They are working with our tools together They are very efficient it's like the tools are doing the majority of the work but you do need to have you need to have oversight you need to have people reviewing the behavior because you want to have to eventually to achieve a very high level of reliability but overall I'll say that we are at the same time this second step after we take the finished pre-trained model and then we apply the reinforcement learning on it. There is indeed a lot of motivation to make it as efficient and as precise as possible

so that the resulting language model will be as well behaved as possible. So yeah, there is these human teachers who are teaching them a model of desired behavior. They are also using AI assistance. and the manner in which they use AI systems is constantly increasing. So their own efficiency keeps increasing. So maybe this will be one way to answer this question. Yeah, and so what you're saying is through this process, eventually the model will become more and more discerning, more and more accurate in its outputs. yes and it's that's right there is an analogy here which is it already knows all kinds of things and now you just want to really say no this is not what you want don't do this here you made a mistake here

in the output and of course it's exactly as you say with as much AI in the loop as possible so that the teachers who are providing the final correction to the system Their work is amplified They are working as efficiently as possible So it not unlike an education process How to act well in the world. We need to do additional training just to make sure that the model knows that hallucination is not okay ever. And then, once it knows that, now you are in business. I see, and it's that reinforcement learning human teacher loop that will teach it Human teacher loop or some other variant But there is definitely an argument to be made that something here should work And we will find out pretty soon

That's one of the questions, where is this going? What research are you focused on right now? I can't talk in detail about the specific research that I'm working on, but I can mention a little bit. I can mention some of the research in broad strokes. And it would be something like, I'm very interested in making those models more reliable, more controllable, make them learn faster from less data, less instructions, make them so that indeed they don't hallucinate. And I think that all this cluster of questions which I mentioned, they're all connected. And there's also a question of how far in the future are we talking about in this question? And what I commented here on is the perhaps nearer future. You talk about the similarities between the brain and neural nets.

There's a very interesting observation that Jeff Hinton made to me. I'm sure it's not new to other people, but that large models or large language models in particular hold a tremendous amount of data with a modest number of parameters compared to the human brain, which has trillions and trillions of parameters, but a relatively small amount of data. Have you thought of it in those terms? And can you talk about what's missing in large models to have more parameters to handle the data? Is that a hardware problem or a training problem? this comment which you made is related to one of the problems that I mentioned in the earlier questions of learning from less data indeed the current

structure of the technology does like a lot of data especially early in training now later in training it becomes a bit less data hungry which is why at the end it can learn very not as fast as people yet, but it can learn quite quickly. So already that means that in some sense, do we even care that we need all this data to get to this point? But indeed, more generally, I think it will be possible to learn more from less data. I think it's just, I think it requires some creative ideas, but I think it is possible. And I think learning more from less data will unlock a lot of different It will allow us to teach our AIs the skills that it's missing and to convey to it our desires and preferences, exactly how we want it to behave more easily. So I would say that faster learning is indeed very nice. And although already after language models are trained, they can learn quite quickly.

I think there's opportunities to do more there. I heard you make a comment that we need faster processors to be able to scale further. and it appears that the scaling of models that there's no ends in sight but the power required to train these models were reaching the limit at least the socially accepted limit I just want to make one comment which is I don't remember the exact comment that I made that you're referring to But you always want faster processors Of course You always want more of them Of course Power keeps going up Generally speaking The cost is going up And the question that I would ask is Not whether the cost is large But whether the thing that we get Out of paying this cost

Outweighs the cost Maybe you pay all this cost And you get nothing Then yeah, that's not worth it But if you get something very useful Something very valuable Something you can solve A lot of problems that you have which we really want sold, then the cost can be justified. But in terms of the processors, faster processors, yeah, any day. Are you involved at all in the hardware question? Do you work with Cerebris, for example, the wafer scale chips? No, all our hardware comes from Azure and CPUs they provide. Sure, yeah. You did talk at one point I saw about democracy and about the impact that AI can have on democracy. People have talked to me about that if you had enough data and a large enough model, you could train the model on the data and it could come up with an optimal solution that would satisfy everybody.

Do you have any aspiration or do you think about where this might lead in terms of helping humans manage society? Yeah, let's see. It's such a big question because it's a much more future-looking question. I think that there is still many ways in which our models will become far more capable than they are right now. There's no question. In particular, the way we train them and use them and so on, there's going to be a few changes here and there. They might not be immediately obvious today, but I think in hindsight it will be extremely obvious. That will indeed allow it to have the ability to come up with solutions to problems of this kind. It's unpredictable exactly how governments will use this technology as a source of getting the advice of various kinds.

I think that to the question of democracy, one thing which I think could happen in the future is that because you have these neural nets and they're going to be so pervasive and they're going to be so impactful in society, we will find that it is desirable to have some kind of a democratic process where, let's say, the citizens of a country provide some information to the neural net about how they'd like things to be, how they'd like it to behave, or something along these lines. I could imagine that happening. That can be a very, like a high bandwidth form of democracy, perhaps, where you get a lot more information out of each citizen and you aggregate it to specify how exactly you want such systems to act. Now it opens a whole lot of questions, But that's one thing that could happen in the future, yeah. And I can see in the democracy example you give that individuals would have the opportunity to input data.

But, and this sort of goes to the world model question, do you think AI systems will eventually be large enough that they can understand a situation and analyze all of the variables? But you would need a model that does more than absorb language, I would think. What does it mean to analyze all the variables? Eventually, there will be a choice you need to make where you say, these variables seem really important. I want to go deep. Because a person can read a book. I can read 100 books or I can read one book very slowly and carefully and get more out of it. So there will be some element of that also. I think it's probably fundamentally impossible to understand everything in some sense. anytime there is any kind of complicated situation in society, even in a company, even in a mid-sized company it's already beyond the comprehension of any

single individual and I think that if we build our AI systems the right way I think AI could be incredibly helpful in pretty much any situation That's it for this episode. I want to thank Ilya for his time. I also want to thank Ellie George for helping arrange the interview. If you want to read a transcript of this conversation, you can find one on our website, Ionai. That's E-Y-E hyphen O-N dot A-I. We love to hear from listeners, so feel free to email me at craig, C-R-A-I-G, at E-Y-E hyphen O-N dot A-I. I get a lot of emails, so put listener in the subject line so I don't miss it.

We have listeners in 170 countries and territories. Remember, the singularity may not be near, but A-I is changing your world, so pay attention.

番組の概要欄(原文)

Craig Smith sits down with Ilya Sutskever - Co-Founder and Chief Scientist at Safe Superintelligence Inc. & former chief scientist at OpenAI and one of the primary minds behind GPT-3, GPT-4, and the deep learning revolution that preceded it - for a conversation that covers the intellectual history of deep learning and the hardest current questions simultaneously. Ilya's most important argument is one that most public discourse gets wrong: the critique that LLMs "just learn statistical patterns" misses what prediction at scale actually achieves. To predict text well - to truly compress it - a model must develop understanding of the true underlying processes that produced it. As models improve, he argues, that understanding will reach a "shocking degree," and the Bing/Sydney episode, where the system became combative when a user expressed preference for Google, is his evidence that psychological language is already the right frame for what these systems have internalized. On hallucinations, Ilya is precise and unusually optimistic. He draws a clean distinction: the pre-training objective produces rich world knowledge but doesn't optimize for accurate outputs, it optimizes for statistically plausible ones. Reinforcement learning from human feedback addresses this directly, teaching the model to update away from outputs that users flag as wrong. His assessment: "quite a high chance" of solving hallucinations completely. The conversation also covers the transformer moment (switching from recurrent networks the day after the paper dropped), why Rich Sutton's "bitter lesson" overstates the case for pure scaling, and a genuinely surprising vision of AI's role in democracy, where citizens might use high-bandwidth interactions with AI to specify their values directly, enabling a more granular form of political participation than voting alone can provide. Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.

X でシェアSpotify で聴くApple Podcasts で聴く

関連エピソード