
社会正義のためのデータサイエンス実践ガイド――「アルゴリズム・リアリズム」とは何か
A Data Science Playbook for Social Justice
A Data Science Playbook for Social Justice
The Tech Policy Press Podcast
要約
ミシガン大学のBen Green助教が、新著『Algorithmic Realism』の中核となる考え方を語った回。アルゴリズムの良し悪しは精度や公平性指標ではなく、現実世界で何を引き起こすかで判断すべきだと主張する。公平性指標や説明可能性の限界、「algorithm in the loop」という発想、データサイエンティストの役割の変化、規制や公共主導のAI開発の必要性にも話が及ぶ。
- ●Greenは、アルゴリズムの評価を技術的指標中心に行う従来の姿勢を「algorithmic formalism」と呼び、現実の影響を重視する「algorithmic realism」を対置している。
- ●核となるのは、アルゴリズムの属性を測る最終的な試験はそれが実際に生む影響だとする「データサイエンスのプラグマティスト格率」である。
- ●公平性指標は統計的な対症療法にとどまりがちだと批判し、暴力防止プログラムなど構造的課題に取り組む現場をデータで支援することを提案している。
- ●説明可能性の研究は実証的には意思決定の改善につながらないと指摘し、既存の意思決定ワークフローにアルゴリズムを組み込む「algorithm in the loop」を提唱している。
- ●AIによるコーディングが進む中、データサイエンティストは組織課題とコードをつなぐ「データストラテジスト」へ役割を広げるべきだと述べている。
- ●Greenは、AIの開発速度の議論に偏らず、民主的統制や公的部門・学術研究への投資など方向性そのものを問うべきだという見方を示した。
章立て
メンフィスでの経験と問題意識
大学卒業後のフェローシップでメンフィス市と取り組んだ都市再生プロジェクトを振り返る。データ入手、運用体制、政策上の資源不足といった非技術的課題が壁になったと語る。
研究の転機と実装のギャップ
博士課程で、指標上は良好なアルゴリズムが利用者の理解不足などで機能しない事例に直面した。そこから研究の軸を現実の影響の解明へ移した経緯を話す。
アルゴリズム・リアリズムの定義
形式主義との対比と、現実の影響で判断するプラグマティスト格率を説明する。データサイエンティストが政治的主体であることを自覚し、責任をもって行動するための手順書という位置づけも語る。
公平性指標の限界と構造的アプローチ
公平性指標では不正義は解消されないとし、刑事司法の例で構造的な介入を紹介する。暴力防止プログラムやSupervised Releaseへのデータ科学による支援を挙げる。
Algorithm in the loopと説明可能性
生成AIの利用でも見られる実装の不一致を指摘する。説明を付ける手法は効果が実証されておらず、意思決定者のワークフローを中心に設計すべきだと述べる。
デューイの知性観とデータ戦略家
デューイの考えを引き、知性を行動を導く道具と捉える。ベンチマーク偏重を批判し、AI時代のデータサイエンティスト像を論じる。
規制とAIの方向性をめぐる議論
自発的な実践だけでは足りず、教育改革や規制が必要だと述べる。開発速度ではなく方向性や民主的統制、公共部門への投資が重要だと語る。
解説記事
ミシガン大学のBen Green助教(情報・公共政策)が、Tech Policy Pressのポッドキャストで新著『Algorithmic Realism: Data Science Practices to Promote Social Justice』(MIT Press)の考え方を語った。アルゴリズムの良し悪しを精度や公平性指標ではなく、現実世界での振る舞いで測るべきだという主張が軸になっている。
出発点はメンフィスでの挫折
Greenはこの本を自身の遍歴を反映した個人的なものだと説明する。大学卒業直後、Data Science for Social Good Fellowshipでメンフィス市の都市再生プロジェクトに参加した。そこで直面したのは、住宅一軒ごとの質に関するデータの入手難、引き渡し後に停止したウェブ基盤を維持できない市の体制、予測に基づいて改修や解体を行う予算の不足だった。
技術者は課題を見つけて予測や最適化を提供すれば社会が良くなると考えがちだが、現場では制度的・組織的・政治的な困難のほうがはるかに大きいと気づいたという。博士課程でも同様の経験を重ね、研究の関心を、指標上は良く見えるアルゴリズムが現実でどう機能しないかの解明へと移していった。
形式主義からリアリズムへ
Greenは、技術的な詳細や指標を中心にシステムを評価する従来の実践を「algorithmic formalism」と呼ぶ。社会的影響が論じられても、脇に置かれるか数理的な形式化で処理されると述べる。これに対する「algorithmic realism」の核が、プラグマティズム哲学に基づく「プラグマティスト格率」で、アルゴリズムの属性を試す最終基準はそれが現実に生む影響だという考え方である。
一見当たり前に聞こえるが、問題設定、開発、導入、評価の細部まで貫くと大きな転換になるとGreenは語る。また、データサイエンティストが自らを政治的主体と認識しつつ、影響を受ける人々の視点に根ざして謙虚に行動するための具体的な手順を示すことが、本書の狙いだとしている。「注意せよ」「権力を考えよ」といった一般的な呼びかけにとどまる議論との違いだと位置づける。
公平性指標と説明可能性への批判
公平性について、Greenは、バイアスへの問題意識が広がったこと自体は成果だとしつつ、多くが公平性指標という形式主義の枠内で扱われている点を批判する。指標は実際の社会的影響を語らず、構造的に不当な使われ方をするアルゴリズムへの絆創膏になりがちだという。
代案として挙げたのは、統計的に公平なリスク評価を作ることではなく、構造的な視点で問題の源流に働きかけることである。例として、地域に根ざして暴力を防ぐ「community violence interruption」プログラムの支援や、ニューヨークの「Supervised Release」のように勾留前の拘束を減らし社会的支援につなげる取り組みを挙げ、データサイエンスで強化できると述べた。
説明可能性についても同様で、説明を付けても人の意思決定は実証的に改善しないと指摘する。そこで提唱するのが「algorithm in the loop」だ。人が監督する「human in the loop」とは逆に、既存の意思決定プロセスを中心に据え、そこにアルゴリズムを組み込む。意思決定者がどんな情報をいつ必要とするかから、必要なシステムそのものを問い直す発想である。
デューイの知性観とデータサイエンティストの未来
生成AIを扱う章では、哲学者John Deweyの知性観を参照する。知性は情報を記憶することではなく、より良い未来に向けて行動を導く道具だという見方だ。Greenはこれをもとに、ベンチマークの成績に偏った今日のAI論を批判し、機械の知性の最大化ではなく、人の判断や学習を支える方向を重視すべきだと述べる。
AIがコーディングを担う時代の職業像については、データサイエンティストは「データストラテジスト」へと役割を広げるべきだと語る。社会課題の理解、組織にとって有用かどうかの評価など、コードと現実をつなぐ役割である。
規制と、AIの「速度」ではなく「方向」
Greenは、こうした実践は自発的な取り組みだけでは広がらないと述べ、教育改革や規制の必要性を挙げる。大手テック企業は公共の利益ではなく利益や市場シェアの最大化という異なる論理で動いているという見方だ。
現在のAI安全をめぐる議論が開発の「速度」に偏っていることにも不満を示した。減速は企業の自主規制を正当化しかねないとし、重要なのはどんなAIを作るかという方向性と、民主的な統制だと述べる。データセンター反対運動や学校・職場へのAI導入への反発は方向性への異議だと捉え、学術研究や公的部門による内製開発への投資を求めた。
まとめ
本書の主張は、評価の基準を指標から現実の影響へ移すという点で一貫している。日本でもAIの導入や評価で精度やベンチマークが重視されがちだが、現場のワークフローや制度との適合を問う視点は実務者にとって参考になるだろう。また、AI規制や公共調達を考える政策担当者にとっても、速度ではなく方向性を議論するという論点は注目に値する。
文字起こし(英語・自動生成)
Your film is now ready to be shown. Good morning. I'm Justin Hendricks, editor of Tech Policy Press. We publish news, analysis, and perspectives on issues at the intersection of tech and democracy. How can you tell if an algorithm is any good? For today's guest, the answer is less than an accuracy score or other technical measure than in what a system actually does once it's out in the world. In his new book, he offers data scientists a step-by-step method for building and evaluating algorithms in the service of social justice. I am Ben Green. I am an assistant professor of information and public policy at the University of Michigan, and I'm the author of the book Algorithmic Realism, Data Science Practices to Promote Social Justice. Ben, I'm looking forward to getting into this book with you today, and maybe just want to start kind of where you start in the book, kind of a little bit of a background around your career and, you know, how you got into data science
and how you started out on this sort of intellectual journey. You talk about an experience in Memphis just before you set off to pursue your PhD. Maybe we can start there. So this book is very personal and biographical for me. It really echoes and reflects my own journey as a data scientist, thinking about the role of myself and other data scientists and algorithms in the world. So when I graduated college, I had studied math and physics and had been learning about data science and was very excited about applying those skills and methods to public policy problems, in particular local government policy, urban planning type challenges. And so my first opportunity to do that was the summer right after graduation. I worked for the Data Science for Social Good Fellowship. It was at UChicago at the time, so I got to live in downtown Chicago,
had a great time. And my team was working with the city of Memphis on a project to help them make sense of data about urban planning and urban revitalization. And I very quickly saw that the challenges of using data for good were not really technical ones in the sense of the types of things that data scientists get trained on. The first issue that we confronted was just getting access to reasonable data in the first place. It was very difficult to actually get data on the quality of every parcel, every house in the city. That's what we needed to kind of try to understand where the city should be investing money in revitalization efforts, which homes or properties should it be purchasing, rehabilitating, which ones should it demolish, what other actions should it take. So that was one of the big challenges that we confronted. Then kind of after, towards the end of the project, and as we were doing handoff, we came up with
other issues. We had created a web platform for the city to work with. And pretty quickly after the handoff, the system just, you know, stopped working. And the real issue was, you know, they didn't have the resources to actually maintain this type of system. You know, they didn't have a huge tech team to sort of take on some custom built website and maintain that for usability for an extended period of time. And then the biggest issue was just that they actually just didn't have the resources on a policy level to actually take the types of actions that the model that myself and my collaborators had built, which was meant to predict like which properties are in need of investment, what actions should the city take for, you know, the set of a set of properties within within the city. And, you know, they just didn't have the funds to actually take these actions of rehabilitation, demolishment, revitalization generally. So this was just one short project,
but it quickly woke me up to this kind of gap between the way that technologists often talk about data improving society and the realities of doing this in practice, where, you know, from a tech perspective, it often feels like, oh, we have all these skills. We can find social problems and develop algorithms to make the optimization or make the prediction or make the classification that provides the correct information. But on the ground, it looked super different, where there are all of these sort of institutional, organizational, and political challenges that are far more difficult to navigate than just figuring out, you know, how to optimize a system. And as I understand it from the book, that landed you kind of on a journey to think about the practice of data science. And you talk about this idea that you ended up passionate, maybe slightly less about developing algorithms to improve society as studying the methods for
how to develop algorithms to improve society. Is that right? I had a number of experiences once I started my PhD that echoed this example in Memphis, where I worked on building a system And I saw, hey, this is actually so much harder in ways that I have never been trained for. The challenges here are about data sharing, are about politics, are about how is an institution like a police department or a city government going to use the information that they've been given? it. So, you know, I really kind of became, it felt to me like just focusing my career on trying to build algorithms for a bunch of local governments wasn't the best use of my time or was kind of coming up against all of these questions and challenges that I didn't know how to answer. So I stepped back and really shifted the focus of my research to be about trying to actually make
sense of the real world impacts of algorithms and in particular the gap between algorithms that look good on paper and what the impacts of those systems are. So as just one example of this, a lot of my dissertation focused on human use of algorithmic decision-making tools, looking at how, you know, an algorithm can seem like it provides, like an algorithm can seem like it's a good system. It satisfies metrics for accuracy, for fairness, the types of things that data scientists tend to look for, but then doesn't work in practice because the intended users don't understand how to use the system. They don't actually want the information that it's providing. They misunderstand the information that it's providing. So that implementation angle was just one aspect of how, you know, this sort of typical technical view of algorithms misses what's actually happening on the ground. So my research became really focused on trying to unpack these types of implementation challenges
and aspects of how algorithms are actually affecting policy. So you say this book is primarily written for data scientists, for researchers who study the impacts of algorithms, but you also hold out it's for legal scholars, policy professionals who are working on AI regulation in particular, who are thinking about how to evaluate algorithms. You say a central concept that the reader has to understand is this notion of algorithmic realism. So maybe give our listeners just a quick one-two on algorithmic realism. Yeah. Algorithmic realism is the methodology that I develop and propose in the book, And it is a way of thinking about and doing data science that centers around the real world impacts of algorithms. And algorithmic realism is in contrast to what I describe as algorithmic formalism, which is the typical way that data science is taught and practiced.
So within algorithmic formalism, data science methods and evaluation revolves around the technical details. You know, what data are you analyzing? What are the technical metrics of an algorithm? And that is typically how a system would be evaluated. Often there's a discussion of social impact within algorithmic formalism, but it's typically treated either off to the side or dealt with through mathematical formalisms Algorithmic realism is an approach that guides data scientists in actually taking real impact much more seriously The core principle of algorithmic realism is this concept that I call the pragmatist maxim for data science. And for interested folks who kind of want to look at this later, a lot of this builds on pragmatist philosophy. But this pragmatist maxim for data science says that the ultimate test of an algorithm's attributes is the impacts it creates
in practice. So if we want to know if an algorithm is well-developed, if it's good for society, if it's equitable, we can't know that by looking at the technical details of the system. We need to actually look at what does this system do out in the real world. And, you know, I think this is the type of foundational idea that sounds kind of simple and banal and maybe obvious. And I think most data scientists would probably say like, oh, yeah, I agree with that. Like, I think I'm already doing that. But if you actually carry it through into the fundamental details of scoping and developing and implementing and evaluating algorithms, you actually can see some pretty these transformational shifts. You do spend a bit of time still grappling with, you know, the extent to which doing data science is doing politics or doing political engagement. It seems like you feel like there's some education to be done on that front. So broadly, this book is an
effort to help politicize data scientists and help them feel empowered to be political. One of the through lines of my work dating back at this point about a decade has been trying to help data scientists see themselves as political actors. And, you know, one of the biggest pushbacks that I get on that idea is I don't know how to do that, right? Like there are certainly data scientists who will say, I'm neutral, I don't even want any part of this. And there's still a current of that, certainly in the field. But a lot of data scientists are like open to the idea, but then say like, I don't know what that means for me. Like, you know, you're telling me I'm a political actor. You're telling me I have all this influence. Also, people are telling me to not be too overconfident and to listen to other people. And I'm feeling caught between all of these tensions of like the influence I have to make decisions, but also the fear that I'm asserting my privilege
and my voice over impacted populations. So this book is trying to both help data scientists see themselves in that light as political actors, but then help them say, here's how you do this responsibly. Here's how you can thread that needle between hubris and fear and act in a way that is oriented towards social justice while also checking yourself throughout the process to make sure that you're being grounded in the perspectives of impacted populations, people who have substantive expertise in the domain you're working in. And so a lot of the book for data scientists is kind of laying out a step-by-step process for how to do this, because I really wanted to get beyond the sort of general calls to action that I have made in my own work, and that is very common in a lot of the sort of responsible AI and critical algorithm studies work that says to data scientists, you need to be careful, you need to think about power, you need to think about social context,
but then doesn't really give data scientists a ton of ideas and steps for how to do that. And that's one of the places where I think this book can hopefully stand out and be really valuable to data scientists is to provide that sort of step-by-step guide with some theories, some examples, some directions for what it looks like to do this. So after you get through the workflows, through the implementation, the scoping, the various kind of mechanisms by which data scientists can go about their work, you stop on algorithmic fairness in particular and this kind of idea. I feel like we spend a lot of time on this podcast talking about issues around fairness, bias, you know, certainly many of the harms people most fear about AI systems and the conversations that we're having right now about AI and the various harms it might produce revolve around questions around fairness and bias and the extent to which these systems may discriminate against people
and certainly against marginalized groups. I don't know, where would you pick up there on this question? What does this algorithmic realism approach tell us about how to think about fairness or bias? One of the things that excites me about algorithmic realism is that it provides us with new ways of thinking about the challenges that the fields of data science and AI face. So, you know, I talked earlier about the sort of step by step guide for how to implement a project. And that's a big piece of it. But when you kind of zoom out from that framework, what you get with algorithmic realism is also a new way to think about these challenges and a new way to approach data science projects. So within the book, I talk about that through the lens of algorithmic fairness and explanations, which are sort of these two big challenges that data science has been confronting and trying to figure out. And the broad angle on fairness is that the field has kind of just gotten it wrong in
a lot of ways, again, by repeating the mistake of algorithmic formalism. So on one level, you know, the fact that so many data scientists are aware of these issues of bias and discrimination is great. Like that's a huge win from where the field was 10 years ago. But they've done that and then they've sort of tried to engage with those questions through the lens of formalism, which means that the primary way that fairness gets dealt with is through, you know, various fairness metrics, various ways of statistically testing an algorithm for bias. But these fairness metrics fall really short. They don't really tell us about the actual social impacts of the systems. They often become a sort of band-aid solution to systemically unjust algorithms that are used in bad ways. So my solution is not to say, hey, I have a new fairness metric, right? I have a new way to optimize algorithms for fairness.
Instead, what I say is, actually, if you care about injustice and you're trying to deal with an issue like racial injustice in the criminal justice system, the answer is not to create a statistically fair risk assessment. There are actually a bunch of other things that you can do that take a more structural lens on this problem. So I talk about things like trying to support using data science to support programs that would more sort of structurally alleviate the reasons why fairness is such a problem in the first place, which is about, you know, inequality in people's characteristics. Right. Like there are significant disparities in the criminal records and social outcomes of people across race and class lines, which is what a lot of these algorithms are picking up on. But then also the impacts of algorithms, the way that an algorithm gets used in practice is often quite unjust, right?
So you have a pretrial risk assessment where if someone is deemed to be high risk, the response in the vast majority of cases is to say, OK, great, let's lock that person up so that they don't commit another crime. That's not an inevitable thing that needs to happen. So what I look for and describe are places where you could use data science to counteract some of these things. I talk about, for instance, community violence interruption programs where you have people embedded within marginalized communities that are working directly to try to prevent violence, particularly gunshot violence. And there are a lot of these programs that are, you know, underfunded, short staffed and really trying to figure out, like, how can they best support the people in their community? There are a lot of ways that data scientists could come in and help them say figure out you know which which of our people in our program need the most support which ones could actually benefit from some additional outreach These are all things that could actually just reduce the crime risk of people within these communities making them less susceptible to being deemed high risk in the
first place. You can also use data science to help programs that are trying to take different outcomes or different interventions in response to these predictions. So I point to one program in New York City called Supervised Release that tries to get more people out of pretrial detention and into, you know, a release outcome with more social support. Again, that program has a lot of challenges that data scientists could apply themselves to and help that program become more effective for more people, especially some of the highest risk people who are kind of at the center of debates about this program. So, you know, broadly, what I would say here is the answer is to really dig in for data scientists and dig into these situations where it's like, hey, I'm concerned about racial disparities in pretrial outcomes. And it's like, great, you should be concerned about that. But rather than just jumping to the sort of obvious decision point,
take a more zoomed out perspective, try to understand where that problem comes from and then look for local institutional actors or programs that are dealing with the structural challenge and then see how you can go help those people. And presumably this is also where some of your thinking on, you know, kind of algorithm in the loop, human in the loop, kind of human algorithm collaboration, as you refer to it, comes in. Why is that such a roadblock for people who are trying to, you know, benefit society by applying technology? This is one of the biggest challenges that has been kind of driving me for many years and actually helped me do a lot of the thinking that pushed me into writing this book in the first place. So, you know, there's this common challenge of data scientists creating an algorithm or creating an AI system and saying like, wow, this should really solve the problem because this provides really good information to make better decisions.
But so often that doesn't happen because the intended users don't actually interact with that system the way that the data scientists assumed. And that's not to say that, oh, the users are doing something wrong. It's to say the data scientists didn't really think about the context of implementation. They sort of often work on this very abstract vision of how a decision-making process actually occurs and then build a system based on that. And then the mismatch between those assumptions in reality means that the system isn't that useful. And, you know, even today we can see this with some of the latest generative AI tools where like, in theory, it's supposed to help with education. But actually, it's reducing student learning. It's reducing, you know, lawyers are putting forward a bunch of briefs with made up citations, things like that. Right. Because people are just relying on these systems. Now, the sort of common approach to deal with this challenge is this framework of explanations,
right? Saying, well, we're going to, sure, we won't just give the decision maker like a judge or a caseworker. We won't just give them an algorithm. We'll augment that system with an explanation. And so there's all of this work trying to build explanations. And again, it's very formal, focused on the statistical details of what a good explanation should look like. But when you actually go and test that empirically, these explanations don't work. They do not improve people's ability to use the algorithm to make better decisions. So again, this is a situation where algorithmic realism can say, okay, well, the issue is not that the explanations are badly developed. Like, we're actually just kind of running down the wrong path. We're chasing the wrong rabbit hole. and the bigger task that algorithmic realism reframes us around is this idea you mentioned of algorithm in the loop decision making where what we really need to focus on is the idea that
we are not developing systems that people are then supposed to oversee we really need to center the human decision making process that already exists and then we are inserting the algorithm into the loop of that. And that's sort of this play on this common idea of human in the loop, which often just means like a human overseeing a process. So for developing an algorithm, what we need to do is really dig into, like, what is the workflow of decision makers? What information do they actually care about? When do they need that information? What types of things are they looking for? And how can I design for that? And so some of the stories I talk about in the book, you know, they're not saying, oh, we gave this person a better explanation. It's saying, actually, we just rethought what type of system we even need to provide that person in the first place. So in the seventh chapter of this book, you get into artificial intelligence, generative AI in particular,
start to ask some questions about how the technology that everybody's talking about at the moment, that may change the practice of data science, certainly change the sorts of approaches and questions that data scientists are asking. And you mentioned the pragmatists. You get on to John Dewey, you know, talk a little bit about his idea of what intelligence is or maybe what intelligence is for. Maybe we could pick up there with a little Dewey. I talk a lot about Dewey in the books. He's definitely a big influence on this one. And he talks about intelligence in a way that is just completely foreign compared to how most people in AI talk about intelligence. So broadly, pragmatism is about making knowledge practical. The pragmatists were philosophers who were concerned about how their field was so concerned about dealing with what is the essential truth of the world? What is the foundational, unchanging principle of justice?
And Dewey and other pragmatists were saying, like, this is just completely the wrong way to think about knowledge and philosophy and ideas. Like, this is we should be thinking about knowledge as practical instruments for problem solving. And these ideas about essential truths are just a fallacy and are actually leading us in the wrong direction. And so he writes about intelligence as something that is instrumental. The purpose of intelligence is not to, you know, memorize kind of like memorize information, but to actually guide action in service of a better future. Intelligence has a purpose. It's not something that can be dictated or evaluated just by, you know, someone's performance on a test. And so, you know, it's very different than how we tend to talk about AI today, where there's such an emphasis on these benchmarks. Again, we can see, you know, the formalism coming up, where the way that we tend to talk
about AI is what is its performance on these benchmark tests? And we have all these metrics for it. But those benchmark tests often have very little to do with, you know, practical tasks that people are actually going to be carrying out in their day-to-day work. and even less to do with whether those tasks are socially beneficial. So what I talk about is, you know, we need to think about the intelligence side of AI, less about the intelligence of the machines and more about the intelligence of people and rethink our approach away from trying to, you know, these ideas of super intelligence or trying to make an AI system that's maximally powerful. and instead think about like, what are ways that we can actually use this technology to augment people's performance, people's ability to make decisions, take better actions, integrate knowledge and so on. And so this builds very neatly off the algorithm in the loop conversation
where the idea is we need to build systems that are actually helping people learn and make decisions rather than thinking that we should just be building the most powerful possible tools and then assuming that somehow that will lead to benefit. So that changes the job of the data scientist, I assume. You talk about this idea that data scientists may become more like creative strategists. Yeah one of the big questions facing data science today is you know what is what is the future of this profession look like given that AI can now do a lot of coding So I think of data science now as having like two different big challenges that it's confronting. So the kind of like ethics challenge is one that it's been dealing with now for about a decade. And it's made some progress, but still it's like a big question in the field of like, how do you make these tools responsibly? And, you know, nowadays that's become in some ways more pressing with all of these AI safety concerns.
But then the other question for data science is not about the building of the tools, but about what it means for them as workers. And I think that algorithmic realism provides a way to deal with both of these things. So I've already talked quite a bit about how it provides a way to integrate ethics into practice more. But it also can help data scientists have a identity that's better suited for a world where a lot of the actual hands on coding is done by an AI system. Because a lot of what algorithmic realism provides is a way to think about connecting code with real world challenges. So the types of things that I'm pushing data scientists to do is not, you know, here's a bunch of new technical methods, but to say, here's how you should understand social problems so that you're building a system that's actually worthwhile. Here's how you should evaluate your systems after you've built it so that you can see, has this been beneficial?
Do I need to scrap it? Do I need to change it in some way? And all of these tasks are incredibly important for making algorithms useful in practice. But they also give data scientists a clear sense of what their role can be that goes beyond just code. And so, you know, I was saying earlier, I'm trying to politicize data scientists. And more generally, beyond the political side of this, I'm trying to give data scientists a sort of wider sense of what their expertise is and what their professional identity can be. So I come to this role, this idea of thinking of themselves as data strategists rather than data scientists, where it becomes less and less essential nowadays to be the person who's actually writing the code. But that doesn't mean that data scientists are useless. You actually, it becomes more important than ever to have someone who can help an organization or a team figure out what is the code that we need to write? How do we implement that code?
How do we test that code? And again, not just test it in terms of technical details, but test whether it's been useful for us as an organization. So I think that sort of the training and mindset of data scientists needs to change a lot to take on more of that connective tissue role. That's less about writing every line of code and more about drawing the connections to organizational strategy and problems. Is there a role for regulation and policy in sort of pushing algorithm realism as a practice? Yeah, this is definitely not something that can come just on a voluntary basis. So I think of the shifts that are needed in data science's AI kind of along these two parallel complementary dimensions. One is methodological. You need data scientists to actually have a set of practices that they can employ to build algorithms that are more reliably good for society. That's kind of the core focus of the book.
But you can't get there without institutional changes that incentivize those practices more and more. Because, you know, we've we've all seen how particularly within the big tech companies, they're not trying to promote the public interest. Their problem is not that, you know, by and large, they're trying to do good for society and then just lack the methods. They're just operating on a fundamentally different wavelength in terms of, you know, trying to maximize profit, maximize shareholder value, market share, all of that. So there is a there is a strong need for educational reform, for regulation that both changes the culture of data science, but also then can mandate, you know, government bodies and tech companies to actually kind of take some of the best practices that I am talking about. So listen, we've talked very generally here in the abstract. I'm interested in, as you kind of think about the headlines,
and I'm talking to you in the midst of this, you know, AI safety, AI risk kind of freak out that has extended for now, you know, a matter of weeks. When you look at the headlines, how do your ideas reflect back to you when you think about how AI is actually being implemented in the world? I think one of the big things that frustrates me about the conversation right now is that it still feels very one dimensional in the sense that the conversation we're having is about the pace of AI development. So often there's this sort of false binary that the industry will tell us of like, oh, if you want to stop, you know, if you want to regulate us or cut down on any of our practices, you just want to stop innovation. So that's like a very extreme version of this frame. but today it's still even in the conversations on pace it's still kind of operating on that framework it's saying okay there is a predetermined path that we're going down and the question is just
how quickly do we get there and you know sure slowing down and putting in some guardrails is good but fundamentally what i'm interested in is laying out a different direction for how ai is developed and what types of AI systems we actually get. And so slowing down the pace in many ways feels like a path toward legitimizing these companies and allowing them to do more self-regulation than actually getting us to an ecosystem of AI development that actually works for the vast majority of people. So, you know, I think there needs to be a lot more conversation about democratic control over AI systems, right? PACE is still very much an elite-driven process where the average person has still really no say in what that looks like. Meanwhile, on the ground, we're seeing lots of meaningful pushback, not just to the pace of AI, but to the direction that AI development is going, right? Data center protests, families blocking the implementation of AI in schools, workers pushing
back against the implementation of AI in their workplace. These are not people concerned about the pace of AI. They're concerned about the fundamental direction. And so we need to open up pathways and ecosystems where more public interest oriented AI can be developed through the types of practices that I talk about in the book and that other aligned scholars have also put forward in their work. And so I think that means, you know, there's not a sort of single silver bullet answer here, but certainly shifting more towards, you know, funding for academic research, particularly academic research in the vein of, you know, community-centered work, human-computer interaction, and often as well working on public sector developed systems, you know, putting more resources into the public sector to build technology in-house that works for them, rather than needing to rely on, you know, contracts with companies like Palantir and
Deloitte and OpenAI to have their own systems in place. Well, we will see whether, you know, maybe more of the world adopts a little bit more of the approach that you're recommending here. in a book called Algorithmic Realism, Data Science Practices to Promote Social Justice. Ben Green, thank you very much. Thank you. Always great to chat with you, Justin. That's it for this episode. I hope you sent your feedback. You can write to me at justin.techpolicy.press. Thanks to my guests. Thanks to my co-founder, Brian Jones. and thank you for listening.
番組の概要欄(原文)
How can you tell if an algorithm is any good? For University of Michigan assistant professor of information and public policy Ben Green, the answer lies less in accuracy scores, fairness metrics or any other technical measure than in what the system actually does once it’s out in the world. In his new book, Algorithmic Realism: Data Science Practices to Promote Social Justice (MIT Press), he offers data scientists a method for scoping, building and evaluating algorithms in the service of social justice.