Revolutionizing AI: Tackling the Alignment Problem
Navigating AI Risk with Nate Soares and Deger Turan
Up next
All episodesDeSci: Blockchains are Powering Science 3.0
Living Forever? The Science and Ethics of Longevity
From Lab to Life: Exploring Synthetic Biology
Zero-Knowledge AI: The Frontier of Cryptography
Can We Fire Gary Gensler? with Rep. Warren Davidson
DEBRIEF - Is Crypto's Bear Market Almost Over?
178 - Is Crypto's Bear Market Almost Over? with Raoul Pal
ROLLUP: BTC ETF Race Heats Up, zkSync drops the ZK Stack, OP going ZK?
Inside the episode
In this episode, we delve into the frontier of AI and the challenges surrounding AI alignment. The AI / Crypto overlap at Zuzalu sparked discussions on topics like ZKML, MEV bots, and the integration of AI agents into the Ethereum landscape.
However, the focal point was the alignment conversation, which showcased both pessimistic and resigned optimistic perspectives. We hear from Nate Sores of MIRI, who offers a downstream view on AI risk, and Deger Turan, who emphasizes the importance of human alignment as a prerequisite for aligning AI. Their discussions touch on epistemology, individual preferences, and the potential of AI to assist in personal and societal growth.
Timestamps
0:00 Intro
1:50 Guests
5:30 NATE SOARES
7:25 MIRI
13:30 Human Coordination
17:00 Dangers of Superintelligence
21:00 AI’s Big Moment
24:45 Chances of Doom
35:35 A Serious Threat
42:45 Talent is Scarce
48:20 Solving the Alignment Problem
59:35 Dealing with Pessimism
1:03:45 The Sliver of Utopia
1:14:00 DEGER TURAN
1:17:00 Solving Human Alignment
1:22:40 Using AI to Solve Problems
1:26:30 AI Objectives Institute
1:31:30 Epistemic Security
1:36:18 Curating AI Content
1:41:00 Scalable Coordination
1:47:15 Building Evolving Systems
1:54:00 Independent Flexible Systems
1:58:30 The Problem is the Solution
2:03:30 A Better Future
Resources
Nate Soares
https://twitter.com/So8res?s=20
Deger Turan
https://twitter.com/degerturann?s=20
MIRI
Less Wrong AI Alignment
xhttps://www.lesswrong.com/tag/ai-alignment-intro-materials
AI Objectives Institute
Transcript
welcome to bankless where we explore the frontier of Internet money and internet finance and today on this episode of our zoox aloe series we are exploring some New Frontiers New Frontiers and new technologies all of which were poised to completely revolutionize the world and change everything about the operating system that Society is currently running this nation today we are exploring the frontier of AI which is actually a frontier that we've already been exploring on Bank lists so if you've been listening to our other AI episodes these will make you feel right at home AI had a big week at zuzalu the AI
crypto overlap everyone knows it's huge and it seems like such a massive Frontier that people don't actually know where to start with it zkml or machine learning models and data that's verified by zero knowledge cryptography was a huge topic of conversation and you'll hear about that in our cryptography episode with Daniel short Phil Diane at AI week gave a killer talk titled Mev for AI people which was this Giga brain presentation about how Mev bots in aggregate kind of presents this
omnipotent omnipresent artificial intelligence and since Mev has been decently corralled and contained maybe we can learn a thing or two from the Mev industry in our approach to managing AI risk there were conversations as zuzalo about how AI can put the autonomous back into Dao's and how AI agents could soon be roaming the ethereum landscape shoulder to shoulder with all the human players out there but mainly azuzalu the AI conversation inevitably converged into the alignment conversation of which you will find two
flavors here in this episode one strongly pessimistic and then the other characterized by this resigned optimism that is uh prevalent throughout all of zuzalu's Frontier Tech challenges unimaginable rewards blocked by seemingly insurmountable obstacles up first in this episode we have Nate stories who is the executive director at Miri the machine intelligence Research Institute of which elieza yudowski founded Nate's perspective on AI and AI risk is definitely Downstream of Eleazar
so we pick up where Bank was left off with Eliezer and bankless Nation it's dark but nonetheless Nate admits that it's less dark than it was a few months ago now that the world is waking up to the potential risk that AI brings to this world following the conversation with Nate is dagger Taran who is charging into the AI Frontier with his head held high with a clear path forward for himself Digger believes that the a AI alignment problem is actually just Downstream of human misalignment and that we actually won't be able to align AI until we align ourselves
this conversation has to do with epistemology what is truth individual preferences and how AI models can help us become the best versions of ourselves because if we become the best versions of ourselves with the assistance of some AI tool we can collectively produce the best versions of our communities and if we do that then our communities can coalesce into the best versions of society all aided by truth-telling AI agents who can help humans navigate through our chaotic world of social organization and politics and social media really a fascinating conversation
that is actually pretty proximate to our conversation with Tim Urban that we had not too long ago really excited for you to listen to these conversations Bank locations so let's go ahead and get right into it but first a moment to talk about some of these fantastic sponsors that make this show possible Kraken Pro has easily become the best crypto trading platform in the industry the place I use to check the charts and the crypto prices even when I'm not looking to place a trade on Kraken Pro you'll have access to Advanced charting tools real-time Market data and lightning fast trade execution all inside their spiffy new modular interface kraken's new
customizable modular layout lets you tailor your trading experience to suit your needs pick and choose your favorite modules and place them anywhere you want in your screen with Kraken Pro you have that power whether you are a Season Pro or just starting out join thousands of traders who trust Kraken Pro for their crypto trading needs visit pro.cracken.com to get started today metamask has something new introducing metamask portfolio metamask portfolio is the best way to view your crypto portfolio from a holistic level see everything across all the trains all at
once in your portfolio metamask will report the aggregate value of all the Assets in your metamask wallets and even the other wallets you import too but metabask portfolio isn't just a passive portfolio viewer it is a place to do all of the money verbs that make defy so powerful you can buy swap bridge and stake your crypto assets so not only is metamask the easiest place to see your wallets in aggregate but it's also a powerful battle station for all of your defy moves so go check out your metamask portfolio because it's waiting for you to open it up check it out at
portfolio.metamask.io arbitrim1 is pioneering the world of secure ethereum scalability and is continuing to accelerate the web 3 landscape hundreds of projects have already deployed on Arbitron 1 producing flourishing defy and nft ecosystems with a recent addition of arbitrum Nova gaming and social daps like Reddit are also now calling arbitrum home both arbitrim1 and Nova leverage the security and decentralization of ethereum and provide a builder experience that's intuitive familiar and fully evm compatible on
arbitrum both Builders and users will experience faster transaction speeds with significantly lower gas fees with arboretum's recent migration to arbitrary Nitro it's also now 10 times faster than before visit arbitrim.io where you can join the community dive into the developer docs Bridge your assets and start building your first app with arbitrum experience web 3 development the way it was meant to be secure fast cheap and friction free Bank of nation we are here with Nate stories and we are starting the AI week here at suzalo and Nate is a AI
researcher is that how you would call yourself alignment researcher uh I would say I'm the executive director of the machine intelligence Research Institute um these days it's less research than I'd like uh I have done alignment research before okay uh can you explain that that institution uh The Institute uh what is that uh so I didn't found it um it was founded by eleazaredkowski I don't know exactly when maybe around 2001 um fun fact uh it was originally founded
it was originally called The Singularity Institute and it was founded uh because Eliezer wanted to make a GI as fast as he could uh and then uh along the way he realized that uh it doesn't go well by default and it doesn't go well for free and so then the organization pivoted to uh trying to make this AI stuff go well and for many years The Institute uh did some research did some like field building
did some awareness raising and so on and so forth uh until uh around 2012 2013 when they pivoted to Pure technical research and this was related to uh some of the field building some of the awareness raising moving to other uh groups as the field got a little larger and I got involved uh when uh they pivoted to the technical research more exclusively and so I was originally involved as a technical researcher and then uh when the previous executive
director left I was The Heir Apparent okay sorry the machine what's the name of the Institute machine intelligence Research Institute AKA Miri Miri okay so it sounds like Mary has its own trajectory of itself that probably runs in parallel with like human understanding with machine learning at Large uh perhaps but also perhaps quite compressed I think uh I don't know the exact years I wasn't around then uh but I think it was only a couple of years of
work uh before Elijah was like hey wait uh this could get tricky right okay so I actually didn't know that about eleazer that first who was like AGI as fast as possible and then he was like whoa whoa whoa ADI as slow as possible uh yeah I think I think I'm not sure as slow as possible but like AI done correctly AGI done correctly um I think you know we were hoping uh for a long time like one of the reasons we do technical research is that uh like you can often like often your levers just like solve the problem right
like screws slowing people down that like pisses people off it's like like slowing people down is sort of a last resort um the the original Hope was can we just like solve the problem in time uh it doesn't look like we're on track to solve the problem and it looks like we have less time than you know I was hoping back in 2014 uh and so uh I think it is with great sadness that uh people like a laser and I are now saying we need we need time we need more time can
you kind of walk us through your your own trajectory how do you how did you become the executive director at Miri uh well if you want to go back far enough um I uh at a pretty young age uh uh realized that the world was not very well organized uh and wasn't uh I don't know I was I was in a civics class and up until uh that particular civics class
um I had some intuition that like there were lots of problems in the world but people were sort of trying to fix them and the reason there were still all these big problems in the world was that we didn't have the technology we didn't have the uh like we were still a young race we were still a young species we hadn't like matured to the point where we could fix these issues this was sort of like an implicit wordless intuition rather than a than a conscious belief and then you know I started learning how the US government works and I was like oh God it's like run by a bunch of monkeys like
it's it's like monkeys invented like monkey systems and it's all like working about as well as you'd expect if it was like invented by people who had no idea what the hell they were doing and then like allowed to run for like hundreds or sometimes thousands of years and just like Korean off into various so I was like okay like obviously from the molok and crypto world we would call this coordination coordination failure totally yeah so uh so I was very uh interested in solving coordination failures uh and more generally in making
the world a better place um and I try my hand at various versions of that while I was like pretty young and bad at things uh and uh uh you know it's it's a hard it's a hard problem none of it worked um and uh the the sort of security circuitous route is that uh ultimately I got a job in Tech while I was like trying to find ways to really move the levers on that problem decided to uh donate a decent
amount of money to good causes uh to sort of keep me honest about actually trying to make the world a better place uh was trying to research where the best places to put money were um donated to some uh like Global poverty type Charities and I started bumping into these arguments that like maybe this AI stuff uh is actually one of the biggest places intervene in the world I sort of read up on it um there were a couple other uh factors in my life that were also causing me to notice
this uh this AI thing I read up on it and I was like oh geez like this is just obviously you know I was looking at the wrong problem uh like the coordination problems are big and they're real but like this AI thing you know like Humanity like lives or dies and gets like a great future or no future depending how they say anything goes and I just like completely missed this problem um for like eight years of thinking myself as trying to make the world a better place and going for the for the
heart of the problems but um so when you ran into the AI topic did you see for did you first see AI as a solution to all of our coordination issues or AI is a problem for all of our coordination issues or did you see it simultaneously at the same time uh a little bit of both um I was I was maybe somewhat primed towards uh understanding some of the issues with AI due to my work on coordination uh problems uh it's like slightly embarrassing but uh I was like
working on like various coordination mechanisms uh that could address the sort of concerns people had at the time like how can a well-coordinated society without coercion address for example like uh concentration of wealth uh in ways that the society as a whole doesn't like and you can you can set up various coordination mechanisms of like uh like uh yeah you can sort of try to think about like what are non-course of ways that a society as a whole can like try
to both have a market system and not let it get out of control in certain ways and while while like messing around with like toy models of this and like attempts to like prove certain theorems I just like couldn't get some of the results I wanted and it turns out that I couldn't get some of the results I wanted because Nothing Stops one actor from being powerful enough that they can just run away with everything right and this was sort of like it was one of those issues where I was sort of like well you know I can get it to work in a lot of cases but I can't get it work in all cases and then with the AI stuff I was like oh that's why like
uh can you can you elaborate on that like why why does the AI uh why is AI like the kernel of the issue uh I mean AI is a is a uh version of that particular issue but like fundamentally no matter how good your Market coordinations are or your Market coordination systems are on earth uh like if somebody has the raw technological power to like uh for example uh get what we call in the in the business a decisive Advantage so uh
like uh like maybe the easiest thing to imagine given that we already know that you know trees are machines that turn dirt and sunlight uh into more trees by stripping carbon out of the atmosphere and building wood we know that like nanotech's possible if you imagine something just like guess the nanotech before everything else and it can just like reassemble you into a uh more willing trade partner that asks for less of the gains from trade suddenly all of your coordination
mechanisms that were like Market based and non-coursome or whatever melt before this thing and like am I saying that literally happens or literally like is in a market framework not particularly but you can sort of see how like I was sort of like trying to put a like collaborative agents interacting framework on a like physical reality where it's just a fact about the physical reality that like things with a sufficient technological Edge just uh can work the table with you
uh if if they have too much of an edge and if they don't care about you simply put is this kind of are you just combining moloch problems and the magnification is pretty familiar with molok problems we've done a lot of content on moloch you combine moloch problems with exponential technology and then you arrive at some sort of like logical end games where humans get their atoms repurposed is that more or less the simple articulation uh it's not a bad summary I wouldn't use exponential in particular I'm not I make no strong claim that it's an exponential curve my guess is that uh
it's it's not and it's worse um and uh you know much of the issue here is if you make something that uh is optimizing the world and it's optimizing the world towards some Target that doesn't have concern for you in it um like I would have I would have much fewer qualms about like uh I would have basically no qualms about uh like human technological development I sort of am very optimistic about uh
Humanities better natures and Humanity like being able to figure out uh like how to make the world more like we would want it to be upon reflection and if we were wiser and rather than like locking ourselves into totalitarian dystopias which I think totally could happen but like if we can just like ramp up human like intelligence and capability and so on without like accidentally killing ourselves I'm like pretty bullish on humanities prospects
and so it's not it's not so much like oh no Technology's coming it's coming too fast we won't be able to handle it it's more like oh no we are like on the brink of building optimization processes that optimize the future much harder faster better than we can and they're optimizing it to a place that has no room for us and so it sounds like AI is one way and that might happen but you're also saying the way that you're talking it sounds like there's other ways in AI or not when the same
hyper optimized future that's not optimized for humans could play out without AI uh like I sort of expect we're going to get to Super intelligence one way or the other um AI looks to me like one of the like basically the only feasible route modulo like if if Humanity can't come together and coordinate to take some other route I think other routes like whole brain emulation uh are probably preferable or to be clear um I'm no carbon chauvinist and I uh
very much want to live in a future with like artificial friends where those artificial friends have like very different sorts of like desires and goals and objectives from me uh I'm not like Humanity must keep an iron grip on the future I want like space for aliens I want space for uh artificial Minds yeah other kinds of Life uh the like concern here is uh like building a mind that doesn't care for life that doesn't
care for fun that doesn't care for uh like uh diversity of experience and like interesting uh arcs and like Cosmopolitan value and like broad inclusive uh like good times and I think that we are in fact uh uh barreling towards that Cliff edge of making something that uh fills the universe not with weird valuable stuff
but with uh non-valuable stuff okay so you've been thinking about these problems for a long time when did you start at Miri what year was that uh that was uh 2014 that I was hired um I once I noticed that the problem existed uh I donated I think sixteen thousand dollars uh which I think at the time put me in the top 10 public donors list uh which uh
uh and they were like congratulations you're now in the top 10 public donors list and I was like what like and they're like you know we're doing we're doing I don't remember the exact amounts but they're like we're doing our fundraiser for 200 000 for our yearly budget this year uh 100 000 of which we're trying to raise in the community and like 100 000 which is like matched from another donor uh and you know it's like three weeks into the four week fundraiser and they raised like 20K of it and I was like oh God like this is worse than I thought like oh my God
that's hilarious that and I was in 2014 or was that that was 2013 2013. yeah so I donated precipitated your arrival actually working at Mary that's right so I donated more money um because I had known it was that bad uh did you end up funding yourself you're on salary uh I I took I took a very big uh pay cut uh moving from Google to Mary um uh and uh I sort of you know I I was expecting not to be very skilled at working on these issues
and you know maybe I'm not there have been a lot of people uh on these issues but at the time uh they uh I was like how how can I help and they were like well maybe if we go to the math you can like come to Summer workshops and so I I came to another workshops and then a few months later they were hiring me and then a year later they were saying can you run the place um so uh it uh I I am largely in this field by Dent of and and this position by identif
showing up early sure uh and like for the love of God people more skilled than me uh like by all means come replace me right so that that's kind of what I was I was leading us into so that was 2014 when you started it's now 2023 so you're almost there for a decade now yeah um now ai's having a moment um very much spurred by chat GPT all of a sudden crypto podcasts are talking to AI people um what is that trajectory like so as somebody who was immediately compelled
by the problem at its very essence so far long ago now fast forward to where we are now and like kind of the problem seems to be on the horizon I don't know how close it is I don't think anyone does it's kind of the problem but like here we are nine years later and now many many many people are talking about it can you just talk about that experience yeah uh it's it's heartening um uh one one thing that I have really enjoyed about it is I've spent many years having
conversations with people in the field uh many of whom sort of don't really want to hear that uh their work by default is uh like barreling towards destruction uh and so I have like these long conversation trees like I have rejoinders to all sorts of counter arguments uh and when I go into these discussions with with uh you know people on the capability side of things uh I have like all sorts of responses prepared and I sort of am like ready to
like go down this long decision tree and then I sort of like nowadays many more people are noticing the issue and uh you know I was invited here um and I think crypto people would actually pretty much really resonate with that where like we have to explain you know Bitcoin uh 21 million heart cap we have to explain all these things like proof of work and like the conversation trees that we have to go down uh you we've all we've built out those like in innate responses those like spinal reflexes and then lately move moving into 2022 and 2023 fewer of those things
we have to explain especially as we just printed out a bunch of money and for covid semi-checks like all of a sudden we have to explain the concept of scarcity a little bit less and so it kind of sounds like a similar experience that the AI people have yeah totally like now I go to people who aren't in the field and I'm like ready to go down all these decision trees and they're like so what's the issue and I'm like well in the most basic sense here's the issue and they're like oh yeah that seems rough I'm like oh man this is such a different conversation I mean that's the first step right the first step is education yeah and then also acceptance of the problem I could imagine for so long you
were saying hey like people would ask you hey what are you working on Nate and you'd be like Oh I'm working on AI alignment and then people are like why the why the hell are you working on that yeah they're like oh that's weird or they're like is that some weird Terminator thing yeah uh and it's been it's been nice to sort of uh I sort of I sort of think that uh like a lot of the basic issues have been pretty obvious the whole time and that we're now seeing people who like don't have uh distorted incentives noticing the issues
um but it's really quite heartening to see uh I don't know where it will go but uh it's been it's been it's been nice to see people starting to notice uh that you know this is a real thing it's really on the horizon like you said I don't think we know how far it's very hard to it's very hard to predict um at least with Precision uh but it has looked to me like one of the biggest issues facing Humanity for a while and it's very nice to see others
start to notice that as well so that leads me to the the question of just like how optimistic you are and I'll ask that in in two phases first the the same question that we'll we ask both Eleazar and also Paul Cristiano I was like all right what what are your chances of Doom what what are your what are your chances of the worst AI problem being the worst the worst the version of itself I mean worst version of itself I think is very is very hard to get but the version we're like we all die like our face worse than death but like the version where we all die I think this is
pretty likely uh I think this happens by more than 50 oh definitely um Paul Cristiano gave us 10 to 20 so you're you're saying uh more more than 50 uh my understanding of Paul is that he has 10 to 20 on the scenarios that I think are like AI takeover and higher probabilities than that on like Humanity completely disempowered um I'm definitely uh I'm definitely more pessimistic than Paul on these counts uh like I would say that on my models and
visualizations on my understanding of the problem there is very little hope uh and most of my hope comes from me being wrong somehow uh and so my probabilities on this destroying everything I know and love are like as high as my probability like they're about as high as my probabilities can go given uh like the the fact that I may just be totally wrong and hopefully am okay so you're you're pretty close to the Eleazar side of things which is like 95 to 99 Doom yeah I mean I think uh
99s are hard to get uh uh but like uh but there's also a difference between like what does the world look like as I see it the world the world looks as I see it like like the place the like as things seem to me or just like you know within a rounding over 100 uh and the the difference between that and my betting odds is and like hopefully the world's not as it seems right yeah so what you're saying is like we don't the the nature of the AI
problem is just a lot of we don't knows and so what you're saying is like the reason why you maintain some level of optimism is because there's like a white swan event that's possible that could save us yeah and you know I I have a bunch of I've thought a lot about various parts of this problem and uh I have you know uh various guesses as to where white swans are more or less likely and for instance uh I it looks to me like the white swans are less likely in my unknowns about Ai and more likely in my unknowns about how humanity is going to
react to the problem um although there are still some unknowns in how AI goes where there could be white swans do you remember when you were first um working on this problem uh I know you weren't as skilled or as knowledgeable back in 2014 to 2017 when you were first working on this problem but what was your level of optimism or pessimism back then and like how has how has your attitude towards the problem shifted over the last Almost decade that you've been working on this uh you know it's gone up and down uh I've rarely had double-digit odds of survival
um but I have I have had double digits of survival when I've been explicitly quantifying and you know most of these numbers are like coming straight out of my butt I let like one put too much but by definition everyone's numbers are coming out of their butt and that's kind of like yeah but some people there's no alternative yeah you're all right yeah uh and I don't I don't spend a lot of time worrying about specific numbers like you know once once it's less than three percent chance we get good outcomes it doesn't it doesn't affect my day-to-day I'm not like staying up trying to calculate significant digits
here I'm like man Humanity does not look like it's up to this sort of task this doesn't you know I've seen Humanity try to coordinate um and yeah we kind of so one thing we have not figured out yeah uh and for the record the reason that uh my uh the the way that I managed to have like high probability this is tricky is not that there's any one part of the puzzle that looks to me insurmountable uh that you know humans are pretty good at solving problems when they put their minds to them uh the way that like the
reason that I'm like pretty pessimistic here is uh it looks to me like there's a bunch of different ways for things to go wrong and there's a lot of things that need to go right for things to go right like um like you not only need to solve various technical challenges you need to have uptake of the Technical Solutions in uh the relevant organizations those organizations need to be able to bureaucratically recognize the difference between a real solution and a fake one uh you need to have them like carrying it all which is not even a fight that like we've won yet there are you know you have like the the heads of
labs that like Microsoft and Facebook like poo pooing a lot of these issues um and so there's like five six seven needles and like this so I want to combine two metaphors where like the stars need to align except the stars are needles that we also need to thread and like that's and that we need all of those things to happen and what you're saying is like that's that that window is small that's that's where you get the uh the difficulty from and to be clear um uh I think it would be a fallacy to say like look I
can give you like six things you need to do and like what's the chance you can get all of them that sort of reasoning doesn't really work like if if I line up all six and then the one that I assign least likelihood two happens probably I underst likelihood probably underestimated the correlation like uh like these are not independent events right like if we can solve the hardest of these issues whichever one that turns out to be probably it's because we turned out to have coordination skill or competence and so on like I'm not saying you can drive the probabilities arbitrarily Low by the fact that I can like line up a bunch of hurdles I'm more
saying it it sure seems to me like there's a bunch of hurdles man and like each of them like like well not all of them but like many of them have a character that like Humanity hasn't really faced before and this all adds up to me being like man I'm like single digit uh probabilities of survival here so with this new or maybe for first surgeons of interest from to the AI problem now things probably thanks to chat gbt thanks for the problem itself how has that shifted your optimism uh if
at all uh it's uh uh I mean it's I I feel a little hopeful about it I feel like a spark of hope here uh it doesn't uh it doesn't like shift my probabilities on the ground too much um like uh this is like a really dumb model but like if you imagine having like three variables each with a 100 100 chance uh and success is like uh multiplying them