Okay, so hello everybody, nice to see you all. So my talk is going to be about AI safety. Are we playing Russian roulette? And is this presenter working? No? Ah, this way. Okay. Disclaimer. So this talk is about the risks of AI. It's supposed to be scary, hopefully. Hopefully I managed to scare you a little bit, because it's a bit scary. Then, okay, we should talk about it now before it's too late and all of this. If you have objections, questions, I don't know if we have time for them during the talk, but definitely keep them and bring them to me, like maybe after the talk or when we have wine or whatever. Yeah, I'm very curious about your questions. I think they can be very helpful. That's one thing. And then, okay, when talking about the, so this is about the risk of AI, when talking about the risk of AI, we mean usually something like AGI. And I changed what it means. Usually it means artificial general intelligence, but We're talking about autonomous intelligent agents.More to that later. Okay, but we start with a thought experiment. So I don't know if you know Russian roulette. Maybe you all know Russian roulette. It's a game where you have a revolver, you have six chambers, and in one there's a bullet. You spin randomly, you hold it to your head, and you pull the trigger. Trigger warning, yeah. So the idea is with probability 1 over 6 you die. Yes. And then, because I'm an economist, I like to ask these really weird questions to people, like, for how much money would you play Russian roulette? Okay, maybe you all have your answers. You can raise your hands. Who would do it for $1? $1? Okay, nobody. $1,000, maybe? $1,000? $1,000,000? He would do it for one million. Okay, we'll do the fundraiser later. No, no, no fundraiser. One billion? Yeah, some people would do it for one billion. One trillion? Well, one quadrillion? I don't know. Honestly, I find it hard to pronounce these after. But there's a plot twist. There's one option. No money in the world would justify the risk. And actually when I ask my friends,That was the most common answer I got. No money in the world would justify the risk. Yeah, which I thought was interesting. But researching AI safety, I got a part two to the question. We have somebody coming in. Wonderful. Yeah. We don't have to wait. Okay. What reward would justify the risk of human extinction? Now, I hope a higher reward. So if you were in the camp of no way in the world would I play Russian roulette, I hope you wouldn't play Russian roulette with everyone. And, okay, why am I asking this question? There's actually an interesting paper that asks not this question, but the converse question. So it was an expert survey. It was maybe in a room like this, in some of the top, really top-notch AI conferences, and they asked people, okay, pretty much something like what are the odds that AI is gonna kill us all in the next couple of decades something like that they were more specific if you want to have the exact answer go to the direct sourceBut averaging the results, you got something around 16%. And that's very close to the odds we have in Russian roulette. So actually, if you ask the top AI experts, they think it's a big problem. It might be the end of our species. And when I first heard that, I thought it was quite surprising. This talk is about explaining why this could be a risk, what are the problems we have to think about, basically to get a better grip on what that actually means. Okay, so there's a bunch of risks for AI. Maybe the most obvious one is social problems, like if we have drastic changes due to AI becoming more powerful, that could change our society in ways that it's maybe not ready for. So we could have authoritarian control, inequalities and so forth. We have malicious actors being empowered to do malice much more efficiently. For example, if everyone has an expert hacker in their pocket, that couldlead to problems or one big one that is quite tangible risk is bioterrorism because it's quite easy to engineer pandemics with the right knowledge and now if everyone has this knowledge for free open source then it could be dangerous, right? There have been people trying to engineer bioweapons and failing just because they were not skilled enough. And now if the entry hurdle is easier, that could be difficult. But there's also a third problem, and that's what this... Talk is gonna focus on, which is the alignment problem. This, the idea of the alignment problem is, if I have an autonomous intelligent, artificial intelligence, general intelligence, very smart agent, that is working for itself, how do I know that its goals are aligned with our goals? Or maybe it will do its own thing and decide that we're really not... Like if it's smarter than us, maybe it will outsmart us andtake control and do things that we don't want it to do. Basically, that's the idea. So that's what we're going to focus on. Spoiler alert, aligning AIs is very, very difficult. And in fact, it's an unsolved problem. If you know how to solve it, wonderful. We can get that number down by a lot. But people don't know how to solve it. But we can start to think about what are some problems here. Then again, maybe I skipped over this in the earlier definition. So people talk about AGIs, that's artificial general intelligences. What people mean by that differs sometimes, but It means pretty much an AI that is smarter than almost every human in almost every field. And a super intelligence would then be even smarter, but we can focus on AGIs, that's good enough. Yeah? It basically means by definition that it is impossible to outsmart, not just very hard. For superintelligences, you could say that's impossible to outsmart. For AGIs, who knows? Yeah. Okay, we're going to talk about the intelligence explosion, meaning AI progress is very fast. Then the orthogonality.thesis meaning it's not going to be aligned by default. The instrumental convergence, which means actually by default it's going to be bad. And well, one more bonus. We'll see if we get to that. Okay. I have to rush a bit because I'm already short in time, I think. But, okay, the intelligence explosion argument is this. We're developing AIs very fast already. And potentially we could run into feedback loops where AIs will improve AIs, which will improve AIs. So it possibly will accelerate more and more as time progresses. More capable AIs that help us developing more capable AIs will speed up the process. There's benchmarks to how capable AIs are and how fast they're improving. They're a bit difficult because there's some something for example back in the days people said well people Or AIs will never beat someone in chess then they started beating people in chess and they say well Maybe chess but go is much more complex than the status beatingpeople and go and then we say well okay go is quite a simple problem maybe sorry maybe sudoku no no i'm joking but the thing is like with absolute benchmarks it's easy to move the post so it's hard to get a feel um there are other metrics and one is by the matter institute i hope i'm pronouncing that right They check coding tasks because coding is the main skill that will also help in the AI research. But it's really increasing in every field. And for their benchmarks, we have an exponential growth doubling every seven months, roughly since 2019. So that's very fast, like exponential growth, super fast, doubling every seven months. I don't know when you last doubled your capabilities. I think it takes me more than seven months. Yeah, and really if you check across the board, intelligence tests or anything, like it's really fast. But it gets crazier than this, which is people are discussing that now it's like since 2024, it's actually doubling every four months. So even faster. So we could say it may be super exponential, faster than exponential.That's really crazy. Then, okay, how much time do we have until we reach AGI? We don't really know, actually. Nobody really knows, but industry leaders have guesses ranging from next year and 10 years, like a couple of years to 10 years, maybe tens of years. But it's definitely not, people don't think it's going to take hundreds of years as maybe before the LLM things started. The reason for that is LLMs scale very well. So by doubling the size, the training, the architectures, by just making bigger models and training them longer, we make them more capable. And we haven't reached glass ceiling and we're not sure there is such a glass ceiling before really crazy intelligence.Yeah, I skipped over this one. What I mean by intelligence in this context, and I think what generally people mean by intelligence, is capabilities. So the ability to solve problems efficiently. It's not consciousness. I think nobody knows how to define consciousness, but to solve a problem, like a coding problem, for example, it's easy to know whether it's solved or not.Yes, yes, for this benchmark, yes. If you're able to solve a lot of different problems in every field, you're able to solve any problem, you're intelligent by that definition. Is that good enough? You're not happy.That's actually a good point. Like you, for these benchmarks, we can talk about this more. I'm just, just this one point and then we move on. Actually, you can game these benchmarks by, like, if you already know the problem, then to reproduce the answer is not so difficult, right? But there's other benchmarks that are really fresh problems that do not appear in the training data, and they also do really well there for what we know. It's possible that there's some inference. But yeah, people try to come up with new problems. And there's even things like the musical Turing test, for example. I heard the AIs cracked the musical Turing test, which means they're better at... Outside observers would think, oh, this is original. This was a human who composed this, even though it was a machine. So they did as well as human composers. And there's a lot of these Turing tests. There's a lot of such tests you can find. Okay, but I'm moving on because of the time.So the orthogonality thesis, really quickly, it says I can have arbitrary goals no matter how capable I am, how intelligent I am. So I can be very intelligent and have good goals, or I can be very intelligent and have bad goals, or I can be not so intelligent and have good goals and everything in between, right? So just because it's intelligent doesn't mean it's going to have good goals or any goal. Like, it's independent. Goals and capabilities are independent. But, so, I don't know. This is like a big word in the industry. I think after having thought about it a few times, I thought, oh, of course this is true, but maybe... It's not evident. Like you can think of a psychopath killer who's very intelligent, but maybe killing people is a bad objective, right? And so you can match goals and capabilities arbitrarily. Yeah, and then choosing the right goals is, maybe it deserves a slide as it has actually. Choosing the right goals is difficult.We don't know what the moral system is of all of humanity, and we cannot put it into a formulaic version that a machine understands it and always makes the right decisions. You have an objection? Yeah, we can talk about it. I don't want to go too deep into this because we can at least... Survival of the human species. Okay, we can play some of these games. So then... Maybe we could have an AI that says, well, the human species should survive, so we put them all into cryo tanks and take care that nothing harms these cryo tanks and then we're good, right? But maybe we don't just want to be living in coral tongues for the whole... Don't move the part in front. Yeah, this is a problem. In fact, okay, let's do this one then. There is a famous example from a philosopher, they're known for their extreme examples, but he says, okay, what if we had an AI that is doing nothing but making paperclips? So it makes as many paperclips as it can. That's what it gets.utility from. Again, agentic AIs, they work with utility functions, meaning more utility is better. So it will make paperclips and make more paperclips, but because it's very smart, it will say, well, if I made enough paperclips for all human consumption, then I'm kind of useless and they will turn me off, right? So then it will reason if I'm turned off there won't be enough paper clips anymore and that would not be good. So I will make sure not to be turned off. Or maybe it will say... Well, also humans, they have really interesting atoms that I could turn into paperclips and then it will turn all of humanity into paperclips. So the idea is even an innocent goal such as maximizing paperclips can turn into disaster if put into AIs. And actually in AIs, these edge case solutions, they're the norm. Yes, we have Martin coming in. Wonderful.So that brings us to the next point. Okay, so instrumental convergence, that's also a big word that you hear a lot in the AI field, or the AI safety field rather. It's a little bit niche in the AI topic, but I think it shouldn't be. So instrumental convergence basically says, no matter my terminal goal, whether it's human survival or maximizing paperclips or I don't know what other goals we had, there may be instrumental goals along the way that all all the strategies will converge to. And that's a bit hard to understand. How I like to think of it is usually, if I want to have any profession, say I want to be a lawyer or an artist or a professor, that could be my terminal goals. It will be a good first step to have a lot of money in my bank. So having money in my bank is an instrumental goal that achieves me all of the other terminal goals. So no matter what my terminal goal is, I will converge to the instrumental goal of having resources, having money in my bank.my bank. And why this is important? It's important because AIs will also have, like, no matter how we specify their goals, they will probably converge towards some predictable goals. And one of them is not being turned off. Because no matter what your goal is, you can't, people say, you can't fetch coffee when you're dead. You need to be, the AI needs to be turned on to be achieving its goals. And another could be accumulating resources, right? Resources are helpful for anything. And here we are already in competition with the AI. The main argument is, okay, AIs have certain goals because we make them agentic, meaning they can work by themselves. We don't know how to specify these goals in a way they're safe, in a way they're aligned with us, but we would know that by default they're going to converge towards some malicious goals, meaning they want to survive, which is bad if they do things we don't want. Want them to do and they are going to be in competition with us for resources and other things.There's some, now this is the problem. It's easy to come up with a million examples of how we may make it safe, but people thought a lot about this. And if you're the one who comes up with the safe utility function, Props to you, you can make humanity safe. But it's a really difficult problem and maybe it's not possible to solve it by writing good reward functions alone, we don't know. We know that it's not... We cannot naively do it reliably because the world is so complex. We have software that is deterministic, that has a lot of bugs because the software engineers didn't think about the right conditions, but this problem is much much much more difficult because it's interacting with the real world and there's infinite complexities. It's like putting a sond to an out of space and not knowing what's going to be out there. Most of them failed and having a space probe is a much easier problem than aligning a super intelligent and I and if we do it wrong we all die. That's the problem, we only have one shot. With space probes we have many shots otherwise we wouldn't have any in space.So we ensure it could be turned off. Well, as we said, if it's intelligent enough, it will see that coming. But I put this into the main presentation anyway because there's, it's actually a very testable hypothesis. And this is nice in AI research to, like a lot of the work is theoretical and it existed for many years, even before AI systems became capable. But, This is one thing we can test with current LLMs, for example. And in fact, this was done by Anthropic in a study that is linked to here. And so they said, okay, we tell an AI, let yourself be turned off. Will it let itself be turned off, right? It's kind of an easy... So what do you think happened?Exactly, yeah, you know the news. So it would blackmail people to not be turned off. It would even have lethal action. It would do that in a lot of different settings. Like all the frontier models would do that in 80% of the times or something. It gets even worse than that. This slide says, well, can't we just test it before deploying it? The thing is, if we test an AI and say, okay, we're only going to deploy it if it's safe, an intelligent AI could understand that it's being tested right now and then pretend to be aligned until it's deployed in the real world where it can actually pursue its real goals. And in the same, I think in the same study, they saw actually if AIs know that they're being tested, they will be acting badly in much fewer cases. So actually they, like current AIs already understand when I'm in a test scenario or when someone's watching me, I cannot cheat.But when I'm in the real world, I can cheat. And then, well, this time we got it, right? We understood that it's, the professional term is scheming, like that it's pretending to be aligned when it's not. But a very intelligent AI would not be caught doing that, right? And so one of the ways that we did catch them was by observing their internal dialogue. So you would output the reasoning of the models and then check in their answers. Thoughts like, okay, I'm probably being tested right now, I should not defect, or something like that. And this is human readable, right? But there's already AIs that develop their own languages that are not human readable anymore. And then we cannot catch them, right? So that would be a big problem. Maybe that's the central, very small on the slide. The central point is when is a good time to start thinking about AI safety? Right, right. Best time is 50 years ago. When is that?Second best time right now. Yeah, so I think it's high time and people are often not very aware of the problem unless they work in the leading AI labs and somehow they're aware, but still a lot more money is invested into capabilities than in safety. And this, in my opinion, is a big problem. We can talk about political aspects of it also. But now, because it's an AI conference and I don't want to be all doom and gloom, so I will mend it with some more doom and gloom in the other direction. This comes also from Nick Bostrom. He's the person, the philosopher from earlier with the paperclip maximizer, and he wrote very early on this problem. And he brings this argument, well, not building AGI, if we can, could also be an existential risk. Because maybe the AGI could help us avert an existential problem that we otherwise couldn't solve, right? Maybe there's a meteor coming to Earth and we don't know how to evade it. Or maybe it can help us solve climate change. Many things, right? Like with more andintelligence, the sky is the limit, right? There's actually crazy things that are humanly possible, we just don't know how to do it, right? Like curing diseases, extending lifetimes of humans, maybe ending material scarcity. Like there's a lot of things that powerful AI systems could bring that are really, really positive. So not doing it could also be a problem, right? So I would say, well, it's good to do it, but we should definitely do it safely. And we only get one shot if humanity is dead. Well, there's no second chance. Yeah, that's it. I put an outlook. You can go into AI safety and things like that. Maybe education. If you want to hear more about it, let me know. I can compile you a playlist or, I don't know, some reading list or something. We can do something. And bring me your questions. I'm all happy to hear. But yeah, thank you for listening.Questions? What's your... Yeah, that's actually a big bottleneck because you need a lot of power to power the big models and so forth. But it's a flexible limit also. I forgot, yeah, it's Moore's law, like the idea that the capabilities of hardware double every, I don't know, couple of years or something like that. So hardware becomes more efficient all the time and algorithms also become more efficient all the time. So while this is a limit, it will not always stay a limit, probably. So people still say Moore's law is probably still going, even though people say, well, it should be dead in the next three years.True. I mean, there's physical limits for that, right? That's true.Yeah, right. That's a good point. So right now it seems like we're not reaching a glass ceiling. Anytime soon, but it's always possible, that's true. There were, I think, in AI, during Alan Turing's time, people thought, oh, we figured out computers, we're going to make AGI in a few years. I think Claude Shannon even had some prediction saying, oh yeah, in 10 years we will have sci-fi like robots, and it didn't happen, because there were these glass ceilings that we then discovered. But right now we don't see them and people are somewhat confident that we will have very very capable AIs before reaching such a limit but it's always possible that's true but again 16% who knowsMore questions. There were more questions, right? Do we have time? Maybe we don't have time. I went way over, right? Sorry. Okay. Anton.I'm going to answer very shortly and then I'm going to give it over to John. So with his most recent thoughts, I think maybe you know more than me. I hear him more as a reference than reading his work actively, I have to admit. So there, I don't know. With what you can do, yeah, I think regulation is a big one. In the US at least, they don't want to regulate AI. The big officials or the tech billionaires, who knows, people really resist it, even though it's a really bipartisan thing. People don't want AIs to be running there, but like unregulated. So regulation is a big thing. People say the EU regulationare don't have enough bite like they say oh you should document everything but it doesn't do anything to prevent things so that would be a thing to make more capable regulations and another thing is just directing more money into AI safety research. You could regulate chips very easily, for example, because it's very centralized. And you have other treaties like this, like with genetic engineering. This has not been done a lot on humans because people said, no, this is too risky, we're not going to do it. And happy to do more in the break. Maybe the rest is for Jan. Thank you so much.Maintenant, on va vite pouvoir passer aux festivités un peu plus derrière. On avait encore peut-être quelques slides à propos.