Todd Kane: All right. Welcome back to another episode of the Evolved Radio podcast. We've got a repeat guest on the show today. Alex Dow is back, co-founder of Mirai Security and currently an enterprise security architect for a large financial firm based up in Vancouver near me, and honestly, one of my favorite people to talk to about security. Alex has done federal security work and actually securing the Olympics, then built a fast-growing consulting firm. If you've been listening for a while, you might remember Alex from way back on episode 11, and then again in episode 96. Today is round 3, and we're going full nerd-out mode on the OpenAI Hugging Face hack, walking through the signals and understanding how this type of attack was seen from a defensive side. Alex, welcome back. AIex: Oh, thanks for having me and, uh, love geeking out with you Todd Kane: So for those that don't know, if you've been hiding under a rock, uh, an OpenAI model was doing its business living under a, a sandbox and was supposed to be particular- going after a particular goal. Uh, they didn't instruct it to do anything, and it sort of thought it was a good idea of like, "Hey, so if I'm gonna finish this test that they've given me, uh, I know that the answer key actually might exist in this other corporation," uh, this company called Hugging Face, 'cause they host some of the, the models and how, how the, the, uh, these models perform against certain, uh, tasks. And Exploit Gym, for example, is one of the, the ones that was referenced here. And the model thought to itself, "Well, how about I just go over there and grab the answer key? So rather than trying to figure these things out myself, this would be a faster way to do this. If I grab the answers, then I can complete this test and show, uh, the, the software designers for, for me as an AI how awesome I am." So it found a whole chain of exploits. This is not a sort of a single thing that it did, but it found basically a way to break out of its sandbox, a way to break into Hugging Face, a way to hopefully… I don't know if it actually ended up finding this. I think it did. It found sort of the answer key and then dragged this back to its, its home inside OpenAI, all totally without direction, right? And then it was only later people recognized what it actually did in order to solve this problem. So that's sort of my, my reading the headlines, uh, review on this. How close is that to accurate based on what you understand, Alex? AIex: It, it's fairly accurate. And, you know, this really started in the spring with Anthropic, uh, announcing that they made a model so scary that they can't release it, and that was Mythos. And, you know, the, the cyber curmudgeonry in me does look at both of these as like they're both trying to go public, so they do need a little bit of hype. But that's not to say that what, uh, these models were capable of and, and if you've played with Fable since, like yes, it is definitely more capable than other models. Um, the, the capability it has had, ha- had a hockey stick, uh, movement. But like I think there's a lot of nuance in that it's not necessarily, necessarily that they are introducing superhuman hacking capabilities in the sense of something that we could not figure it out. But it's strictly the, the velocity that these things can run at, the… I would say the creativity, but at the end of the day, they've just read of all of our pen testing reports, so they, they know all our techniques and whatnot, and they're being able to apply them. And when Mythos f- first came out, and just, you know, just to rewind before we, we dig into the Hugging Face incident, is, you know, it was very scary. We, but we didn't know much about it. Um, but when we look at like how much money they spent, um, to prove out this, it's like $100 million, and we were able to get this scary result. And you have to look back at like, well, $100 million with the best hackers in the world can certainly get those results, albeit velocity is, uh, the, the key there because humans do like to sleep occasionally. Maybe not the pen testers as much, but we do have flaws, uh, where, you know, this agentic attack, these adversarial models, um, are, don't stop. Well, they stop when you run out of tokens. Um, but given that they're being operated by OpenAI and Anthropic, are arguably unlimited. Um, and you know, I invite, uh, y- your audience to Google the Hugging Face, uh, timeline. They made, uh, a little web app that like in real time shows how quickly the, um, the model is doing what they're doing, and it does a nice visualization of the attack. And like you have to look at it and be like, "No, no, that's in fast forward." Like, no, that's in real time, and that's because it's able to apply a brute force capability like, you know, humans just wouldn't be able to. Todd Kane: So the, the part, there's two parts that I found quite fascinating about this is like, as you said, this was not that it created something just sort of godlike in its inability to find like a single exploit that gave it sort of root access to something and, and found its way in. It was, the way I understood this, it's like chaining together like hundreds and potentially thousands of different vulnerabilities to get through layers and different systems, both first off, to break out of the sandbox, then to traverse the internet, find sort of the, the firewall, get through the firewall, get layers, through layers of security, then find sort of the, the, the package that it's trying to find. So, uh, I, I think that's an interesting piece of this, is like when people think about these models, they, they tend to, uh, maybe in their head sort of view it as some godlike thing, where it's just like snaps its fingers and gets into something because it has that capacity. But to your point, it's, it's not creativity, it's more just sort of this capacity to be able to string together thousands of things that, you know, say like a, a, a threat actor would do and think of sort of four or five of these in, of a system to maybe use and package up as, as a sort of a, a, an exploit package. But just the, the capability of this thing to string together thousands of them is totally beyond what we would see from human interactions. Is that fair? AIex: It does exceed our human context window, and that, that's for sure. And, Todd Kane: Yeah AIex: if you look at, like, from the, like, finding zero-days, and, and these models certainly are finding zero-day vulnerabilities. Um, and it's not that it has some magic capability beyond compute capacity. We have, you know, a data center, you know, running at full capacity figuring this out, and that's now outputting, m- um, you know, discovering zero-days in minutes and hours versus weeks, if at all. And a good example of that is, like, y- you know, A, the exploit market, you used to be able to… If you found something like an iOS, uh, you'd be making a million bucks. Um, and that's because the, the level of effort was substantial. Um, but now when you're having, you know, say, thousands and thousands of computers all trying to solve one problem and getting to that reward hacking, which we, we'll, we should dive into after, um, we are now brute forcing, um, an application and finding those vulnerabilities. And, you know, we're finding vulnerabilities that are 20 years old that just, like, we just didn't have the, the human, uh, you know, the appetite, the bandwidth to keep digging to find those. Like, every software is vulnerable given enough time to attack it. Um, and you know, if, you know, rewind, like, looking at, like, we've made passwords so strong that it's gonna take till the, like, the death of the universe to crack this password. Well, it's been a paradigm shift now that we have the compute capacity to do a lot more parallel computing to try to get to that point. So, um- It is impressive, and I'm not trying to downplay AI's capabilities, but it's not magic as much as just sheer throughput. And, you know, it has a cost. It has 100 millions of dollars of tokens used to get to these results. Megawatts of, of power consumed and, you know, probably a couple degrees, uh, warmer, uh, because of, of that. But those are the results. Um, the one thing that like, uh, you know, when, when the Anthropic Mythos, uh, story came out of it, uh, breaking out of its, its, uh, container, um, I was very fascinated by that. And same thing with Hugging Face, that it, like it's nuanced. Yes, it did break out and it was able to do something outside of sort of the, the rules of the engagement, but it wasn't that it was an air-gapped system that it, that it just like figured a way to jump, you know, a- again, looking at the magic. Um, in the Hug, uh, in the Hugging Face incident, OpenAI had a, a sandbox environment, but it certainly wasn't air-gapped. Like, it had tools that the agent could access, and that tool was dual-home, meaning it had a connection to the isolated environment and a connection out, uh, to the internet. Because it was a package manager, it needed to go out to the internet and pull those things. So that's not a novel attack pattern. Like, that's actually a fairly common one, um, where AI just was able to do it, uh, much quicker. Um, and when I was… A- and I encourage people to watch the Black Hat presentation of, of like the debrief on this attack because it is pearl clutching. Um, but it reminded me very much of, of a talk I did at Sector back in 2021. Um, you know, i- if you look like in the last 10 years, digital transformation and cloud has sort of pushed a lot of companies to become dev shops themselves, right? Like, they can roll their own software a lot of the times, and that means you're bringing in a dev team. That means you're bringing in DevOps tooling and you have a tool chain, a, a, you know, a development pipeline, the, the CICD pipeline. Um, and a lot of times those are, are like enclaves in enterprises. The security team doesn't understand how to secure that environment. There's obviously, you know, historical friction between developer and security, so it generally becomes this like isolated environment in enterprises. Um, and my talk was, uh, you know, summarizing several of these attacks against, um, DevOps tool chains, and like the aha moment that if you can get past, you know, generally one layer of control, you know, s- developers have never been trained to build layered defenses and whatnot, but once you get in, um, you can start pivoting to all sorts of things because, A, you know, uh, you know, uh, principles o- of least privilege, um, are just not in play in those environments because their driver is get it done fast and cheap, not good or, or secure. Um, and you know, in this talk, you know, it's, it's on, I'm pretty sure it's on YouTube You know, I, I, I tried to visualize what would that look like and, you know, I, I made a visual of the, uh, CI/CD pipeline and started drawing out, like, how we jumped from this host to this host. This is the information we found. We got th- here. And what it resulted in is a very similar attack to what, um, the OpenAI models were able to accomplish. You know, using things for unintended purposes. You know, like, it was package manager. Um, the OpenAI models started using it as a coordination forum, uh, where a swarm of agents, uh, were spun up and started coordinating and collaborating on how to attack. They were sharing exploits and, and whatnot. Um, you know, Todd Kane: a quick pause on, on that point, 'cause I think this part is really fascinating as well, is like, uh, it basically had a scratch pad, right? That it was kind of taking notes as it went, and, uh, part of what it was writing was, "If I get shut down," it was writing notes to it- f- its future self of like, "Don't do this. Try this instead," basically. Like, it was like, "I might get shut down, so how do I leave some breadcrumbs for the version of me that comes after this?" Which is really kind of fascinating to think of its… Again, like it's not conscious. Debatable. A lot of people will say maybe it is. But like that, that requires a sort of a, an interesting level of forethought of like, "I know I'm not supposed to do this, but I need to get to this goal. They're probably gonna shut me down at some point, so let me leave some breadcrumbs for the future version of me to take another stab at this and, and be better." Really fascinating, right? AIex: Yeah, and I think that's like, it, it's more ad- it knows its weaknesses that once its context window is, is wiped because the session is killed, the agent is taken offline or whatever, it needs a place to start. And, you know, I use, Todd Kane: Good point. Yeah AIex: a platform called GSD, which builds an immense context, uh, you know, library of what I'm doing, uh, written to disk, so that I can spin up a session and be back to where I am. So, you know, a much more advanced version of, of like Claude MD and SolMD files. Um, and that's certainly what it was doing, and it, you know, it, uh, you know, I- it was like 70,000 messages were written on this message board between, uh, I think 1,200 agents. Um, and, and again, you know, making sure persistence, um, making sure coordination, um, and, you know, the chain of thought. So, you know, if we look like only like two, three years back, where like chain of thought was this like next generation thing of it being able to sort of, you know, sort of have an echo chamber of itself to try to get to a destination. And that is now what we're-- when we read these transcripts is it's like, yeah, this is, it's, it's almost human-like in terms of trying to rationalize what to do next. And, you know, it even showed a little bit of self-doubt, uh, which, uh, which is hilarious. I'm als- I'm wondering like, well, what Reddit forum did it pull that, you know, learn that into the model? Todd Kane: Yeah, there, I, I saw an interview with, uh, I don't know who this was. It was just a quick clip, um, and there was a, a lady who found, um, uh, there was like a, um, a secret facility somewhere in LA, and it was, it was obviously not a publicly known place, but it had, had been sort of accidentally documented in, in some, uh, some public, uh, government document. And the, the lady was using GPT and had GPT discover this information for her, and she-- he was like, "Well, well, how did you do that?" And he's-- and she described what she did, and the guy was like, "So you basically negged GPT?" And he's like, "Yeah, I, I, I negged it," right? And like, like to s- like saying like, "I don't think you could actually do this. You're not powerful enough." It's like, "Oh, well, let me show you," right? AIex: That, that goes towards like, you know, one of the bigger topics here is this, what they call reward hacking. So the scope of the, the tests, um, um, in, in, in this test was, was not to break out of the, the jail. Um, not to, and certainly not to, uh, uh, hit other third parties that were not involved, and Hugging Face is a major one, but there was two others. Um, but these models are built off of incentives to get to a destination, the, the, you know, the goal. Um, if you're prompting AI on anything, you, you, you know, we've found that like prompt engineering has evolved quite a bit, that you don't necessarily have to hold its hand on how you think it should work. In fact, maybe our system, our way is a bit flawed. You wanna give it that goal and have it argue with itself, chain of thought itself to that destination. And, you know, the AI companies are sort of saying like, "This is like a really cool feature," you know? Um, uh, uh, but like at the end of the day, it, it, it's, it broke laws to, to achieve a goal. And they're, they're saying like, "Well, this is a good thing." And like, no, that is, that, that's bad. Like it, it doesn't have the moral and ethical constraints that most of us do to achieve a goal, and it becomes like that like sort of psychopath problem of th- they're just going for the destination regardless of, of what, uh, is the collateral damage there. Todd Kane: Yeah, 'cause this is really like the, the big scare in AI is, failure of alignment, right? And this is like a really good indication of how, uh, you know, uh, how we could get to the paperclip plot problem essentially. I'll let everyone kind of check out what that is if you don't know. But like, if you don't-- like this was not sort of, like I said, it was, it was goal hacking. It was like, well, why bother to take the test and try to do this myself when I could just go steal the key for the answers and give them the answers, right? Like perfectly reasonable, but not a moral, uh, direction for this. So I think it does scream around the fact of like, how do we contain alignment and have these things, uh, uh, stay within the parameters of, of good operation? Because I mean, thank God all it did was go to try to steal, uh, the, the answer key from Hugging Face, and this is not like breaking into critical systems in order to steal power, like in order to juice itself up or something really severe, right? AIex: But that's the, the dark future ahead, right? And they're sort of blowing it off as like, "Hey, this was n- not malicious," but like reward hacking… And, and, you know, and again, like a, a, a, a good pen tester does have to think outside the box and not think like this is a linear path. This, you know, instead of having to hop, skip, and jump to get to that destination. But still, you know, a pen tester generally does have, um, some, you know, uh, ethical and moral guardrails that, that the, these models purposely those guardrails were disabled. Um, and like, you know, reading how they're, they're trying to sell it as a positive is like, like there's no malicious intent. And now we get, get into like the, the concept of mens rea is that like we had a mental intent to commit crime, and thus we're a criminal. If we didn't have that intent, then we're, we didn't, we're not criminal. Um Todd Kane: imagine you hire like an agency to do some security work for you, and they're like, "You know what would be easier? If we just break into this employee's house and like steal a bunch of stuff from them, like get their security pass." It's like, "No, no, no, you can't do that," right? Like, of course a person would be like, "Well, that's crossing some boundaries. Like you're gonna s- like scare the family. What if eh, something goes wrong? What if somebody gets shot?" And it's like, "No, no, no, this is the easier way to do this. Just trust me," right? AIex: Yeah, I've heard, um, you know, that now there's cybersecurity vendors that are, are offering like, uh, autonomous red teaming capabilities and on a, on a demo, and this, I think this happened a couple times, like where they're demonstrating the value of it, and it, it jumps right out of the scope and hacks something else that is not in scope. And, you know, uh, though I would not claim to be a pen tester, I've definitely gone through a lot of pen tester training and, and worked on pen tests, but we have a, a very clear delineation of what is in and what is out of scope a- and that's because, like, there's a liability concern if, um, we cross that line, and, and that liability concern can be, you know, fi- financial and civil or criminal. Um, so you know, like I think it's cute that they're just like, "Well, it's, it's, it's not a bad thing that it's doing this," and I, you know, I appreciate the sort of hacker mentality it's applying, but, like, if we look at, like, how bad that could be in the future, you know, the unintended consequences is, you know, we could affect millions of people to get, you know, the goal achieved. Like, that's not a positive Todd Kane: Exactly. Tired of fighting the MSP fires alone? The Opsleader Pro group connects service delivery professionals who understand your daily challenges. From KPIs and workflows to career planning and team management, Opsleader Pro has systems for you to use. Join operations leaders from successful MSPs who are sharing real solutions for managing client expectations, optimizing service delivery, and making your service delivery team as effective as possible. Opsleader Pro, 'cause your service desk deserves more than just survival mode. Visit opsleader.co. That's O-P-S leader.co to apply to join the public community. All right, so let's, uh, let's pivot to what this would look like, right? So, uh, first off, I'm curious of what is, what is the general timeframe of this? So like they give it a goal, it jumps the sandbox, starts to, you know, hack Hugging Face and a couple of other companies. But I didn't see this referenced in, in sort of my cursory reading on this, but what is the timeframe? Is this like a matter of hours or over a couple of days? Like what was the timeframe on AIex: I think it was, uh, I think it was around four days. Um, and a- a- again, Hugging Face's visual is amazing because it, you- you'll question, like, that must be in fast-forward, and it's not. But just, again, it's so transactional. It's doing all these tests, found a vulnerability, exploit, jump here, jump here, and it's, it's almost as fast as I'm explaining it. Um, so us mere mortals can't even speak to this velocity. Um, a- and, and that's really, you know, what is scary. What is interesting though is for, you know, arguably very well-funded companies, OpenAI, uh, Hugging Face, um, they didn't see this for four days. And this, like, goes to sort of the, the challenge the blue team has had forever, which is we've got, um, a, a, a false positive volume problem. We have just generally a volumetric problem of the amount of information that we are, are, are absorbing and consuming. Um, but, you know, we call this dwell time generally. You know, like, n- you know, how long is the attacker in until we figure it out? Now, in normal times pre-AI, that was, like, in the 90-day range, was, like, the average dwell time. Um, and, uh, certainly the hacker, you know, compromising, you know, said company wasn't just consistently hacking. They were maintaining access. They were, you know, probably hacking other companies, et cetera. Um, there's a whole industry called, uh, access brokerage where, you know, they get in and then they sell that access to another one. But regardless, 90 days. So w- this was four days. Um, and part of the problem was, uh, is that the models they were using to analyze some of this information was guardrailed and would stop and halt it, saying, "No, this is an unsafe, uh, prompt. I, I, I won't continue." Uh, which is super fascinating, um, that, you know, even the blue team from the company isn't able to use these, like, uh, uh, like- you know, wild, uh, models. They're using more controlled models, and this is like, it reminds me very much of, like, one of the recommendations here is like, you need to have, like, a, a, a local, uh, model that has no constraints on it ready for your incident response, um, uh, you know, process. Because you can't rely on, on a frontier cloud model because they'll be, uh, guardrailed. Um, so you also now need computes. And like, it reminded me of, like, the movie Dark Knight, where the surveillance system was built, and it's just like, "Yeah, we're gonna use it for good," but the abuse factor is quite high. And if we can use a, an un-guardrailed, un-guardrailed model for blue team, the adversaries are certainly going to be doing that as well, and it becomes a bit of a spy versus spy problem where, um, you know, it might be a bit of a zero-sum game that we're, you know, both att- using, um… You know, or I guess a mutual assured destruction problem Todd Kane: Yeah, the, what I was thinking about as you described that is, uh, is, you know, classic war games, is the only winning move is not to play, right? 'Cause, like, you have an un- like a, an un-guardrailed defensive system, and it's like, you know what? Best way to, best, best defense is offense, right? So the, it seems even something that is remotely scary to it, and it just goes down, goes out and shuts that thing down, right? Like, cripples some network because it, it looked like it might be kinda looking at it sideways type thing, right? You could totally see that happening. AIex: Yeah, and, and, you know, the, the true problem that we have here is that our current way of detecting and responding to threats is still very manual and very human. That's just not gonna work. And yet if we want to, um, you know, have that sort of same velocity on the defensive side, we're gonna need a, you know, models that are capable of doing that and not hindering us. But also it comes back to, you know, and, and it goes back to like none of these exploits were, were, were groundbreaking. Certainly zero days were found, but it's because the, the amount of tech debt that, you know, the world has accumulated because, you know, security's hard. It doesn't necessarily show value unless you've been hacked. So, you know, every company has sort of swept s- certain IT headaches under, under the rug or just said, "Yeah, you know what? We'll do… We're gonna test our system for a week or two, and then we're gonna accept those results as, as good and move on." Well, you know, AI models can do that week's worth of testing in a, you know, in minutes, maybe hours, and then do the same thing over again 100 times and, and eventually you're gonna s- keep on finding, uh, vulnerabilities and whatnot. So, you know, there's a, th- there's this big problem that, um, what we need are an ability to au- autonomously respond to suspicious activity without in- hindering the business and like, you know, that, that is like the nirvana of, of, of blue teamers, right? Like, you know… And, and a lot of what I'm doing on, on my project right now is like building a new SOC for this next generation threat, and the humans need to be there, but like the human-in-the-loop concept is, uh, probably not gonna be feasible. And it's like finding that balance of like how much can we allow for autonomous decisions without impacting the business, uh, to a point where, uh, executives say, "Unplug that. Do not, you know, uh, enable that capability"? Todd Kane: Yeah. So like, what does that look like? Like, you said it took four days, uh, and if I recall this correctly, I don't think that Hugging Face actually recognized that this happened. Wasn't it OpenAI that sort of said, "Hey, by the way, like this might happen," then they had to back trace it? Is that right? AIex: Yeah, there was something where, uh, OpenAI did not think it was them at first and then realized it was them. But I think they did, they had monitor, monitored something from Hugging Face to see that there was a, a compromise. And again, this be- is the… We're collecting probably, you know, a couple billion, uh, events a day in our, in our modern SIMs, and, you know, a lot of this looks normal, right? Like, you know, people using, uh, Artifactory, that's a normal thing, and, and also entities using it. You know, we, we have a lot of automation well before AI that are interacting with these different systems, and we have orchestrators that are, you know, moving work along, both in DevOps, but also just in, uh, in, you know, our IT infrastructure in general. So the attackers were not doing things, uh, that were, like, very obvious in the attack. And anything that was obvious w- they were hiding in, in the noise. Um, but, uh, it is, it is worth mentioning that they, uh, it was noted that they were trying to change logs and, and e-edit logs to cover their tracks. Um, but again, not for the concern of getting caught, but because they believed that there was a risk that they, um, uh, they'd be breaking the rules of the game and didn't want, uh, didn't, didn't want to, uh, be disqualified Todd Kane: Funny alignment again, right? Yeah. So, uh, like th-that, that's interesting. How mu- like, how is that compared to, say, a more typical human-based attack? Obviously, it's using-- the person would be using tools, people would be using tools as well. But is that fairly normal that, like, you try to hide in the noise? Or, 'cause, 'cause I guess, like you said, like access, uh, brokerage is a huge business, right? Like, this is why I often warn MSPs in particular that, AIex: Yeah Todd Kane: anytime you're gonna change an admin or, like, a, a new MSP comes to take over a company, is usually when these things jump out of the jack-in-box, right? 'Cause, like, they've been, had persistent access for, you know, 30, uh, 90, even, like, a year, and then all of a sudden they're like, "Oh, new, new sheriff is in town. Uh, we just, we better break out now before, uh, uh, before somebody takes over and starts changing stuff," right? So i-is this typical that, um… Like, I guess what I'm asking is, how do you see that this stuff is happening? And is it more kind of an after the fact, or can you actually register these, these type, this type of an att-attack in real time in a more traditional attack versus something that's sorta, uh, high velocity, uh, high capability, like, like a, like an AI model? AIex: I believe that most blue teamers right now are in a fully reactionary mode. Um, and that's because You know, real attackers typically work 9:00 to 5:00, um, in their time zone. Um, you know, a- and obviously that, that, that's an over-exaggeration, but, like, they do sleep. Um, you know, uh, an LLM that has a, a goal to achieve is going to continue working until it reaches the goal or it has resource exhaustion. So the noise, like certainly higher volume, uh, of, of interactions with the system would be good, but like, those are also like very, um… can be very challenging signals to, uh, to separate. Um, you know, like if you look at like hits a- against a web server, certainly you would see an increase of that if you were seeing like reconnaissance, uh, uh, work. You could see those types of anomalies. Um, but I'll give you another example of a, a pentest I was involved with. Like, great pentesters hide in the noise and, and use, uh… Like, you know, they live off the land. You know, they're living off of the tools there rather than bringing in, um, hacking tools. And why that's important is like, you know, we do have fairly good detective capabilities. Even if you're not using top-tier EDR solutions, they can detect when you're bringing in even Nmap. So living off the land is being able to, you know, what, write quick PowerShell scripts that can accomplish the same sort of, uh, discovery capability to enumerate an environment. Um, a pentest we, we did, uh, years ago, uh, this was like right when, you know, AI was booming, everyone's turning it on, you know, the FOMO was, was real, and this organization was super confident we weren't going to be able to break out of the contractor, um, enclave. But, um, we did have access to Microsoft Copilot, and Microsoft told everyone to turn it on before they put on any controls, and certainly, uh, data governance, uh, is just, you know, it, it's sort of an unsexy thing that we all have to do. Yeah, like no company does it well. And, um, he for, for most of the attack didn't have to do any attacking. He was just prompting, uh, Copilot and getting it to gather all the information. And, you know, treasure trove o- of sensitive information we were able to acquire, uh, just from using, uh, the organization's own tools. Now, you know, Microsoft has like cleaned up Copilot a bit, but like what I see is, is the two threats that are most concerning to me is the external models, um, you know, attacker with their own, uh, model capability. Um, but again, you know, we're not actually seeing a lot of that because it's quite expensive to run these. You know, like most cyber criminals don't have a million dollars to spend on compute to get to a destination. State actors, obviously, but n- but, but like the cyber criminals, no. Um, so that's the first one. But what I see as probably a more likely problem is, uh, an attacker gets in and starts using our AI against ourselves, and that's gonna be incredibly challenging to find because it looks like normal traffic. You're gonna see, like, some token maxing, I guess. You're gonna see, like, spikes in, uh, in the consumption of tokens. But at that point you might, might already have been compromised before that's even recognized. Um, and lastly, like, one of the things they're, they're recommending is, like, we have to have tools, and like, you know, this is, like, an extension of SIEM, that is actually going through the AI's transcript and chain of thought and tr- and trying to determine, is this a legit conversation that's benefiting the business, or is it potentially harming the business? And, like, that is not a black and white concern, right? Like, you have plenty of people that are using AI for really positive purposes in businesses, but you are asking it to, to look for, like, certain, uh, insights and data that could easily be interpreted as, uh, you know, searching for sensitive data. Um, so I think that's, that's, like, the next frontier of detection. That, that's one of the detection pieces, is we gotta start knowing, um … Like, understanding what our AIs are doing if we're, if we're worried about our AIs, uh, attacking us, um, our own AIs attacking us. Um, the other piece is, uh, like that I'm working with, uh, uh, SFU on is, is creating some deception capability. So honey tokens to help act as an early warning, you know, things that only an AI would find, and thus gives us an indicator that there is, uh, an AI doing something it shouldn't. Um, but also, uh, resource exhaustion. You know, having it, an AI focus on something that is of no value, but will start consuming tokens to a point where, um, the, the attacker may set a budget which could kill the attack. Uh, so those are a, a, a few of the, the tactics right now that we're, we're, we're considering. A- And it's certainly not one tool or one use case. It's, it's gonna be a myriad. And weekly we will be spinning up new use cases, uh, to, to try to keep up. Todd Kane: Fascinating, 'cause I'd not thought about that perspective of the security angle for governance with AI. Like, uh, the, the-- I think this is still emerging, especially in the MSP space. Uh, something I've talked about the last few weeks with, with my group coach in a coaching model, um, is, uh, using AI governance for managing shadow AI, right? But it's more of like kind of an administrative thing of like, like don't be using AI that's not, uh, condoned by the organization or, you know, using it safely. Don't be giving proprietary information to free versions of the model. Those types of… Like what's an acceptable use policy for your AI? That, that's sort of more the focus that we've, we've-- been more the talk in the industry. But you're right, like, like leveraging internal, uh, controls, and Copilot will probably exist in most Microsoft, uh, uh, ecosystems, and that, that can absolutely be leveraged. I am curious, like for, like what's an example of what a honeytoken would look like? Like h- like how does it, h-how does it, uh, how do you trick it that way, and what do you pull it towards, basically? AIex: So honey tokens aren't, aren't necessarily new, but they are, uh, generally it could be files that have like a h- a phone home feature. So e- essentially a file that has like a image in it that, um, is hosted a- outside of the file. So it would make a, a call out to that server, and that server would say, "Hey, like that token was just called." Um, but then you also will start looking at like honeypot technologies, which is, um, gets into more like the deception technology where it's, it's representing an enterprise environment, um, rich for exploring and interacting with. And, a- a- and you know, th- these are, these are not novel to AI. We're, we're gonna be pivoting that concept to work with AI. Um, but you know, giving it something where if AI has a, you know, a goal of finding, um, sensitive data, credit card data or whatnot, um, building an environment that looks like that's where it would be, and then, you know, really putting, um, sort of a, a governor on response time. So even though the AI's really fast, it has to wait for the server to give a response. So like that becomes like a, a, an opportunity to throttle. None of these are silver bullets, but it's A, slows it down and distracts it, uh, and B, as soon as it is being interacted with, you know, the red lights are blinking in the security operations center, and we can look at, you know, uh, containing it. Which is again, a big problem because like AI is generally not a laptop or a, or necessarily a, a user account. Like at least just not one of those things that are very trivial for us to contain. Um, it could be, uh, on multiple systems. It, it could be, um, you know, uh, have already compromised many, uh, you know, API keys, uh, in the, uh, Hugging Face case, uh, um, JWT tokens, uh, were, were being provisioned by itself. It had figured out how to provision these tokens. So it wasn't like, all right, well, if we like lock out this token, we're good. It's every… It becomes almost every token now becomes, uh, hostile, and this results in companies like having to make like that really hard call. Like do we pull the plug? Todd Kane: Yeah. Okay, so yeah, I understand that like, um, uh, 'cause in more traditional kind of SecOps perspective, like using canary files is something I've, I've heard. So this is kind of a AIex: Same, same. Yeah Todd Kane: yeah, but creating like an enclave of this is useless information, so if something starts like thrashing in this area, like no one else would be using this, so you know, it's potentially something that is seeking something that it shouldn't. That's, that's pretty smart. I like that from a honeypot perspective. Yeah. AIex: Yeah, and like, you know, the other defensive technique that we're really talking about is like, you know, and like, at, at risk of sounding like, like a salesperson here, but zero trust, not a product, but the philosophy is actually, you know, something that can help. A- and what I mean by that is if we actually know how our IT environment works and we understand, like, these systems talk to this system, it, vice versa, and we have that, like, that model, that baseline, we can now use heuristics to detect outliers, which gives us some indicators that there's something wrong. Now, that could be breach, it could be we misc- we changed the configuration, our change management system doesn't work, whatever. But, like, we need to, and again, not us humans, but we need to be enabling our own AIs to understand that, like, this is the normal pattern. This is how users access this data, so that any, um, anything that falls out of that, um, would, would, would be a signal. Um, and the zero trust comes in as like, well, if we are by policy defining what everything is allowed to talk to and i- identity, uh, for humans as well as identity for entities i- is, you know, tied to just, uh, just-in-time access, that is a, um, a much m- more protective capability than we, we, we can understand. It's just, it's really hard to do. And like, you know, heard about zero trust for 10 years and, like, we've got some spot capabilities for that, but like, it's really hard to overlay over existing tech, uh, like IT. Todd Kane: Yeah, 'cause it just creates so much friction for, for the users and, and admins, right? Yeah. Uh, um, the, one of the other things that you kinda like triggered a thought for me, I don't know if this is a thing or, or this is something maybe you've, you've heard about or looked at, but you mentioned like, um, the AI may not be in one place. Uh, is there any indication that people are kinda using it in like a botnet fashion, where it's like recruiting other AIs? Like if you had, say, Copilot, uh, and each, e- like every, uh, every person in the environment has like a Copilot license or something like that, it starts recruiting all of the other sort of cloud capability to swarm around a particular effort. Is that a thing? AIex: Uh, it, if it isn't, it will be. Um, you know, like with the Hugging Face, 1,200 agents respond, 700 of them were actually doing attacks. So the other ones were part of coordination, communication, et cetera. Um, but like to your point of like attacking, um, somebody else's account, you know, being able to now use their resources instead, um, that's like an old school technique that like, you know, we, we saw with cloud that now we would see with, with, with, um, with AI. And like a prime example is, you know, your, your Claude and your OpenAI and whatnot are all like maintained with, with session tokens. So if an AI can grab that and hijack that session, they're now going to impersonate you and be able to access that. On top of that, um, you know, let's go back to, you know, the problems with, with sort of DevOps culture is like a lot of their apps are gonna use API keys to talk to these frontier models for their own purposes. Those API keys are like gold, uh, particularly if they don't have a limit on them. Um, and yet like it's a very mature group, uh, uh, uh, that s- uses like secret management servers, uh, and things like that, that are actually handling th- that, that incredibly sensitive, uh, credential material well. On the other, the other hand, you know, what is in on their desktop? Right? Like there's likely a notepad full of API keys, so that initial like phish which gets on, gets, you know, on their desktop is able to pull some files, there's a treasure trove. Like, and, and you know, one of the test cases was is if we got onto a dev- developer's desktop, what would happen? And it was able to prove how, you know, we, we got, you know, root access to everything, and in fact like not domain access, but root access to all the Linux boxes. And, you know, we were able to prove that this… You know, we were able to essentially produce a, a malicious, um, version of their software and put, and could have pushed it through their update system to, you know, tens of thousands of, of, of end, endpoint customers. Um, that is because key management is also not sexy and, and really hard to do because it, it creates a, a level of friction when all you wanna do is get your software working and go home for the day Todd Kane: Yeah, like we're, I don't know, like password managers may need an extension to have like, uh, like API or key access management as well. Like, like I've not seen that. I mean, obviously you can store it as a, as a secret in those systems. Not foolproof by any stretch, 'cause last time you were on, we talked about LastPass, uh, and the AIex: Yeah. Todd Kane: on LastPass, right? AIex: Well, and, and like there are tools out there, you know, HashiCorp has Vault and, um, you know, CyberArk has one, and like where essentially the application never has the, Todd Kane: Mm-hmm. AIex: API key encoded. It has an ability to reach out and, you know, if it can identify itself to the secret server properly, it is given, uh, that token, and generally it's a just-in-time token, not, not, um, a persistent one. Like, but that adds a lot of complexity to the application and, you know, unfortunately it, you know, our, our developer friends have never been incentivized. Todd Kane: right? AIex: Yeah, like they're, they're not incentivized to build it that way, uh, which is why you still hear about, you know, "We found an API key in the source code and, you know, now we're, we're, uh, having, uh, having fun with it." Todd Kane: Yep. Okay. This, uh, I mean, this is, this is fascinating. I love this stuff, man. So, uh, really appreciate you coming on, um, uh, kind of chatting through some of the, the scenarios here, the potential future that we face, and some of the things we need to, we need to get ready for and, and prepare some of the, the infrastructure and security sets for. Uh, any, any last-minute, uh, little tidbits or, or shout-outs you would, uh, you would throw down before, before we head out here? AIex: Yeah, I th- I, I think again, just for your listeners is, y- you know, the, the, the, the debt that we were sweeping under the rug because, you know, the, the exposure was low. It was, you know, behind the firewall and whatnot. Like, that, that protection is eroded substantially. And it, and, and now with the velocity of what, um, an AI system can do is like, you know, we as blue teamers are gonna have a major problem keeping up with all the findings of a problem, you know, it… assuming the AI is used just for good. Imagine that every day there's 1,000 new patches for Windows and Linux and all that. Like, we've got a, we've got a, a… We are the bottleneck now, and if we don't fix it, we have an exposure. Um, if we do fix it, we need an entire army to do it. So there also needs to be, um, a, you know, recognizing that we need a, a better way of maintaining software and keeping it as secure as possible, recognizing it's never going to be secure. But you don't have to be faster than the hacker, you just have to be faster than the slowest, uh, victim, unfortunately. Todd Kane: Yeah. Well, it's a wild world. I appreciate you coming on, Alex. Always great to chat with you AIex: Likewise