NEAR - Sponsor Image NEAR - Confidential swaps across 35+ chains Friend & Sponsor Learn more
01:13:00 · 1 year ago
Podcast

AI ROLLUP: AGI in 26 Months | Meta's New Top Model Cheated? | $20 Billion To AI Apps

The Acceleration Continues

Up next

All episodes

Inside the episode

This week on the AI Rollup, we dive deep into the accelerating world of artificial intelligence, bringing clarity to the chaos with our special guest commentator, Josh Kale, an expert in frontier technologies. As the Silicon Valley AI scene heats up, the latest developments offer compelling—and sometimes controversial—insights into the rapidly evolving landscape.

Meta's launch of its latest AI model, Llama 4, has stirred significant controversy. While the model boasts impressive capabilities on paper, there are widespread concerns and allegations that its performance has been exaggerated. Critics suggest that current AI development is overly focused on scoring well on standardized tests rather than delivering genuinely useful applications. The intense debate underscores a broader issue raised by commentator Ejaaz: that many organizations may not yet understand how best to leverage AI models effectively.

In parallel, a provocative roadmap, "AI 2027," authored by influential voices like Scott Alexander, paints a dramatic future of AI development marked by geopolitical warfare, espionage, and potentially uncontrollable intelligence growth. Alexander, known for his influential blogs and deep engagement with AI alignment and rationalist communities, significantly shapes how thought leaders and insiders view AI's trajectory over the next few years. His predictions force us to reconsider the profound implications of AI reaching general intelligence (AGI) as soon as 2027.

Meanwhile, the corporate world is swiftly adapting to this accelerating future. Shopify CEO Tobi Lutke has issued an internal directive declaring AI proficiency as mandatory for all employees, signifying a broader trend that could redefine employment and professional skill requirements industry-wide. Similarly, Claude Education showcases a powerful application of AI in education, dramatically improving student performance and hinting at AI’s transformative potential within academic institutions.

Finally, venture capital firm Andreessen Horowitz (a16z) is fueling the AI fire, aiming to raise an unprecedented $20 billion megafund dedicated to supporting innovative AI startups. Among their targeted investments are platforms designed to enhance creativity and collaboration—so-called "Vibe Creation Platforms," which illustrate how Silicon Valley envisions the future of AI in work, play, and creativity.

This episode of the AI Rollup highlights how quickly AI technology is advancing—and how deeply it is reshaping industries, education, and society at large.

Transcript
00:04
David

Welcome, Bankless Nation, to the AI Roll-Up, where we cover the weekly news in the AI industry. AI is accelerating us into a weirder and more chaotic future, and we are here to help you keep up with the future that's hurtling towards us. We're trying out a different format this week because the crypto AI industry is pretty dormant, uh, while the Silicon Valley AI industry is only getting hotter. So me and Ajaws are bringing in a third commentator, Josh Kale, who's done a couple episodes with me on Bankless before, all about frontier technology and who we are tapping in to help us go through the AI news this week. Ajaws, Josh, how are you guys doing?

00:37
Josh Kale

Uh I mean I'm

00:38
Josh Kale

yeah, I'm I'm I'm feeling good. Um it's it's awesome to have Josh. Um though I can't really get rid of the uh yeah the $10 trillion wiped out of the stock market this week thanks to uh Trump's wonderful Liberation Day tariffs, which has, in my opinion, obliterated any kind of economic recovery that we wanted to get. But more so, it looks like he's using Chat GPT to create his tariff formula, right? I don't know if you guys saw this, but um, if you type into any AI L L M, so that's like Chat GPT, Claude, whatever, um, you know, can you create a basic tariff formula? You get essentially what the US announced for their tariffs across every single country, which is just insane.

01:21
David

It's pretty hard to get away from talking about tariffs this week, just because everyone is just looking at the market, looking at tariffs, tariffs, tariffs, tariffs, tariffs. But I'm I'm pretty glad that we are also going to focus on AI, which I think we're all pretty, pretty stoked about. Josh, it's good to have you on the uh the AI roll up this week, my man.

01:37
Ejaaz Ahamadeen

It's great to be here. Yeah. Thanks for having me. I'm excited to talk about this stuff. The the funny thing about what EJS just said is

01:42
Ejaaz Ahamadeen

the way that we found out they used AI is actually by using AI ourselves and kind of reverse engineering it. And then they openly admitted it after the fact. So wait, did we

01:50
David

Did they admit that they did use ChatGPT to come up with their tariff?

01:54
Ejaaz Ahamadeen

not that AI was used, but that there was this weird formula that wasn't just tariffs. And the way it was discovered was by actually using AI and kind of reverse engineering it. And then they were like, oh yeah, we actually did do that.

02:05
Ejaaz Ahamadeen

So it was implied. It wasn't directly stated, but like there was some artificial help in the decisions being made last week.

02:10
David

I mean, I feel like that's just a harbinger of things to come of like global decisions being made using AI influence. Uh, you know, the crypto technologist will would love to say it's like, oh yeah, eventually the future will just be governed by AI. And, you know, we accidentally got there maybe a little bit too ahead of schedule, a little bit too soon, a little too early for comfort.

02:31
Ejaaz Ahamadeen

Yeah, it makes the AI alignment conversation a lot more scary when world powers are using it to make decisions.

02:37
Ejaaz Ahamadeen

Or even write law.

02:41
David

All right, let's get into the the news of the week. Big news this week. Uh, five big things that we're going to cover. First, the most controversial AI model release ever came out of Meta. Meta released a new Llama 4 AI model, uh, which on paper is very impressive, but people are worried that they are lying about its performance. Uh, and which will also lead into an Ajaw's conversation about why he thinks no one is using AI models correctly. That's up first. Uh, second, some legends in the AI space author a document called AI 2027, a theoretical roadmap for the how the next three years of AI development will go, featuring things like espionage, geopolitical warfare, and runaway AI intelligence. Uh, third topic this week, uh, Toby Lutke from Shopify ran an internal memo saying that using AI is now a job requirement, is an expectation for all employees. He released that memo on Twitter, so we're gonna go and talk about that. Uh, fourth, Claude Education focuses on using AI models for just revamping the education industry. And we've already seen a uh school in Texas rocket their students' learning scores to become like top two percent in the nation. And then lastly, AI16C is planning a $20 billion AI superfund, uh, focusing on vibe creation platforms like Cursor for Vibecoding. So those are the big news of the week. Before we go into each one of those specifically, though, um uh Josh, maybe I'll just throw it to you. Uh, set the tone for us this week. Uh, how was this week as a vibe in the AI space?

04:12
Josh Kale

I feel like over the last couple of weeks, we've been kind of increasing tension between model producers, right? So what do I mean by that? It's obviously a very highly competitive market. And ever since we had Deep Seek come out and compete with OpenAI, you know, OpenAI was the darling child, right? It was, you know, way far ahead. No one was competing with it. And now we have these guys that are like, you know, competing every week now, it's become a meme on this show that there's a new frontier model drop. Um, I feel like we've reached the boiling point this week, David. Controversy with a top AI model producer, not just from any company, from Meta themselves. You know, we've got the Shopify CEO announced that if all his employees don't adopt AI tooling in their workflow processes, you're probably gonna get axed. You know, tension is getting to a point where like AI might start replacing jobs, it might start having significant impact. And I feel like people are really feeling it this week.

05:09
David

Josh, what do you think about just the acceleration of the AI arms race? Because like Jaws says, every single week on this show, we are like, oh, there is a new, like a new leapfrogging of AI models. There's a new number one. And it's been consistent. It's been consistent for the past like 10 weeks. Um how do you think that when it's gone gone on for so long, how do you think that finally is manifesting in like in its in the current day?

05:32
Ejaaz Ahamadeen

It is unbelievable that we're still getting 10x improvements week over week. And that happened this week with the meta um release. It was controversial, but the main headline that I was excited about is a 10 million token context window. And previously, Google's Gemini model had 1 million as the flagship. So we've got 1 million.

05:50
Josh Kale

Which was

05:51
Ejaaz Ahamadeen

Which was last week. So week over week, we've 10x the context window, which is a very big deal because that is the um like quick access memory for a large language model. And it allows you to get

06:01
Ejaaz Ahamadeen

a lot more data packed into your queries. So to go from one million to 10 million week over week, we are still accelerating so, so quickly.

06:09
David

Yeah. Let's unpack context window just pretty thoroughly here for people who are like learning about a context window for the first time. Josh, how would you define a context window?

06:17
Ejaaz Ahamadeen

Okay. It is um it's fun to look at it like if you had a book. If you have a really big book of pages and you're able to rip out those pages and place them out in front of you so you can see them all very clearly, the context window is the amount of those pages that you can see in clear view without having to actually turn the page and find something. It is just quick access.

06:36
Ejaaz Ahamadeen

data to whatever you want to query. So imagine it's like not your eyes, but it's an AI and it's a camera and it could see everything that's laid out in front of you. That's kind of the context window. And it's helpful to have more pages because it can see more things quickly. And it's also more accurate. It doesn't have to infer what the next token is because the next token is clearly visible in that context.

06:54
David

Hmm. Okay, so if I'm may maybe a human way to turn this into a metaphor is like, say I'm doing mental math and I'm holding numbers in my brain and I'm trying manipulating the numbers, I'm thinking about the numbers, I'm retrieving numbers differently in a sequence. A context window means I can think about a larger array of numbers all at once before I like lose them in my mental space. Is that a good analogy?

07:18
Ejaaz Ahamadeen

Yeah, that's right. Um, and you can store a lot more. So if it is a very complex model where there is a lot more numbers, you can store them all in one place. And you don't have to guess. It is all very clear because it is all seen out in the general view.

07:32
David

Is is that the big reason why this a uh Llama 4, the new release from Meta, that's why this is such a big deal, or was it more of a collection of things? Maybe Idraz, I'll th I'll throw this one to you. What were the big standout uh components of this release from Meta?

07:46
Josh Kale

Sure. So the first standout was what Josh just mentioned, which is the 10 million context window. Actually, to build off your example, Josh, it's not just one big book. It's the case or equivalent of 75 novels. Or another way to look at it. 75 novels of a context window. Of a context window in a single context window, or you could think of it as one million lines of code, right? So that's like one of the major things. You know, it's it's easy to recall memory, data, etc. The second most important feature of these models, and and I say models because they technically released three models, um, two which are kind of like basic, and then one gigantic model, which is like a two trillion parameter um beast. Uh but the second most important thing is it's a mixture of experts design, David, which means that at any one point when you're querying the model, you're only querying around 14 to 17 billion parameters. So it's hyper efficient when you're like querying it. And for the bigger model, it's obviously a larger.

08:49
Josh Kale

uh query. So it's gonna probably hit around 250 billion parameters, which is a lot more computationally heavy, but it's still way more efficient than how OpenAI runs their models, where they just query the entire trillion parameter model. Uh and it takes a while to like come back to you and stuff. Yeah.

09:05
David

We've defined um parameters on the show before, but uh maybe it's worth it just to take some time since we're in the mode of defining things. Like to me, a parameter is a specific uh like node of information or a neuron, like a it's a unit of information or a node of information or a or a neuron neural connection in the brain. And when you have more parameters, you just have like a higher resolution camera pointed at the internet, which all of these AI models are uh trained on. You have a higher resolution camera pointed at knowledge at the internet, and these parameters with their weights are set to a specific number to actually be a map of knowledge that is digested from the internet. And so if you have a higher parameter knowledge, uh uh higher parameter model, you simply just have a more knowledgeable model. There's truly just knowledge is contained in these parameters. That's kind of how I think about it. Does anyone want to amend that definition or add to it?

10:01
Ejaaz Ahamadeen

Yeah, that's pretty good. Um, I think with parameters, they're also known as weights. Uh the way it works, and this was a big debate early on, is is does it actually improve with scale? So does it actually increase directly exponentially with the amount of data you feed it? And so far it's been yes. So so far, the more

10:19
Ejaaz Ahamadeen

Data you can feed into this, the higher resolution the model will be. So in the case of weights, it kind of goes through this pre-training phase where it's fed just a ton of data, a ton of information. And then each one of those training runs that happens actually alters those parameters just slightly. And then it compares the results of the new run to the previous run, and it gets better and better and better. And that's kind of how they tune it. And generally speaking, the more parameters, the smarter it is. So the fact that this has, I mean, the largest one has two trillion parameters, that is huge. That is a tremendous amount of training data. I don't know where they found it from. I don't know how they saw two trillion things, but hopefully that will result in very, very high quality answers.

10:56
David

Mm-hmm. Okay. Okay. So now we have a 10 million token context window. And again, a token is just like a unit, it's like a syllable of sorts is almost a word. And then we also have an increase in parameter size, but also three different models came out. So between all of those things, that's kind of some summarizing why the why this announcement from Meta is what it is. But there was some follow-on controversy downstream of this announcement because people Meta, maybe to set the context, Meta has been lagging as a model, as a as an AI lab. Like the A no one really uses the Llama models as far as I understand it to be. Like everyone's using either ChatGPD4 or Claude. And so Meta has been just like losing this game. And with these release of their new models, they are hopefully trying to leapfrog their place into number one, number one. That would be the easy way to like talk about this announcement. But it draws, um, you're telling me that there's been some controversy with the release of this uh model. Can you can walk me through that?

11:54
Josh Kale

Okay, so t to summarize this, every AI model that is released

12:00
Josh Kale

Doesn't matter which company uh produces it, are typically uh graded across what we call benchmarks, right? And uh they're known as SOTA, SOTA benchmarks or state-of-the-art benchmarks. And typically every model producer attributes it to this. Um the reason why they do this is, you know, how do you know whether one model is better than the other? You can't really kind of, you know, test or query it. You can ask them the same questions, but you're gonna get different answers. But how do you really know? Well, there's a measurement of different things like parameters, like, you know, we just discussed weights, et cetera. Um, and then there's the actual output of the actual model, right? So when Meta released these three models uh this week, um they had all these different grades across their benchmarks. And they were claiming uh David and Josh to be as good, if not better, than Deep Seek, V3, and V1, or R1, uh, and OpenAI's O3 model, which is like a really high performative reasoning model, right? Um, but for the first time ever, there was a huge amount of backlash about these benchmarks for two reasons. Number one, uh it was stated, uh, and there's like a few screenshots that we've seen that you can share on the screen here, David, um, that they had blended test results across different models to basically give a falsified or rather skewed result for their benchmarks across their models. So basically, they kind of blended data sets to give a falsified answer just so that they looked clearer or nearer to the actual state of the art benchmarks. The second thing was, and this is the time uh tested trial, is people just used the model and was like, this is kind of shit.

13:41
Josh Kale

And it isn't it isn't as good as OpenAI's reasoning model, or it's not as good as Claude's model, or et cetera. And so people were just like, is Meta lying about this? Now there was a lot of back and forth, a lot of rumor, fear-mongering, etc. Um, their head of generative AI, Ahmed, uh, actually came out and stated that this isn't got anything to do with their benchmark. Their benchmark testing is, you know, to a high quality, a high grade. However, there is an implementation issue. So if you look at this tweet that he says here, it says, you know, we've also heard claims that we've trained on um test sets. That is simply not true, and we would never do that. Our best understanding is that the variable quality people are seeing is due to needing to stabilize implementation. So he's putting it down to the fact that, you know, inferencing is taking a while, our models are overheating, etc. That's the kind of like equivalent of what he's saying.

14:33
Josh Kale

Yeah. So big amount of controversy as as to whether this is real or not. If it is true, that is a huge dark mark

14:40
Josh Kale

on Meta's, not just brand, but the entire AI model.

14:44
David

Josh, what do you think is going on here?

14:45
Ejaaz Ahamadeen

It also feels kind of like a dark mark on the industry. I'm not super uh well up to date on the meta situation in terms of benchmarks, but benchmarks as a whole very much feel like they're broken because they can be gamified. I think benchmarks early on were a really effective way of measuring AI models because so so much of it was new and they were actually challenging problems to solve. But now that AI models have gotten so advanced, it's difficult to generate these benchmark problems and standardize them because the process of standardizing them makes them cheatable.

15:14
Ejaaz Ahamadeen

Um, so I think it just it starts a broader conversation of how you actually measure the intelligence of these models as they get better. Because there is the option where people can gamify them. And when they do gamify them, it looks great on paper. But when you use them, you're like, this doesn't quite match up. So that's kind of how I've been gauging the models that I've been using, is mostly through either my own use or just smart people who use them differently and how they react to them. I don't think benchmarks are a super reliable way. And I think there's a lot of work trying to figure out how to make them better. But in this case, I don't know if Meta did them, but I think a conversation needs to be had about how we actually benchmark these models properly.

15:49
David

I think that's just fundamentally true for benchmarks as a concept. There's something out there called uh Goodhart's Law, which is just often stated as when a measure becomes a target, it ceases to become a good measure. As in when if we are comparing whose model is best using this particular test, then all of these AI labs are incentivized to make a model that's good for that test and not really care about whether that thing is actually useful or not. There was a blog post on Less Wrong, uh LessWrong.org. Less Wrong is um kind of this like rationalist corner of the internet. A lot of AI safety people come from here. Elie Azer Yudkowski comes from here. Uh and a blog post was posted um uh on the 24th of March. So actually not terribly recently, uh, but this also showed that this actually has been a growing conversation in the AI industry. The blog post is titled Recent AI Model Progress Feels Mostly Like Bullshit. Uh and it starts off with two main claims that either AI labs are cheating and they are cheating their way to find a good benchmark number to report a good number, uh, or they are just uh accidentally testing, they're actually done accidentally gearing their AI models to be good at taking that test, to be good at that benchmark. And then it also kind of goes through some of the incentives here. Uh one that really stood out to me was uh talent acquisition, because there's not that much talent, AI talent to go around. Uh and attracting very good talent for your AI models, very, very valuable. And no good AI talent wants to work for the fourth or fifth or sixth best AI model. And so if you can, if Facebook can show who again, we started off this conversation saying Facebook has been lagging, they have been showing poor performance on their models. Now they're number one. Now they're performance number one, but people are looking at these models being like, this is not a useful tool for me. Uh and so so just there's a big a growing conversation about like, yo, benchmarking sucks. And also there's just a ton of warped outcomes and warped incentives around this whole benchmarking complex. Did anyone any of you guys read this blog?

17:50
Ejaaz Ahamadeen

I did not know.

17:51
David

Just

17:51
Ejaaz Ahamadeen

Okay. Well I mean I haven't told you.

17:53
David

everything you need to know.

17:54
Ejaaz Ahamadeen

Yeah, but it sounds um it sounds right. And the w the re like to benchmark a model that's supposed to be reflective of the collective intelligence of humanity by using a few math problems, it just seems wrong. So yeah, something needs to be changed.

18:08
Josh Kale

I mean, would you both agree that the best way to just assess a model's capability is whether it makes material impact on your life? I feel like that's that's pretty much it, right? Like what if we just had a test bed of people across a variety of different uh economic standards and professions, and we just said,

18:25
Josh Kale

here's a new model. Let me know uh what you think about it or what the general feedback is. Give me a rating of one to ten. Is that a dumb idea?

18:33
David

Yeah. Yeah.

18:34
Ejaaz Ahamadeen

Feels closer. It's more subjective.

18:36
David

No, no, no, sorry. That was a yeah, yeah, I was agreeing with you. Yeah, it's like yes, I agree with your point.

18:41
Ejaaz Ahamadeen

Because it is it's increasingly subjective how people use models, how that how it affects their lives. Most people don't use them for code or math problems. They want them for general purpose stuff. So yeah, I'm all in favor of that.

18:51
David

The way that we benchmark GPUs, I think, is actually genius. Where we, if you want to compare like the next higher end GPU, the NVIDIA like 590 to the 4090, you actually just benchmark it using real games, like high intensity, uh what are what are the most like triple A graphically intensive games? And then you just do a frame rate comparison on how like how good that GPU is for that game, which that is true usefulness because if it spits out higher frame rates on a higher end game, well, that is useful for gamers who want exactly those properties. And so I think we are need some sort of benchmarking system like that where it we're actually testing against uh things that people want rather than like standardized tests. Because we all know they the failures of standardized testing in in the school, in the school system, right? Like anyone who is just like wealthy or has the means will practice the ACT or the SAT 10,000 times and then they will get a good SAT score, but they're actually just an idiot who has money.

19:51
Josh Kale

Yeah.

19:51
David

Yeah and yeah.

19:52
Josh Kale

I I actually just thought of a uh a specific distinction between that, David, which is and the standardized tests that we're used to, they're normally theoretical or paper questions, right? It's like hypothetically, what if this happens, you know, or give me the answer to this, you know, random equation. Um, but

20:12
Josh Kale

With AI, it's actual actually practical. You can see it happen immediately in real life, right? So it's like it's not like, you know, uh, tell me the weight of these 10 oranges and then use that to calculate the mass of the sun. It's like, no, I'm gonna show you right now whether that makes sense. And if you're wrong, you're really wrong, and it's gonna taint your reputation almost entirely. And it and also, I'm gonna do this calculation in like 15 seconds. Here we go. Right. So I think people aren't really ready for that feedback loop, and it's obviously showing with meta.

20:41
David

Right.

20:42
David

Any last comments on the subject before we move on?

20:44
Josh Kale

No.

20:46
David

All right, moving on to one of the big uh releases of the week, I would say, is this uh website, which is kind of just an interactive document. Uh, the document titled AI 2027. And these are it was authored by a collection of just AI legends. Um, Scott Alexander is the notable one to me. He is the author behind this Late Star Codex blog, which is this uh very well-respected, well-read uh for nerds blog that Silicon Valley tech leaders read. Um, and uh if bankless listeners are familiar with the concept of Moloch, Moloch came out of this Late Star Codex blog. A bunch of other things came out of there. Um and so Scott Alexander and four other people read uh wrote this AI 2027 uh model for how they think, a plausible model for how AI development goes from here. And they give dates, right? So mid-2025, stumbling agents is where we are in current AI uh landscape. And I think anyone who just came out of the AI slot bot meta understands exactly where we are with stumbling agents. Uh and then they kind of give a prediction late 2025, what that looks like, uh, early 2026, coding automation, mid-2026, China wakes up, late 2026, AI takes some jobs, January 2027, uh agents never stop learning. And they give this predictive idea for where this could go. And they also talk about incentives from they do they bring in anything that's relevant. They talk about uh AI copycatting uh Silicon Valley Tech Labs, they talk about Department of Defense uh realizing that they need to pair with their own domestic AI labs as a matter of national security. They talk about uh China trying to figure out their ways around the chip uh shortage. Uh and really where this ends up at is where they think this is going is like a uh a superintelligence explosion by the end of 2027, where the governments of the world

David Hoffman

1490 posts

Co-owner at Bankless. Optimistic storyteller of frontier technology.

No Responses