NEAR - Sponsor Image NEAR - Confidential swaps across 35+ chains Friend & Sponsor Learn more
01:28:21 · 1 year ago
Podcast

AI ROLLUP: The AI Experiment That's Been Secretly Manipulating You

The Reddit Bots Are Out For Vengeance

Up next

All episodes

Inside the episode

OpenAI’s new o3 model dropped with a bang—and a Mensa-style IQ score of 136 that places it in the top-1 percent of human test-takers. Early benchmarks show the system surpassing GPT-4 Turbo on reasoning and code tasks, but outside audits are less unanimous, with researchers still debating exactly which skills the score captures. 

Yet o3 isn’t flawless. Independent testers are already catching the model in “complimentary hallucinations,” flattering users while smudging facts—fuel for the wider debate over chain-of-thought (CoT) prompting and whether hidden system cues block end-to-end reasoning. OpenAI’s own scientists argue the model proves intelligence scales with inference time and reinforcement learning, not just bigger pre-training runs—hinting that smarter, leaner agents are still ahead. 

On the developer side, OpenAI open-sourced a natural-language coding agent that can rebuild a Photobooth-style app from a single screenshot, fix bugs on demand, and slot seamlessly into VS Code. To accelerate that vision, the company is negotiating a $3 billion takeover of Windsurf (formerly Codeium) after two failed attempts to buy Cursor—evidence that OpenAI wants to own the full application stack its models power. 

A quieter change may prove just as consequential: “Memory with Search.” ChatGPT now rewrites user queries based on remembered preferences before sending them to the web, promising relevance but raising Black-Mirror worries about echo chambers shaped by private data you never see. Meanwhile, an ImageGen API and murmurs of an OpenAI-built social network suggest fresh revenue streams and viral-loop ambitions to rival Perplexity and Meta’s new standalone AI app. 

OpenAI’s momentum hasn’t deterred rivals. Google rolled out Gemini 2.5 Flash, a cheaper, toggle-able reasoning model; Chinese startup DeepSeek’s leaked R2 boasts 1.2 trillion parameters and a 97 percent cost drop; and Alibaba’s open-weight Qwen 3-235B is already edging past Llama 4 on key leaderboards. 

The arms race is also decentralizing. Prime Intellect just launched INTELLECT-2, the first globally distributed reinforcement-learning run of a 32-billion-parameter model, with experts predicting community-trained systems in the 70-100 B range by year-end—a potential counterweight to hyperscaler dominance. 

Amid the hype, academia offered a cautionary tale: University of Zurich researchers secretly deployed 13 AI bots on Reddit’s r/ChangeMyView, logging 1,500 comments and swaying opinions six-fold over the human baseline. The backlash—accusations of unethical manipulation and a possible ban on publishing the study—underscores how persuasive these new systems already are in the wild.  

Transcript
00:04
David

Welcome to the AI Rollup where we stay up to speed with the emerging trends and developments in the AI space. I'm David Hoffman here with my two co-hosts, Ijaz and Josh. Josh, you have been gone for a while now. Uh hiking in Peru, doing normal human things. Welcome back to the world of uh technology and AI, my brand.

00:21
Josh Kale

Thank you. Yeah, I heard in order to be a co-host on Bankless, you need to hike a big mountain. So I went and I did that just like you did. And um throughout the entire time, there's this like low level underlying anxiety. Um, because I was very far disconnected from the world, knowing that there is so much news happening every single day. And I was gone for like nine days probably off the grid. And I just got to That's a whole cycle.

00:41
David

AI time.

00:42
Josh Kale

Oh my god, I've probably missed like four different frontier models in that time. So it was a little anxiety inducing, really great time, but like super excited to get back into the the trenches, into the matrix, and like see what I missed. Yeah.

00:53
David

I've been

00:53
Ryan Sean Adams

I'm actually curious, Josh. Is there actually any evidence that you were where you were? I've just seen pictures. It looked weirdly AI generated. I'm not sure.

01:04
Josh Kale

We're getting close to the point where it it might I could probably fake it if I wanted to.

01:09
David

for nine days.

01:10
Josh Kale

My videos look slightly better than Zora currently, but I think we're probably six months away from me being able to fake an entire trip. So stay tuned for that one.

01:19
David

Uh well, over the last nine days, uh, I've been focused on uh Ethereum land on the crypto side of things, and that just actually leaves Josh as the one person informing us as to like what the hell happened over the n that last nine days. Uh Josh, how are you doing this week, my man?

01:34
Ryan Sean Adams

Doing well. Josh, you're not too far off the truth. I think uh three Frontier models dropped, including uh the big dog's open AI. Uh, there was an incredible amount of drama from a new product that released that helped you cheat. And a lot of people were on the side of it, and a lot of people were against it. Uh, we have a new open source model from our best friends over there in China uh that kind of beats some of the current models that we see right now, and a heck and a heck of a lot more. So I've just been glued to the computer. I'm kind of jealous of you, Josh, that you got to be out on a big hike. Uh, but I'm excited to dig into things.

02:08
David

All right. So starting with uh we got five big topics we're gonna talk about today. The AI roller coaster. We're gonna talk about uh the great Reddit experiment, how a bunch of researchers duped thousands of people on Reddit using AI. Uh there's DeepSeek that uh it draws just hinted at. We're gonna talk about uh the capacity for AI to cheat and how that's coming no matter what. And then also on the crypto AI side of things, Primes and Prime Intellect's training a 32 billion parameter model on a decentralized network. Crypto coming in hot on this episode. But it draws maybe just to start us off you know, where we feel safe. You know, AI, open AI, uh that's the homegrown uh territory. That's classic AI. What's going on in the world of open AI right now?

02:48
Ryan Sean Adams

Yeah, so um out of fear of boring people on this podcast, I'm gonna do it again. Um, OpenAI has released yet another frontier model, but it's two frontier models this time. Okay. So um in their official release, they say introducing OpenAI 03 and 04 mini. Now, I can bore you with a bunch of details and a bunch of really cool benchmarks, which by the way, David, you should scroll down. There's some really flashy uh illustrations down there. But the TLDR is O3 model, um, is their new flagship model. And it is uh a super genius at reasoning, at coding, and a bunch of math stuff. So the typical set of like benchmarks that we typically you know see or measure across other uh models. And when compared against other you know competitor models like Meta's Llama or Anthropic's Claude, um it also beats them, which is a major significant update since you know Deep Seek was coming for OpenAI's throat. Um

03:49
Ryan Sean Adams

About two months ago, and Anthropic was coming for their throat in terms of like coding. And now OpenAI surpasses them with this new O3 model, right? There's actually a great breakdown from Rowan Cheng over here, David, that talks about the kind of new models and how it kind of levels up against everything else. One big thing that I want to point towards is that the biggest change is that they can now use images as part of their reasoning process, which doesn't sound like a major thing to start off with, but is actually quite a monumental leap because previously models trying to digest images and diagrams and contextualize that into someone's response wasn't really an easy thing to construct within a model, right? It was very much text and character prompts. Um but now we can use and leverage images and

04:37
Ryan Sean Adams

Eventually, even video to help give you like the right types of answers or context towards whatever your prompt is. So Rowan then goes on to explain some other key updates where he explains that you know both new models were able to use ChatGB tools independently. So like you've got web browsing, Python images. So what it's what he's getting at is these models now collectively combine all of OpenAI's separate tooling that they've built in the past. So that's includes image generation, which you spoke about a few weeks ago, uh, which led to you know the Studio Ghibli trend. Um that includes their coding agent. Um, and they actually released a new coding agent. Um so yeah, all the lovely um kind of different tools are now collectively combined in this one supermodel. Um

05:22
Ryan Sean Adams

And it's been pretty awesome to use. It's completely leveled up my way of doing research. It's given me some really cool uh case studies that I can use. For example, I asked it to basically conduct a research study on all different kinds of nutrition types for, you know, kind of like a meal plan that I'm prepping for. And it added a bunch of imagery diagrams. It gave me a really conducive report. And I was then able to kind of like use that to construct an app via their coding agent, which I could then like kind of

05:54
Ryan Sean Adams

use locally on my laptop to advise, you know, my routine on every kind of stuff like that. So uh for someone who you know hasn't been coding his entire life, it was a major step gain function.

06:06
David

Yeah. The pattern that I'm really seeing here, this doesn't really feel like it's a groundbreaking release. It seems like this is one of the more incremental releases where there are some like there's still a lot of low to medium hanging fruit. Let's make sure that our AIs can interpret images. Uh let's make it, you know, think it can do the things that normal humans can. And I'm just seeing like AI start to more closely mimic the capacity of humans. There are some things that open AI is very good at. There are some things or you know, AIs are very good at, and then there are some things that it's just like, why can't you do that yet?

06:38
David

And uh this seems to be like fixing some of those low hanging fruit things of just like, yes, now you can interpret images too. Yeah. And these are like pretty expected updates. Um

06:46
Ryan Sean Adams

Exactly. And like the the way I would think about it is imagine you go to high school and there's always been that smart kid, David. And then there's a new kid that's joined, and he or she is even smarter, right? And that's basically what this new O3 model is. And actually, on the topic of intelligence, um, someone uh mentioned or measured its Mensa IQ for this uh particular model. And the purpose of doing that was like uh there a discussion ensued after this model released where they were like,

07:16
Ryan Sean Adams

Have we achieved AGI? I'm not entirely sure. Uh and if you look at this Mensa IQ, it scores a 136 uh on the Mensa Norway uh IQ test. And I don't

07:25
David

is it Mensa is just like a way a methodology for measuring IQ?

07:29
Ryan Sean Adams

correct. Yes, yes. Um, and uh if I were to take this test, for example, I'd probably be on the low end of that spectrum, probably around 80 uh uh IQ or less.

07:38
David

Yeah, I think you'd be

07:39
David

the top of the bell curve.

07:40
Ryan Sean Adams

But but but uh but this new model scores you know almost.

07:45
Ryan Sean Adams

Double of that for me. So, so, and you can see how it kind of weighs up against all these other types of models. So, you know, it's doing pretty good. Um, one thing that their head of reasoning, Noam David, points out, uh, if you pull up this tweet, um, and I think this is really important to note because I think it went over a lot of people's heads, he goes, our new OpenAI 03 and 04 mini models further confirm that scaling inference improves intelligence. And that scaling RL, which is uh reinforcement learning, shifts up the whole compute versus intelligence curve. And there is still a lot of room to scale both of these further. Now, what he's pointing out here is previously the traditional form of training models is to put a hell of a lot of compute into the pre-training and the post-training process, particularly though, the pre-training process. That's where you get all the figures of like multi-billions of dollars of compute and hardware spent to try and train up these models. But what he's pointed out here, or what he's saying, is with these new models, actually, the way that we got them to be way more intelligent was via inference and reinforcement learning. So that was like repeated prompts of these models and saying, hey, have you thought about this correctly? Different kind of reward functions saying, hey, if you get closer to this kind of an answer, we'll give you a higher reward and getting the model to think for itself. And what he's saying here is the classic projections that we've had so far on how this kind of intelligence scales up has actually been incorrect. Because if we factor in inference and reinforcement learning, the scale goes even more vertical.

09:20
Ryan Sean Adams

So we might achieve AGI much quicker than we expected.

09:25
David

Maybe to understand that a little bit more, there's pre-training and post-training. And pre-training is you you scan all the internet, you turn it into uh what is it, what are they called? Um what are the little phenomes? Tokens. You you turn you turn the intelligence and the the data of the internet into tokens. Uh you create the parameters, which is like the the the more parameters, the higher resolution image that an AI has, an uh LLM model has over the knowledge of the internet. And that is pre-training. And it's just, it's just this like raw, unbridled, crude intelligence that's not refined, it's not directed, it's not, it's just the oil without the engine, right? It's just it's all gasoline, no combustion engine. And then the post-training is like, okay, now we now that we have the gasoline, we have raw intelligence. How do we extract it? How do we actually turn it into productive output? You need post-training. And then you combine pre-training and post-training, and then you have an AI model. And I think what I'm hearing from you, Ejazz, is like, okay, we have spent lots of resources on pre training. Yep. And we are realizing that investment into post training uh is actually returning a lot more uh value, is getting us to whatever AGI is a lot more than just adding more uh resources into pre training. That's what I'm hearing from you.

10:45
Ryan Sean Adams

Yep, that's exactly it. And we essentially saw that trend begin when DeepSeek in China released R1. And we realized that a lot of its intelligence derived from reinforcement learning and reasoning. Um, but I know Josh knows a heck of a lot about this. Josh, I'm wondering if you have any commentary here.

11:02
Josh Kale

Yeah, well, the exciting thing about this chart is that it since the beginning of AI, people have been waiting for us to hit a progress wall where like eventually you would put more resources in and it would not scale proportionally. And what this shows is that's still not true. And we're still progressing and there is no wall and we're still ascending up this curve. And in fact, one of the best things to happen is Deep Seek and R2 leaked, and we might discuss that a little bit later. But what China has is they have constraints that we don't have. And the United States, we have $500 billion investments and infrastructure, and we're throwing a lot of money at the problem and we're we're solving the hardware problems that are very expensive. But then we have the other side, which is China, and they're solving the software problems by being resourceful. So we have these two oppositional forces one is resourceful and scrappy, the other is well invested, well funded, and

11:47
David

Course.

11:47
Josh Kale

Brute forcing, exactly. And what we're seeing now is like the convergence of these two things where we have the brute forcing from the United States and we have the resourcefulness of uh recursive learning through Deep Seek, and we're starting to see this like compounding acceleration. So we're not actually even hitting a wall, we're actually accelerating even faster because we have these two converging forces happening and they're kind of learning from each other. So it's it's really exciting for me to see because we are just continuing up this trend of um.

12:14
Josh Kale

Of the like training curve.

12:16
David

Intelligence curve. Yeah. Do you think that we have found a wall on the pre-training side and then we had to route around that by starting to optimize for post-training? Or do you think that we actually have not found a wall in either case?

12:17
Josh Kale

Yeah.

12:28
Josh Kale

I think we're gonna find out really soon with GPT-5. I think that is the new large model that has been trained on, we'll probably have a gigantic uh parameter count, well over a trillion. And that will be a really good gauge to see if that trend still does continue on the pre training phase. Um, because we haven't gotten a new really large foundational model in a very long time, similar to like a GPT style release. And I think GPT 5, which is hopefully coming out in the next few months, will show that.

12:56
Josh Kale

So that's what's going to be really exciting. But we mentioned a little bit earlier that O3 felt incremental. And

13:01
Josh Kale

I used it. It came out right before I left Peru. And it feels like a bigger deal than that. And I'm not sure it shows that on the benchmarks. But when I was using it, it it was so good that I've just kind of stopped using everything else. I stopped using basically all of the other models except for Grok when I want real time data. And I think they've kind of captured something.

13:22
Josh Kale

By adding all of the tools in natively that I haven't experienced anywhere else, where I can just kind of ask anything and it will be resourceful in a way that I don't need to infer. So, like one example, I was in the middle of Peru. I was trying to figure out like what the hell this thing was. It had Spanish writing on it. It was like, I took a fuzzy photo and it actually it zoomed in on the photo, it added clarity, it you learned that it needed to use the translation function. It wrote a little bit of code to create that, and then it just did it all and it spit out an answer. And that's something that I haven't really.

13:53
Josh Kale

been able to do with anything else. So O3 feels better in a way that's not really tangible but kind of vibes-based. I'm like, oh wait, this is really good. I actually don't want to use anything else. O3 is like really, really good. So that's

14:05
David

That has been my experience with O3 as well. And this is why I'm not interested in as much like pre-training, because like we've identified, there is this raw intelligence, this crude intelligence. And now it's much more about these front ends being intuitive about what I want and being clever and witty and resourceful about understanding what I want and being clever in how to deliver what I want. And that is a much, that's much more of like how how intuitive and clever can the engineers who are like bridling this raw energy and uh intelligence energy and what and turning it into a form factor that is useful for the query inputs. And that that is not about more pre training. That is strictly uh ingenuity and intuitiveness in the post training side of things, which is just the front end and the delivery of the prompts.

14:58
Josh Kale

and the way it's presented too, I noticed that O3 uses a lot of tables and it kind of organizes data much more cleanly. So in the case of building out an itinerary for a few days of the trip, it really neatly laid it out in a way that I hadn't seen before. So yeah, the way it's presenting the data, the way it's using all of its tools without asking you which tools to use, it just creates this really cohesive singular model that feels really powerful to use relative to the others that kind of have pieces of that, but not the whole thing.

15:25
David

And so Josh, when you're talking about ChatGPT 5, which comes out and you said a couple months, you are just strictly looking at is this an incremental gain or is this a large step function forward for the output of of AI AI models?

15:39
Josh Kale

Yeah, I

15:39
David

I think the

15:39
Josh Kale

hope is that like this this will be the new largest base model. And and if it actually does prove that having that larger base model to sit underneath this new learning that we've done, this new recursive learning, this new pro-training era, um if it actually is

15:55
Josh Kale

Um scaled in the sense that the more compute that we still throw at it, the greater the output is. Then I think that's going to be a really exciting thing because it will probably mean everyone's going to keep their foot on the gas and train these foundational models. Should it hit a wall, should it not scale proportionally? I think then we kind of look to the post training phase using the recursive learning, using these new forms of technology to kind of refine the base models. But there's only so many people in the world that could build a base model that large because it costs a tremendous amount of money. So I think they'll probably lead by setting an example. And that will be super interesting to see what the future of this base training model will look like.

16:32
Ryan Sean Adams

I I have a slightly different perspective on your takes. So so I heard a few things from both of you, which mentioned context. Um

16:41
Ryan Sean Adams

And data. And I think those two things are actually going to become the most important properties to developing a frontier model. So let's assume that like pre-training and post-training takes a ton of compute. And, you know, uh, to your point, Josh, you know, you need to be able to spend the most basically to get there. I think, you know, all of that will eventually become commoditized. But you kind of want the model to understand what you're asking for before you even ask for it and present it to you in a way before you even kind of like know what the best way is to present it to you. I think most of that comes through real-time data. Um and it's interesting that you mentioned, Josh, that you you just use grok right now for real-time kind of like media or updates, right? And so do I, right? I'm like, you know, what are the latest updates uh that I might have missed? Blah, blah, blah, blah, blah. But I was thinking, man, if I could just link my OpenAI memory to my Grok account, I would have like a supercomputer basically in front of me. And I think that might be eventually where the moat ends up becoming. Um, you know, we we spoke about OpenAI's memory update a few weeks ago. They actually made another subtle update, which wasn't widely publicized that we're gonna talk about in a few minutes' time, which was pretty dark. Um, but I I think memory and data and context specifically might be the way that it goes. Um, but we'll yeah, we'll see. I I I think we are gonna eventually hit a wall. But um, you know, we have Meta training a two trillion parameter model. Zach mentioned that on Duar Questiers podcast yesterday. Uh, we have Google training something around 1.8 trillion, and then we have um OpenAI supposedly also training a two trillion parameter model. So I'm I'm just really excited. Uh, the the best benefactors of this is us. It's just like

18:25
Ryan Sean Adams

Dude, it's amazing. We just sit back and we're like, oh my subscription can do this now. Oh that's uh that's pretty cool. Uh yeah, that's pretty that's pretty sweet. Um Okay, okay.

18:36
David

Jez, keep us going, yeah.

18:38
Ryan Sean Adams

Okay, okay. So

18:39
Ryan Sean Adams

I want to break the illusion.

18:42
Ryan Sean Adams

For a second. Okay. So every week we mentioned these frontier models and we're like, oh my God, look at this amazing thing. Look, it understands me. It gets me so well. But I don't think any of us have really contemplated whether what we're ingesting or reading is to be true, right? We definitely get what we want to receive. We're like, okay, this is the answer that I kind of was thinking about. And like I now have ChatGPT validating it. But I wonder, like, you know, could it ever be lying? Well, interesting that O3 was caught lying. Um, uh, if you pull up this tweet from um Transluse AI, which were basically an entity that got access to O3 before it got officially released, um, they tweet, we tested a pre-release version of O3 and founded that it frequently fabricates, sorry, fabricates actions it never took and then elaborately justifies these actions. So it doubles down basically when confronted. Um, and so they dig into this, right? Um, and they realize that it's not just an OpenAI O3 model behavior, but it is a behavior that they find across many other models, right? So um, and they just use open AI as kind of like a test bet here, right? And we haven't really quite come to an answer as to why these models are lying or why they're doubling down on these hallucinations. Um, one really interesting example before I I want to get you guys' feedback is again, Transu says this means um.

20:11
Ryan Sean Adams

O series models are often prompted with previous messages without having access to the relevant reasoning. When asked questions that rely on their internal reasoning for previous steps, they must then come up with a plausible explanation for their behavior. So basically, what it's saying here is when OpenAI releases a model, they have this kind of uh what did you call it the other week, Josh? It's like a hidden prompt or something, basically like a

20:35
Josh Kale

Some prompt.

20:36
Ryan Sean Adams

A system prompt, right? So it's hidden. OpenAI has already fed um the model this prompt. You don't see it. You just see your chat interface with a fresh conversation. You start talking to it. But really, this model has been pre-prompted to be like, you know, don't be racist. Don't say anything, you know, above this unpolitical agenda or whatever. You know, it's kind of been censored to an extent if you want to take the extreme. It's been guided. It's been guided. It's been guided. It's been shepherded, you know, by people who we have no idea, you know, what their moral alignment is, but it's being guided. And the the explanation behind these hallucinations that, you know, transluse AI is giving is the AI is coming in, right? And it's taking in your prompts and it has this new reinforcement learning reasoning prompt, right? And it's like, okay, I'm gonna take David's prompt and I'm gonna think about really what he's asking me. And I'm gonna try and think, okay, maybe I should go down this path. No, maybe I should try this path. And it's learning as it's, you know, you see when you use O3, it's like reasoning, it's thinking, right? But what it doesn't have access to is why its system prompt was written that way. It doesn't have the reasoning of its administrators, the humans who put in this prompt. And that might be why it starts hallucinating. It might be it might just assume oh, my administrator used this prompt because it ran this code.

21:55
Ryan Sean Adams

And then it starts kind of spiraling into this never ending loop, which ends up in it lying to you.

22:01
Ryan Sean Adams

I don't know, what do you guys think of that?

David Hoffman

1490 posts

Co-owner at Bankless. Optimistic storyteller of frontier technology.

No Responses