Chapters
Questions from this episode
- Why did Jensen Huang receive the Business Idiot verdict?
- Alex Stenlake named Jensen Huang the week's Business Idiot for declaring that AGI had arrived after OpenAI's Astra posted a claimed 99.9% Arc AGI 3 score. The Business Idiots hosts considered that result dubious because it relied on OpenAI's proprietary harness. Will agreed and suggested Huang's announcement served NVIDIA's interest in selling GPUs.
- Why did the hosts question Astra's 99.9% Arc AGI 3 score?
- OpenAI reported a 99.9% score for Astra on Arc AGI 3 using its own provider adapter harness, which preserved state and context between API calls. The Arc Prize Foundation later tested Astra with a harness that did not preserve that state and recorded 62.7%, prompting the Business Idiots hosts to question the headline result.
- What did the hosts think of OpenAI's Astra in practical use?
- The Business Idiots hosts found OpenAI's Astra highly capable in practical agentic work. Jim Lovell said it could take defined requirements, work independently and return a result without repeated prompting. Alex Stenlake had not tested it as deeply, but praised the built-in migration skill that transfers workflows from Claude into OpenAI's coding environment.
- How could AI bot traffic increase website costs?
- Cloudflare warned that AI bots could drive enormous increases in website traffic and hosting costs. The Business Idiots hosts discussed crawlers making speculative requests across niche sites, rather than sending targeted human visits. They said businesses may need bot protection, while weighing whether blocking crawlers could also reduce their visibility in AI-driven search.
- What did Alex notice about AI advertising in San Francisco?
- After moving to San Francisco, Alex Stenlake observed that bus and street advertising was dominated by SaaS and AI products. Many ads promoted tools designed to make agents more reliable, reduce hallucinations or add knowledge. He saw this as a sign that the market is heavily promoting patches for unresolved AI problems.
- Why were the hosts frustrated with Anthropic and Claude Code?
- The Business Idiots hosts criticised Anthropic's recent Claude Code changes for making established workflows less predictable. Jim Lovell said classifier and permission changes forced him to rebuild workflows and cost him working time, while also adding model usage. Alex Stenlake argued that repeatedly introducing and removing features makes the product difficult to trust.
- What did the episode say about OpenAI's Navier-Stokes proof?
- The episode discusses OpenAI's claimed Navier-Stokes proof, produced with thousands of agents and an unreleased internal model. Alex Stenlake highlighted formal proof-checking tools as essential safeguards against imagined solutions. The hosts also covered researchers' allegations that OpenAI used their distinctive approach and pressured them over authorship. The episode presents those accusations as allegations, not established facts.
Read the full transcript
Will: Hello, everyone. Welcome back to the Business Idiots podcast, where a few of us Aussies are just breaking down everything that's happening in AI and technology. Now, this week, we're moving to an online format now that Alex has landed in the States. So, Alex, welcome back, mate. How's life in the Bay Area?
Alex: It's been an interesting and very fast-paced first couple of weeks. Day two, database surgery while the CEO is on a sales call. You know, good times. If you've never done database surgery on a code base that you only saw for the first time the day before, you've never lived.
Jim: I think that is the definition of living.
Alex: Yeah, well, I mean, look, this is the sort of thing that anyone that knows me would be like, hang on, why are you doing that?
Unknown: You hate that stuff.
Alex: It's like, yeah, but needs must. Series A, you know, it paces everything.
Will: And so tell us, like, what are the biggest differences that you're finding just in the overall, I suppose, environment in San Francisco at the moment?
Alex: you know people would probably say oh well is it you know everyone talking about ai yeah there's a decent chunk of that or you know the robo taxis yeah but i've lived in china it's actually the advertising the biggest difference between here and brisbane is it's the advertising ecosystem From the second you get off the plane, there's kind of like SaaS vendors selling their, in particular, agentic solutions, but data warehouses, sort of observability layers, anything. Like if you've been in tech, all that stuff that you talk about at work is suddenly on the sides of buses, right?
Jim: Yeah. And so is that just because there's so many tech people in San Francisco or you think that's the reason or because they just know they can hit everyone and get a return?
Unknown: Yeah.
Alex: So this is one of the other things I kind of noticed just chatting with some people. Like it's worth noting at this point, I've got very limited engagement with the wider tech ecosystem here.
Unknown: I'm kind of just...
Alex: scrabbling to kind of get a foothold on the problems I've got to do for work, et cetera. But even in the conversations I've been having, there's a real circular... nay incestuous set of financing and sales that kind of happens in San Francisco to the point where I was joking around with one of my co-workers and he was shocked that anyone in the rest of the world would buy half this stuff he's like I thought this was people from San Francisco selling to other people in San Francisco because there's so much silly money here that you know bad ideas get a very long sort of runway and Yeah, that was kind of like digging into that a bit and sort of looking a bit more at the way, you know, what passes for...
Jim: But is that a bit of an explanation to the hype cycle? Because a lot of people would say AI is overhyped and the rest of the world really know nothing about it. Is that a bit of a look behind the curtain as to, oh, well, even everyone in San Francisco is thinking that no one else actually, they're the only people buying this stuff. You know, I think that's actually really interesting because the more and more you talk to businesses in Australia, they, again, think being able to chat with ChatGPT is integrating AI, you know?
Alex: And I think I had this kind of telling experience. You know, I walk to work. It's, you know, a couple of blocks from here.
Unknown: And...
Alex: I count ads on the sides of buses and sort of classify them in my head. And I reckon sort of 70, 80% are some kind of SaaS or tech product. And of the SaaS or tech products, I would say 30% of them are ads for, you know, preexisting SaaS giants that kind of do good work. that have bolted AI onto their product. PostHog and Sendry are everywhere, for example, but Airbyte had a really good one that made me laugh. It was like AI, because our board demanded that we add AI to our product or something, and that was literally their ad on the bus, right? But then the other sort of 70% of those techie categories are things like... Tools to make your agent more reliable. Tools to make your agent hallucinate less. Tools to give your agent knowledge. So not just selling AI technologies, but very specifically selling into people that are selling AI products to patch the problems. Key, to patch the problems.
Jim: That's a digital product with funding. This is...
Alex: If that doesn't scream red flag to you... That's unbelievable, yeah. Yeah, and this is like selling to...
Unknown: I know...
Alex: I don't know, but I think Harvey AI have a presence over here because there's an event next week where I'll go and meet a bunch of people doing this sort of stuff. And there are people... I've had one or two conversations with people doing genuinely interesting work. Using AI agents and stuff like that. But you can kind of feel the gap between the sales and what it takes to actually make this stuff work when you see these ads. Because there's all these people clearly going crazy. crap this thing doesn't work and it needs to work by the end of like q4 etc that we'll see the bus go past and be like yes i'm gonna jump on a sales call with those people right that the branding work actually pays off so i don't know that to me was like a very huh signal there's not a secret over here they're just patching it harder
Jim: That's it. That actually makes me feel a lot more optimistic about where Australia is at. And I think the three of us have been a bit, I say down in the mouth, but you know what I mean, in that we have all these conversations and no one really truly understands what really integrating AI is or doing it properly is. Turns out all the hype masters in San Francisco and Silicon Valley aren't doing it either. They're just patching the problems in the same way.
Alex: I mean, full disclosure, next week I'll probably have a bit better information. But just the fact, the prevalence of advertising is often a good bellwether. The composition of the ads everywhere, yeah, skewed very heavy. Hey, here's how we solve that one problem that you've been dealing with since like 2024.
Will: So when I visited San Fran a few years ago, I got a similar feeling here as well, which is that the advertising is, you know, in Australia, you're kind of advertising to try and get people to come and adopt and buy your tool. I don't think it's necessarily the same in San Francisco because... There is really, there's people building with these AI tools now and then there's people who are funding them. And I actually think a lot of the advertising that you see in San Fran, the physical advertising, the billboards and the sides of the buses are actually about trying to create exposure for a lot of the VCs and the funding in that space so that they're familiar with the brand come series A, series B.
Jim: Which also tracks as well because they all seem to know all of the products. which you would never be able to stay across everything the way they are. But if it is advertised to you everywhere, you would.
Will: Yeah, that's a good point. Talk to anyone from San Fran and they know 400 more products than you know.
Alex: There is a decent chunk of true believers out there. I was chatting to a guy at a current company about my age and, you know, oh, what's the meetup scene? like here like if you try and do some networking and sort of put down some roots he said honestly i wouldn't even bother until the ai boom goes out the door this is full of 23 24 year old kids you know spouting the latest sort of marketed tool and the thing that'll really annoy you is some of them have been given 100 million dollars and they'll treat you like an idiot for being skeptical And so it was like, well, you know, fair cop. I'll encounter that, no doubt, in my own time.
Jim: It's also the same here. I can't believe how often I'm compared to just some idiot who's decided he wants to get on Instagram and sell AI. All of a sudden, 20 years experience goes out the door because this guy who's been here for 20 weeks comes on and goes, oh, well... AI and they have absolutely no idea or they're just bolting a tool in that they don't really understand. I was having a conversation on Friday, just exactly pretty much verbatim to what we had in the podcast a couple of weeks ago where we spoke about data privacy. And I was literally explaining, here's your privacy policy, what you're doing with this data by sending it to Gemini the way in which you are. you are in breach of your own privacy policy. And you could see the penny drop in the meeting and he just went, hang on, we've just paid a fortune for this, you know, and it's all complete waste. But going back to what you were saying before, I'll be really interested to see your impressions of the Harvey team, you know, or those guys. Because, again, Ligora seem to come up to the same level as Harvey, but all the discussion now is Harvey, Harvey, Harvey. So it's whether it's a real drive by them or they've found another way to unlock something more, you know.
Alex: Sorry, we might have to cut this in and out. I'm just trying to get a list of the invites for this thing next week just to give you guys some idea.
Will: I might just respond to Jim's Harvey thing while you're looking at that then. So, yeah, Jim, I was actually able to have a look at how some of these tools work recently, the AI legal tools kind of behind the scenes. What was interesting to me is a lot of them right now are offering a feature where you can kind of request a lawyer from that company to come back and review all of the advice that the AI has given in that chat. So there's like a button to click to request that. But the other interesting thing was... They have some automated checks going on as well where regularly after you finish the session an hour, two, three hours later or overnight, you'll actually get an email from the company saying, hey, we'd like to provide some revisions to the advice provided by the AI. And it's a list of everywhere where the AI was wrong and then what the updated advice is.
Jim: Do you reckon that is driven by an actual lawyer reviewing it, or do you reckon that is a long-running container that's actually been able to take some time where they haven't had to get a quick response or a user-need-driven response? They can actually go and run a batch container overnight. actually do some proper, you know, like council of agents scenario where they can come to a better response? Or do you reckon it's, like I said, a human lawyer giving it a review? It'd be really interesting to see that.
Will: Yeah, so I think it's probably a hybrid of both. So if I was on that team and I was building this, I would have AI reviewing the transcripts to look for things like disagreement between the user and the outputs. And then also... depending on how much kind of compute we had access to, would be reviewing all the transcripts with like a team of other agents that are acting antagonistically towards the advice that was given. And then anything that kind of flagged a certain threshold would then be maybe drafted a response by the AI and then have a human lawyer reviewing that and then actually kind of sending out the final piece of advice.
Jim: No, for sure. And it's actually really interesting in terms of, you know, like that's ultimately what everyone should be doing. You know, is it whatever is created by whatever agentic workflow you have, you should really be having some sort of review process happen based on that.
Alex: I don't think that's a good idea. I don't think you can scale that. And, you know, you've kind of reintroduced the bottleneck you were hoping to avoid. I don't think the tech is there.
Unknown: I don't think the tech is there.
Alex: I think we do need to do it. But if all people are doing is reviewing outputs of LLMs all day, They're going to stop reading, so you kind of lose the advantage anyway. You just get someone to strangle when things go wrong.
Jim: No, don't get me wrong. I wholeheartedly agree that I don't think the actual tech is there, and you're just going to get a slightly less hallucinated answer, but it's still better than the answer you're getting.
Will: I think one of the interesting things is we, one of the early predictions with AI was that we wouldn't need as much kind of middle management layers and all this kind of obfuscation in the middle because people could kind of just, them and their AI could just do much better work. We don't need so much management of the work anymore. But I think that's like almost the opposite now because we've moved the bottleneck from the creation of work to the reviewing of work. And all of a sudden now we need like way more middle managers to be reviewing all the outputs that the staff is submitting.
Alex: I think a lot of that is because we're still thinking like we've got an office with like inboxes and in trays and paper and printers and stuff, right? Like using AI, for example, for reporting purposes, to recontextualize work for different audiences rather than having a team spending 25, 30% of their week. recontextualizing numbers and slide decks to present information to the various stakeholders. Just connect everything and let the stakeholder get their situation updates from the system. If something comes out that's, you know, incorrect, you know, fix it then. But, like, usually those compression processes are massively lossy anyway. So, like, you know, those... Saving that level of middle management, the information flow, I think is good. But when people are talking about middle management, people are rarely complaining about, hey, they're making things more efficient. I think that side of middle management is very hard to find, but it's the actual valuable bit. The moving of information and sitting on the information flows and the sort of compressing and recontextualizing, that sort of thing, I think, is just gone and good riddance.
Will: Yeah, yeah. So given what you're saying earlier then, Alex, about the marketing being kind of the key difference you're seeing, how far behind do you think Australia is at the moment?
Alex: Yeah, quite a way, but that's attitude. I don't know enough about how things run here in granular detail to make authoritative statements, but I'm going to. And I'm going to say the biggest difference between here and there... is people here just knuckle down and get things done because there's an incentive to get things done because the ecosystem kind of understands that if you pay a bunch of big money for expensive specialists, the worst thing that you can do is start standing in their way. And so there's like a real flywheel to actively strip out dumb management practices and counterproductive organizational structures. And so you end up with super lean structures.
Jim: That's actually really insightful. Is it, you know, the same, is that why they sort of, you know, there's a delay in terms of... I guess, government and regulatory interference as well. Whereas, again, here, we all sit on the sideline until there's government interference, and then we go, oh, well, we can't actually do anything, and it's all a complete waste.
Alex: I think that there's elements of that people are much more aggressively... leaning into or taking part in the regulatory process i also tend to think the government over here you know makes australian bureaucrats look confident competent and well-funded so there's that but like to give you an example of one of the changes coming down the barrel that you probably don't or you probably haven't heard of there's a 10 sales tax being applied to intangible goods i.e sas products in the state of california That means anyone that's selling SaaS in California, the government's going to take a 10% clip. It's expected to raise about $2 billion. So that there is the sort, like these things do happen, but they're just different things. It's a different game, I guess. In Australia, the incentive is not to start something new. The incentive is to sit on the biggest pile of gold you got and not fall off. Whereas over here, it's like, well, if you haven't fallen off a pile of gold before, you're clearly not trying hard enough to grow the pile of gold.
Will: Yeah, yeah. It's fundamentally like the values of the culture, which is Americans are far more reward-driven and Australians are far more risk-adverse. And you just end up with two economies and two technology markets that just reflect that value.
Alex: But it only takes one walk down market street to kind of see the darker side of that system too, right? Yeah, there's a lot of homeless people here, a whole lot of homeless people. And there's that joke about, you know, we'll write HTML and CSS for food. I don't necessarily think that's a joke. There are people in our sort of demographic that I've seen kind of kicking around the streets. And it's like, well, you might have been a developer that was just in the wrong place at the wrong time.
Jim: Yeah.
Alex: There's no safety net here.
Jim: And you really can fall in that sort of hole where what you've been working on for the last couple of years is then effectively useless because something new or something different comes out. I was having a discussion just a couple of weeks ago just with people in that... you know about sort of data bricks and fabric and that oh well everyone is using one or the other and if you've if your experience is in a different data platform then most of the corporates in australia really have no use for you know because you don't fit their mold they can't think outside that mold and so what are they you know what can you do
Alex: Yeah, it's a 1970s model of, I think I've mentioned this before, that model, that hierarchical model is designed from 1970s business schools studying 1960s manufacturing concerns. And we apply it to knowledge work and wonder why it doesn't work. You know, over here, that sort of foolishness causes enough problems, particularly if you're an early stage company, that the way that the company's organized just looks very, very different in practice. Like no one, yeah, literally no one that I can think of on my new team is a specialist in any conventional sense of the word.
Unknown: They're all kind of like,
Alex: yeah, I've done this, I've done this, I've worked here, I did this for a bit. And then I took up, you know, playing guitar and sort of teaching Spanish in rural France. And now I'm back in tech, you know, like there's just a lot more. I hate to use the word diversity because, but like, that's it, you know, like within the, within a single individual, there's more diversity of like thought and experience.
Will: Less people being in a box. Yeah.
Alex: They don't like boxes here.
Will: Yeah. And to Jim's point, like in Australia, in that data space, like you have to put yourself in that box of I'm a data engineer who works on Databricks or you're not going to get a job.
Alex: I will say one thing though, sorry. While you're here, don't publicize yourself as like an AI person. They don't like AI people here at the moment because rents have gone up about 50% in the last six months. And so, yeah, if people are asking, do you do AI? Well, you can play that one how you want, but you've already put your foot in it if they've figured that out.
Will: You'll love this, mate, because you get to finally say, I work in technology and I avoid AI as much as possible.
Unknown: Well, yeah.
Alex: I mean, like, that's the beauty, right? Like, when do I use it?
Unknown: When it's useful. When do I dislike it? Most of the time. Why? Because I dislike everything.
Will: I had one more for you, actually, Alex, which is... So when I went to San Fran, actually, no, sorry, this was a conversation with someone from San Fran came over here to visit about a year ago. And she was telling me about people in San Fran were like kind of on the border of like tools down because we're all going to lose our jobs soon anyway. So kind of just dial it in at work, quiet, quit, and just wait until the inevitable happens, which they were all saying was three to six months away and that hasn't happened yet. So I'm curious what the vibe is at the moment in terms of optimism and doomerism and where you kind of see each.
Alex: I can't speak meaningfully of that at this point in time. I can speak to my immediate environment and that is people aren't like that because they kind of know the tools aren't coming to replace them. A lot of time is spent swearing at the tools and, you know, we're in a kind of more traditional... tech company than like an AI company or a professional services firm however if I was working at Salesforce I would absolutely be bricking it I met an Accenture consultant the other day and I've been meaning to catch up with him and sort of get a bit deeper take but Salesforce they're the guys that I reckon are all feeling the crosshair on their back at the moment because not only is a lot of the work that they do kind of automatable, not all of it, but a decent chunk of it. But the powers that be at that company are very, very vocal about, we intend to replace our workforce with AI.
Jim: He's also been out there saying, oh yeah, you know, we want to switch to be AI first and all this sort of, and they made all the good moves. to you know you can run salesforce headless and all of those things like that's all really smart but when you've got a head when you when it's a very ui user-focused tool and a lot of your developers work on that side of it and you're now promoting a headless version let me tell you about how it's all fine to spend whatever you're spending now on in terms of tokens on your developers because half of them ain't going to be there very sure yeah i think salesforce is a special example because of the nature of what it is right
Alex: It's not winning any prizes for looks or innovation or anything these days. It's the new IBM. It's the new sort of Oracle. And in that sense, you know...
Jim: I think it will be interesting in that Benioff moved faster and first, or I say first, but earlier than the other... generic behemoths you know is that you don't hear as much about hubspot doing the same you don't hear as much of you know like at monday.com doing the same you know and so that's sort of where i start to wonder is there going to be an element of salesforce that survives based on a network effects and b that they did move fast enough you know
Alex: Yeah, I honestly have no idea where that's going to shake out.
Jim: Neither do I. I just think it'll be interesting. Yeah, yeah, yeah. I think it's all going to be interesting. Like, no one, everyone's out there, you know... giving spicy takes on podcasts, but none of them really know. Well, again, that's a boring way to spend your Sunday, isn't it?
Alex: It's Saturday here, mate.
Jim: Yeah, well, again, 17 hours behind. Do you want me to tell you how the future works? I'm 17 hours ahead, mate.
Unknown: All right, I'll take that.
Will: I think the strategy there, Jim, with like a sales force is recognizing that there is a reasonable chance, whatever it is, that a lot of their customers will stop renewing contracts because they want to replace a lot of those workflows so they can build themselves now. Now, if that is to happen, let's say 20%, 30% probability that's like significant, then their costs are going to be way too high for their revenue. So then they go, okay, well, we're going to need to lean up the business just in case of that eventuality. And hey, leaning up the business looks very much like actually replacing all of these human-led workflows with AI workflows anyway. So it's actually kind of like the safest hedge right now for like a big SaaS company would just be to replace a lot of those workflows with AI workflows.
Alex: If the net effect of all of this is Salesforce gets an API layer that's half as powerful as HubSpot's API layer,
Unknown: Net win, I say.
Alex: HubSpot is decent to use. It's not winning any awards.
Unknown: It's fine.
Alex: But Salesforce, as an engineer, is awful to try and work with. And what year is this? It's not 2011 anymore. You don't get prizes for having a web UI.
Jim: I agree wholeheartedly. I think it's all going to be that there's these transitions and these evolutions, is the right way to describe it, in the way in which they all behave and that they have to service their client to provide the value. And I think... agree no matter what seat counts are going to go down because businesses just aren't going to need them or the agents will be able to retain that knowledge. But if they can refocus in that they are the source of truth and they're the data, they're the data layer. And like you say, Alex, Salesforce produces an API that's usable, then net win, everyone's moving forward. And all of those other workers find another role somewhere else because there's still going to be a need for humans to interact with it. So there has to be some sort of UI.
Unknown: It's just not going to be click here, enter information into this field, press... That being said, if you've been a Salesforce integration specialist for the last 15 years, you've made out like a bandit.
Alex: Those guys get paid stupid money to interact with Salesforce.
Jim: But also, you've already got your nest egg. Stop complaining.
Unknown: Yeah, yeah, fair.
Alex: Go learn a real skill.
Jim: You always spoke about going and sitting on a beach. I'm hustling now because I want to go and sit on a beach. Go and sit on a beach. There's more important things in life.
Will: There's an interesting sort of, I think, story there as well this week before we fully get into the big ones. The Cloudflare CEO was talking about they're expecting and because they're starting to see AI traffic to websites, like some websites are receiving sort of a thousand times more traffic than they ever have before. Actually, it wasn't even some websites. I think it was like an average across the whole sort of portfolio of it. And because of that, the potential of that causing a thousand times the cost flowing down to a lot of websites, right? So at the moment they said it's around $3, I think, and 16 cents to serve a website and they sell it for about $12. But if the average website goes from being kind of $3 to serve to $3,000 to serve, how does that affect most businesses? If we talk about, let's say, maybe small and medium businesses in Australia.
Alex: Can I just add another point there? The problem is so bad that Cloudflare was talking about delisting Google to prevent crawlers from being able to find websites. Yeah, I think that's how it worked. Cloudflare was saying we're going to block Google Origins so that Google can't take, like, this is, that would be fairly nuclear if that were to happen, but that's how serious this issue is.
Will: Like Google, I would probably understand like origins coming through from open AI or anthropic, but they're probably used to the level of Google crawling right? Or is Google like massively increased with Gemini and their AI overuse?
Alex: Google doing the indexing, making it possible to find the websites in the first place. Like if you're a big website, you're fine. But if you are a niche website, like I know I read Naked Capitalism all the time and the web admins on that website have been complaining about bot traffic because... you start to Google around to talk on a particular issue or your bot starts to talk around and it starts and going, hey, all these random websites speculatively. They're not hitting for targeted traffic. They're going and hitting speculatively for agent work and then scraping.
Unknown: Good.
Jim: i've even i've even noticed it on you know on the websites like mine and on client websites you know there's all like there's recently been a big spike in traffic from areas you know like where sort of traditional data center hubs you know like singapore central europe like west coast us and when you look at the bounce rate On the traffic from those, the bounce rate is significantly higher than everywhere else because it is clearly just bots pulling data and then reproducing it. I've then started to notice that there's things I've said in articles appearing elsewhere. I just go, why are you using mine? I just find it ridiculous, A, that they're using mine, but B, it's a cheap and easy...
Unknown: way to get information if you really think about it you know compute is cheaper than tokens i mean this is a good opportunity for cloudflare right they've got their ddos and bot protection that you can put in front of your website
Alex: And so to your question, Will, there around costs running through, if you are a Cloudflare customer, not that they're sponsoring me to say this, but there's other services that do the same thing, then you're probably not going to see a 1,000-fold increase on your traffic. You're probably going to see a 1%. or you'll see 1% of that, say 1,000-fold, 1%, 10x. So it runs from 3 to 30, which is still not ideal, but that's a lot better than 3,000, right?
Jim: It then comes down to a question for each business and each Cloudflare customer. Do they put those... preventions in place and exclude themselves from AI search, or do you lean into it and go the other way? You know, it's, and, and again, we're well aware that Cloud Bay isn't sponsoring, but if they are open to sponsoring is that Alex will wear a t-shirt.
Alex: I will, I will wear a t-shirt and I will show you mercilessly.
Jim: Will happily do an ad read.
Will: I think Alex would wear any t-shirt. I've been reading the top of his t-shirt and I'm kind of a bit too scared to ask what's off camera.
Jim: What's the other type of person in the world?
Alex: Well, you can see there's a line here, right? And that's where I've cut along and the rest is just like bare belly.
Jim: It's actually a tank top.
Unknown: Yeah.
Jim: That's great. Let's get into our stories. Well, but Cloud Player Alex will also wear a tank top.
Alex: Tassels also. If you can get me Cloudflare tassels, that's a mental image that no one needs, but I'm going to put there anyway.
Will: It's probably not in the merch shop, but they're going to have to make it for you. All right, guys, let's get into our stories of the last kind of two weeks, actually. So, Astra. OpenAI's latest model, ChachiBT, Astra, people are loving it. The model seems pretty phenomenal. I think before this, we hadn't really seen a really big OpenAI model release in some time. Like we'd sort of talked about... What is it? 5.6? Am I getting that right? Yeah, 5.6. You know, being very good and being useful, but it still sort of felt like the fable mythos family models were still like a jump ahead. And I think Astra is a big kind of launch into that space. We've got Jensen Huang, CEO of NVIDIA, congratulating OpenAI saying AGI has arrived. And that's a pretty big statement, I think. So yeah, guys, have you used Astra? What are we thinking about it?
Jim: I've used it a fair bit in that a little bit out of frustration, which I think we'll talk a bit more later. But I find, again, it is an incredibly capable model. And what I didn't realize, and I didn't realize that I'd developed this fatigue in that there's so much... forcing Fable and Opus to do what you wanted to do. Whereas what I've found in switching over to Astra and Sol is it just goes away, does the work and comes back. I was telling, you know, sort of just anecdotally saying it to Will during the week is that I got a bit of anxiety. Is it because I hadn't checked, you know, I hadn't had anything come back from the agent. in over an hour. And I started to go, Oh, hang on. Usually I have to interact with, or when I'm using the, the anthropic models, I have to go and force them to move forward or, or answer a question for them because they're, they're being overly cautious or, or overly restricting what it is I'm trying to achieve. Whereas Astra was just taken away. Is it, they've got, they'll have to do something in terms of, the tokens they give in the subscription or they subscribe in the subscription because, again, you go through your subscription far quicker. But it's an incredibly capable, and I think that's where Jensen's coming from, is that this is a model that just goes away, handles it, argues with itself, and brings back a result that is in line with what you asked it to do. And to me, that's been the biggest... biggest takeaway is that i define requirements and it comes back with the work and that is what i'm looking for out of a out of an agentic workflow yeah i think it's fairly compared with fable in terms of both
Alex: how it sort of does work and you can just kind of send it off to do its thing and also how quickly it rips through your quotas. I haven't really given it a chance to properly thrash it out yet. Before recording today, I was setting some things up and just playing with it and sort of trying to see where it breaks a little. But the one thing I love that... they shipped with out of the box was a migrate from Claude skill that you just run and it takes everything from Claude and drops you straight in nice and smooth. So even if their harness requires a bit of work or anything like that, the fact that the lift and shift to kind of try it out is pretty low says that they've got some pretty smart product people. And I like these well-designed.
Jim: I particularly love the timing, you know, is that it's, you know, as everyone was getting a little bit frustrated with Opus, you know, being too verbose and it just generally, you know, Anthropic restricting everyone. Open, you know, Sam came out and, you know, he's been quite... public in terms of the interviews he's doing. And he's just been matter-of-factly talking about getting AI to work, doing work, what they want to do. And he's back to likable Sam. You go, I don't care how you feel about him. He's a very, very successful individual. Hearing him over the last couple of weeks as they've released this model, you just go, well, that's why he's successful is that he's able to flick that switch and just become this charming and rational person again.
Alex: I don't know about that. I think they've just elevated the chemicals in his bloodstream temporarily, allegedly.
Jim: Well, whatever is doing it, however they're doing it, they're doing it right at the right time. And that to me is... how you build big companies, how you do everything right, or you get a product into delivering value for the clients. I'm feeling like, again, I don't like that I have to learn new syntax. You both know how much I whinge about that sort of thing. But Codex is... a straightforward enough product. And again, like I say, it takes the requirements. It takes the work, the work that you've assigned to it goes away, brings it back and it's done. And that's what I'm looking for. That's the solution that to me, AI brings. I don't have to go and go, you know, like so much with, with Opus. I come back and it's somehow purely in the way that it summarizes what I, the requirements I give it. before it actually goes and executes, it gets into a point where it's, oh, well, the user has asked for specifically this, and it generalizes into, oh, well, it's got to be something like this. And that's not what I'm asking for. I'm asking for something that is exactly this. And so I think that difference alone is going to make a lot of people take up the migrate from Claude option.
Alex: But Jim, I don't want to talk about practical stuff. I want to talk about benchmarks. I want to talk about this claim of AGI because that has been...
Jim: But that's where I just start to ignore all that bullshit, you know, because it just becomes this noise.
Will: Hang on, Jim. Let me read out the benchmarks and then you can tell me how much you hate them. So OpenAI have claimed 99.9% performance on Arc AGI 3, which is one of the big ones that Frontier Labs have been working towards for some time. So I've completely saturated the benchmark on this one. Now, interestingly, To achieve that score, they're actually using their own provider adapter harness, which allows them for that model between different calls, API calls, to actually preserve its state and its context between it. The Arc Prize Foundation, who I believe actually runs this benchmark, they then ran Astra again later using their harness, which does not allow you to carry state between the different API calls. And Astra only scored 62.7% there. So it's very interesting, like we've talked in previous pods about the value of the harness now and how much of the intelligence of these models is actually starting to be funneled into the harness because these model companies we're finding don't have a lot of moat, especially after a lot of the distillation that we've seen going on. And the moat is moving into the products that they embed alongside the harness, alongside the model, which is their coating harness.
Alex: Thinking back to my days of doing data science and anytime I had a model that out of the box that gave me 99%, even 99%, right?
Will: You put in the bin.
Alex: Well, no, first I'd pat myself on the back and tell myself I was a genius, but then my sense of sanity would reassert itself because I'm not a genius. And I go, I've done something wrong here. I look at this number very suspiciously. I look at it very suspiciously. And I think the fact that they're willing to publish it suggests that Jim's attributing all this good judgment to Mr. Altman. And I don't think letting this number go live... necessarily speaks to good judgment. This is time to send the engineers in and have another one.
Will: But then Jensen Huang, who we've got so much respect for, is saying AGI has arrived based off this benchmark.
Alex: Yeah, because he can send more GPUs. How big is his account receivable at the moment?
Will: But OpenAI have been a challenger for them with their jalapeno chips, right? So OpenAI have been the doghouse of them.
Jim: And only in inference. And how much did NVIDIA invest in OpenAI when all the circular deals were getting done? There's still an element of Jensen going, oh, well, I've got to talk my book and support all these people. Because you think about it also, the week before, he bought Hugging Face. And the week before that, he basically said open models are the way to go.
Will: Yeah, he's like tripling down on open source and then gave a little gift to Sam Altman this week.
Jim: But coming back to timing, you know, and who's handling messaging at OpenAI, like, not that I actually do believe it's Sam, Alex, but... I'm blaming him. Let me tell you...
Alex: He's Mark Zuckerberg's less charismatic lizard twin.
Jim: But let me tell you, Greg Brockman definitely didn't check that 99 number before it went out, you know, and a lot of it... There's still a lot of very, very... intelligent people in their research and science team who didn't properly check that number before it went out because you guys go, you know, you go, oh, if you get a 99, geez, if you get anywhere high 90s first off the bat in a model, you've absolutely done something wrong because it's nearing.
Alex: Yeah, and also the nature of this being their super secret recipe, trust me, bro number. Like I've got the Alex harness. The Alex harness does like 99.999%. I have outperformed OpenAI's new model. Trust me, it's based on my proprietary harness where I invent a number.
Jim: You don't need to test this against your harness, which is what the benchmark is based on.
Alex: No, just use my number and I'll take my multi-billion dollar bonus check if you wouldn't mind.
Will: So I've been doing a little bit of work the last maybe two months on trying to make my cloud code more adhere to the way that I want it to behave based on who I kind of personally am and the way that I do projects. And I want to measure how it improved, right? So I started to kind of like measure the baseline and then start to see how I improved it. And then thought, you know, I should probably be using actually the benchmarks because there's a lot of work that's kind of gone into making those fair and, you know, a good kind of sample across a whole bunch of different problem domains that I would work on. So I started working with the benchmarks and like within a few hours, it became abundantly clear to me how easy it would be to game that system. Like I had, you have this benchmark and then you have this model and your harness and there's this big space in between where you can do whatever the hell you want and no one from the outside is allowed to see what you've done because that's proprietary. But whatever you do inside can radically change your performance scores. And you sit there and you go, well, there's a scientific side of me that says, you know, I only want to put through the improvements to my harness that is... truly going to make my model better because i do want it to be better and then the marketing side of me goes well but hang on if i add in this other little feature here i get an extra 15 bump in you know reducing hallucinations and it won't work for everyone in fact it probably doesn't even really work for me but i've got it i've got a experiment result here that got that result and i can publish that if i want to
Jim: That also brings us back to what Alex was saying about the differences or the real realizations going in that the marketing is right up in your face. There's so much of it. really does seem about being intellectually honest and not caring you know like because like you say is that you want it to work because a you're intellectually honest and b you're actually looking for more value out of this you want a higher return on the tokens you're investing And to then shuffle the numbers or have, like OpenAI did, have the state continue from session to session. you're just not being intellectually honest with yourself. And so none of us really know how it was set up and all that sort of stuff. But the way in which they reported it and the way in which they discussed it, you just sort of go, okay, well, you're gaming the system. Anyone who's worked with a model in their life goes, if I got high nines, high nineties in any scenario, I've got to go back and have a look at something because it just doesn't happen.
Alex: Can I demonstrate that this number is allegedly bullshit? I think I've got a clean little experiment that'll close this case quite enthusiastically.
Unknown: So...
Alex: There's another topic that Open... Another thing that OpenAI have done this week that we'll talk about in a tick, but as part of kind of reading through a particular maths result, I was looking into some claims being made by various parties, and I came across a chart published by OpenAI. OpenAI was saying that this particular model that they were using for this particular problem was... was this super secret internal model that, you know, while Astra had, you know, a performance of like 10% on these internal benchmarks, this had a performance of like, you know, 70% over the top. You know, so Astra is just a big dum-dum compared to this new amazing model that only we have access to. Keep your eyes on this space. Now, the ARK benchmark is meant to be a benchmark of AGI. It's meant to be a benchmark of actual sort of like these systems are intelligent and they're at the point of self-improvement, etc. If Astra is at 99.9% or whatever the quoted figure is on this benchmark and they've got this new super secret one, either their internal benchmarks are wrong or the AGI benchmark is wrong. So, like, there is an inconsistency here. There cannot be that much headroom in a generalized test of intelligence. And the fact that they dropped this in the same week as they've been making a lot of other big claims says that I think what we, you know, cool. Astra, cool product.
Unknown: Keen to see it more.
Alex: I'm officially off anything that their marketing department says. And I think OpenAI has gone into the bucket of do not trust a damn thing they say.
Jim: No, I agree. And well, I think they're both, no, no, but I think they're both in getting into that bucket. You know, is it, I can't, I can't trust anything Anthropic says either.
Alex: Yeah, 100%, especially with their chopping and changing. Sorry, this was a conversation sort of pre-recording. We were talking about the Claude Code harness and how it's just you repeatedly develop new workflows around new features that they push really hard. And then two weeks later, it's yanked out from the product with sort of no warning. Yeah, I think Anthropic's got its own issues with the truth, as they say.
Jim: And but it also just expanding on that a bit, you know, while we're doing it, is it must be something in the way, like it must be an internal battle between the model and its abilities. the doomerisms or the we've got to save the world from what these could possibly be. And the poor product people who are just trying to ship something that everyone loves and wants to keep using, which is what they should be focusing on. But there's just so much, there seems to be so much distraction that there's now this internal fight and I can't use the product.
Alex: There are two wolves inside every AI giant company.
Jim: But it's becoming abundantly clear that is the case. And if I can't, or if I lose two and a half days of work because I have to rejig all my workflows that I've built around your product And I now have had to change everything because of a change you've made to your product that you didn't make explicit to me. I'm off you. You're done. I keep going on about it, but it needs to be predictable, repeatable, and cost-effective. Well, you've just lost me two and a half days.
Alex: They need to make good products and stop with the over-inflated... The overinflated press releases and the silly trial balloons and the speculative declarations that everything's changed and then a week later everything's changed again.
Jim: That's it.
Unknown: Some consistency.
Jim: But it's also you need to do it intelligently. They brought in the classifier for auto mode. But they've done it in a way where it has general rules and then Sonnet decides whether the agent is allowed to execute the tool use. But I don't want it to be probabilistic. I want that to be deterministic. I need to know that the agent can do what it's doing and I've given it permission to do that because that's how it becomes predictable, repeatable, and... and cost effective for me is that the moment you, you impose something on that, you've, you've lost, you've made that, that scenario not real for me and I'm done.
Will: Yeah. The last two months have been interesting watching OpenAI and Anthropic. I think OpenAI have very much still been doing the, the product first approach of your users are the thing that matters most and let's give them the best possible experience. Anthropics seem to be traveling down this pathway of like, we know better than you. And it's starting to come across in their language, in their products, and even in the models themselves now. I think Opus 5 was probably one of the first examples of that where they... it kind of degraded the user experience with how verbose the model was because it performs better on some of the, I suppose, the benchmarks, the internal benchmarks that they're trying to aim towards. I suspect that they know that wherever AGI is or maybe, you know, the two or three internal models ahead that they are, they've realized that there is some level of kind of saturation of the language space that is required in order for you to find the right result. And we see that with like the agent swarms. It's like... one out of a thousand of these will be right. So the best way to get the right answer is to just generate as many tokens as possible and then find the right answer. This is a far worse user experience, but it actually pushes them closer towards what they would define as super intelligence. And it's very clear these two parts are diverging.
Alex: Well, isn't it just a little bit suspicious that they tend to bill on tokens at the same time? Part of me, I hear what you're saying, but it's like, well, bill on outcomes then. It'd still be a shit user experience, but at least it won't hit you in the hip pocket.
Jim: Well, that's what ended up triggering me the most is that it turns out there's scenarios where they're billing me for Sonnet to decide whether my tool can be used or not. I just went, are you kidding me? You've changed this and it's costing me more.
Alex: Speaking of costing a lot and spinning up thousands of agents, the other big OpenAI result this week, another Millennium Prize problem has fallen.
Unknown: Huzzah!
Alex: This is exciting for me. I was wrong as to which direction this is going. Not that I know bugger all about maths, but this is the most likely to fall, and it's kind of fallen in a way that I kind of didn't see coming. So, yeah, this is really cool. However, well, no, let's start with the good news. Will, do you want to run through the press release on this one?
Will: Yeah, yeah, sure. I mean, I love that Alex says he knows bugger all about maths and he would absolutely kind of swamp 99% of people on the planet. I feel like a sixth grader right now. Anyway. So, yeah, OpenAI generated a proof this week for the Navier-Stokes equation, which is in fluid dynamics. It's a 90-year-old unsolved problem. I've read it. I don't understand it. So Alex can give us some explanation if he likes. I read the word vortex and I saw some images and I thought, hey, that's cool. I don't know what that's for. But I suppose one of the interesting things about this was, I suppose, the impact of it. A problem that was this complex that was kind of held in high regard in the mathematics community as something that was very hard to solve and important to solve. And we now have the proofs coming out of our AI labs where researchers and mathematicians are working side by side. And apparently they use thousands of agents using one of their internally unreleased models. to get to you know kind of this result so we just have to firstly just like kind of there's a lot more to this story but i think we just got to park on that first of all and just talk about the incredible implications there is now that we have the ability to have thousands of ai agent models working collaboratively towards solving math problems that have real world impact
Alex: I think there's another, like, this is the big headline one, but there was another similar result that came out of Anthropic just a couple of days ago. Fermat's last theorem was a famous unsolved problem up until about 1994, stood unsolved since the 1600s. And it's just been formalized in this thing called Lean. And what's Lean? It's a programming language for maths proofs. And I think that's when we're talking about these new maths proofs that are coming out. Yeah, agents are kind of enabling this, but the unsung hero are these like... formal proof checkers that allow agents to kind of try ideas and see if they actually solve the problem not just hey does this solve the problem on paper but give some sort of programmatic guarantee that hey this theorem that i've just stated it's actually provably true So, you know, yeah, both this Fermat's Last Theorem and the Navier-Stokes kind of regularity problem. I don't know if it has an official, like, theorem title yet. They're both relying on this sort of underlying architecture of deterministic code that makes sure that the agent isn't just imagining a solution.
Jim: But, yeah, huge... Well, I'm just... really interested is to see if it'll improve or help the f1 teams go faster you know is it because sure surely all that like isn't isn't airflow part of like competition or fluid dynamics isn't that really you know an application of it i think so i can't remember if it's for compressible fluids air is a compressible fluid the detail
Alex: The details of what the problem is basically saying if you've got a big body of water and you move... If you were to move it in a special way, can you accidentally generate infinity somehow? right and that's that's really what the question said we didn't know if it would or not why is this because it's a thing called a dynamic equation oh sorry a differential equation differential equations are a pain in the ass no matter which way you cut them you can only really solve the simplest ones all the other ones you've got to simulate So will this lead to new F1 equations or F1 models? Probably not, but it'll probably add some new if-nan-then-do-something guards in the code that does this. So it might make them marginally more efficient.
Jim: But that's where I like it. But in an area where... a, you know, a microsecond improvement in efficiency can actually, you know, in each turn can affect, you know, the time at the end of the lap. I, yeah, that's, these are the sorts of things that, you know, cause like a lot of the reporting was, oh, you know, oh, well we can make drones fly faster and all this sort of stuff. And, and, and I just go, well, that doesn't excite me. You know, I'd much rather, I'd much rather have it, you know, have it that, that one team applied it better and, and got a marginally improved, improved result. I'm trying to, I'm literally, I'm literally, running through in my mind is that who actually has open AI sponsorship, you know, this, this may like to talk like real talk for a tick.
Alex: this may actually open the door for like the thing with Navy Stokes is this really simple set of equations.
Unknown:
Alex: But we knew nothing about it, right? We know that it shows up in all these places. In the case of things like airflow around an F1 car, it's kind of like a chaotic system. Chaos is going to dominate that. Local perturbations due to like track conditions or whatever, they're going to dominate those equations.
Unknown: So like...
Alex: Does this solve the problem today?
Unknown: No.
Alex: But this simple equation that we just knew nothing about, like does this ever go to infinity, is like a pretty basic question about a set of equations, right? Like you see an exponential function?
Unknown: Yeah, it goes to infinity. Sine wave?
Alex: No, it doesn't go to infinity.
Unknown: Straight line?
Alex: Yeah, it goes to infinity. We didn't even have an answer for that. We now have an answer for that. And that's kind of like the way in which they've studied it and constructed it. And this is kind of the problem with these AI-based proofs. You end up with this massive machine code generated proof that doesn't give you intuition into how you got there. It's just this brute force thing, right?
Jim: and that was also i think that was the interesting part about the reporting they you know they really spoke about how many tokens it took how many agents were involved how much money it cost well last time i checked you set the price of the token so how does that how you know how you then put a monetary value on it you know and anyway, it was just interesting that was the reporting. Oh, we managed to solve this and it just cost this. Well, we have all this money. Let's solve everything then.
Will: It's like incredibly cheap if you think about the implications. I'm also reading here as well just some other use cases that have been identified. The first one is in weather and climate forecasting, so understanding turbulence and extreme fluid behaviour. Help to improve the predictions of storms, jet streams, ocean currents, heat waves and flood risk.
Alex: It won't, though.
Jim: No, no, but I think what you're saying, Alex, is that it won't in the current state. The way I sort of understood it was that once we then break down how the jet stream interacts with this mountain range at this temperature of the day, and we can actually break down all the scenarios, then it actually becomes something of value.
Alex: No, again, it's like the Navier-Stokes equations are an approximation, but they're a very useful one. The advantage of this finding here is it says this is where the approximation breaks down because we don't see little black holes appearing in a cyclone, right? And so we're not getting infinities in the real world. We're getting infinities in the equation. the equations aren't quite right. But the equations themselves are incredible. They're already incredibly successful at predicting things like weather patterns. They're just computationally expensive, one, and driven by chaotic dynamics, two, which means... it's really slow to get your predictions at a really fine grain level because we can't do a closed form. We have to simulate time step forward, which leads to any small deviations in measurement at any point in the process sort of like exponentially compounds as this chaotic process.
Will: Can I say that back so you can tell me if I get you right? So you're saying that... Knowing where it breaks down, which is kind of much further along the axis than most of these things operate, they're going to break down due to chaos before they break down due to limits.
Alex: This doesn't take away from it whatsoever, but the reporting around it is like, oh, we're going to understand how neutron stars work now.
Unknown: No, we're just not.
Alex: The numbers that we have are really good for the things that we use them for today, and most of our limitations are in terms of measurement. The reason we want to understand these equations is because they are so successful and they're so general. This in and of itself isn't going to like... I mean, it's weird that the infinities crop up... Like, again, I thought they would have found in the negative, no, infinities can't come out of these equations. But the fact that they did, and the fact that they did it in a very particular way, says that there's probably very, not a lot of... real-world application for this beyond, okay, now we know the equations are imperfect.
Unknown: Yeah.
Will: I'd really like to know which unsolved problems could have the most material impact.
Unknown: P versus MP.
Alex: P versus MP. If P equals MP... that would materially break everything in awesome and terrifying ways. Cryptography, for example, would suddenly stop working.
Will: Yeah, so it'd be really interesting to see them throw more than $6.5 million of their trillion-dollar valuation at a problem that could actually make a real dent.
Alex: Well, the reason they threw the money at the problem they did is two reasons. One, it was the most likely one to fall. It was the most tractable. The other problems are hard. Like, the Hodge conjecture is like...
Unknown: super interesting.
Alex: The Riemann hypothesis, super interesting. P versus MP, I don't think we'll ever get an answer on P versus MP. That one's hard, right? There's probably not enough... The second reason that they threw this, they chose this particular problem... was some other people had been working on it with a very particular approach. And they happened to be using ChatGPT to kind of help verify parts of their theories. And in the days before OpenAI's big announcement, they may have gone public themselves and said OpenAI had not only... had not only refused to confirm whether or not their research had been nicked, but actually threatened them if they didn't play ball with OpenAI's story. So, yeah, some spicy takes coming out of the lizard world at the moment.
Jim: But that is going to be one of the things moving forward. It's sort of partly why everyone's worried about them all talking about the regulatory frameworks and all this sort of thing. Is it because then they control everything and it just becomes literally the Spanish conquistadors arriving in South America and wonder why they can take over everything.
Will: Yeah, it was interesting, their response to OpenAI. So they received a lot of accusations of ideas being leaked and potentially who has access to what data. And OpenAI's response to this was very much to demonstrate where they had shown really good intentions throughout this whole saga. But the big outstanding question that they just refused to answer is, was the other researchers... data conversations in the training set for the models that they were using. And they just refused to comment on it. And that should be scary for every CEO in the world who's using these models at the moment, because these companies will say that they will not train on your data, but they have not ever commented on the fact that their internal teams still have access to all of your chat to improve their products. And maybe their products counts as their research as well.
Alex: In this particular case, I'd say this is at least partially on the researchers for not reading the terms of their contract, right? There's no way they're using a commercial agreement here. They're going to be using one of the pro agreements, assuming that protects your data. And the reality is it doesn't.
Unknown: It doesn't.
Alex: If you are using any individual license on any of these platforms, you're going into the training data set, son. And that's why you're getting massive discounts on tokens.
Jim: rates etc but and that's that's absolutely the case with opus 5 and fable on bedrock all of your logs even if you have it in a completely sealed vpc all of your logs are still going back to anthropic for 30 days just because they can then see what you're doing with it to see if it's all right with the world. But that's my, or my clients, in my case, my clients' proprietary information is that we had to switch a client completely off any anthropic model on Bedrock because they couldn't violate, because it was being stored on servers in the US, it instantly violates their data policy.
Alex: And this is one of the kind of telltales in this whole saga, right? Like most companies worry about protecting silly little operational details that don't matter, right? You know, like, oh, it's our secret sauce.
Unknown: No, it's not.
Alex: No one cares, right? Most companies tend to think that they're these unique, beautiful little snowflakes. And it's like 80% of the DNA is the same. Maybe 10% is differentiation and the other 10% is stuff that does differentiate you in a negative way. Like there's a solid center of this distribution that just doesn't matter if it goes out the door. But in this particular case, this approach that these researchers were taking, the non-open AI ones, was so particular and so left of field and essentially never been tried before, never really showed up in the literature before. that is the sort of thing you want to protect. That is the sort of thing that OpenAI has now kind of demonstrated. You can't put that stuff into an OpenAI model and expect it not to be used, and worse, used against you. The kind of... The nastiest detail in this whole story wasn't just that they nixed it, because they did say, you know, yeah, we'll put you on as lead author. But they absolutely refused to let an anthropic employee be on the paper. The Anthropic employee did a lot of the hard work, was there throughout the entire journey, used a mix of Anthropic and OpenAI models while kind of doing his work. So he's not a partisan. He's not even a big name at Anthropic Labs. But that was the thing that OpenAI was willing to go to the mattresses on and literally threatened the academic involved.
Unknown: What was his name?
Alex: one moment here buckmaster tristan buckmaster they said oh this will have negative effects for your career and he said i'm an academic you know like well threats for my career i'm a tenured academic
Unknown: What threats do you have?
Jim: That's the definition. That's the definition of a threat to a career.
Alex: It's the definition of job security as well, right?
Jim: But it's also, just as we've been having this discussion, I've sort of gone back to what Will said early on, is that it's really quite a small amount of money and time they've invested in it. You know, like 88 hours, what's that? Three and a half days? Just over three and a half days? is that six, six and a half million dollars. What is that? That's nothing in, you know, in this. And so, and so you start to look at these guys have done, you know, these, you know, the, the non-open AI guys have done all this work, tested all these things in the lead up, then open AI is just thrown three and a half days at it. And then, then come out with the trumpets blaring, you know, it really is like, it really does, you know, come to, okay, what are we actually going to learn? What intuition are we going to learn from this? if not zero.
Will: Yeah, yeah. I think that's, sorry guys, that's all we've got time for this week. So let's ask our most important question of the week is who is the business idiot?
Jim: I've had a tough week, and so it's got to be the product people at Anthropic losing the fight over restricting Claude Code or making the changes to the classifier in Claude Code. I cannot see, unless they're going to release Mythos before the IPO, I cannot see any reasoning. to what they've done because all it's done is just made their product unusable. So that's a business idiot to me.
Alex: Jensen Huang declaring AGI arriving, putting his foot thermal in his mouth and taking a big old bite. I hope that tastes good, Jensen Huang.
Jim: It may sell you some GPUs, but... You think he read the number and just went, oh, well, 99%, that means AGI. He didn't actually look into it.
Alex: Look, there are interesting ways to torture credibility, and this is certainly one of them.
Will: Yeah, mate. Nah, you stole my one for that one. So yeah, I've got to say Jensen Huang. I feel like when we eventually feel like we're on the cusp or actually reaching AGI, it's going to be really funny to talk back and we go, I remember in September of 2026 when Jensen Huang made that sweeping announcement just to talk his book. And it's now been 40 years before we actually got super intelligence. And actually back then it was just LLMs generating tokens.
Alex: Gary Kasparov famously did this when Deep Blue beat him at chess, right?
Unknown: Yeah, that's right.
Will: There you go. So in the timeline of people announcing super intelligence, Jensen Hong's just put himself on that.
Unknown: Now got a seat.
Jim: That's where I come down on the Moonshot podcast with that one, is that the number of times they drop it in every single conversation, I just go, guys, it's becoming clear that you don't use the models because they're in no way super intelligent.
Will: Yeah, they're wrong all the time.
Jim: That's it. Fantastic. Good work, boys. Alrighty. Thank you, guys.
Will: Thanks, listeners, and we'll see you next week.
One email a week. No courses, no funnels.
Sign-up is not wired up yet. Check back soon.