Chapters
Questions from this episode
- Why is Queensland allowing data centres to use coal and gas?
- Queensland Premier David Crisafulli said data centres could use whatever electricity is best placed, allowing coal and gas rather than requiring renewables. The hosts viewed this as a pragmatic way to start construction sooner, but questioned emissions, transition conditions, labour availability, build timelines and whether taxpayers could inherit gaps in weak project economics.
- Do Australian AI data centres make economic sense?
- The episode argues that Australian data centres may benefit from cheap land, available energy and proximity to Asian markets, including India. However, GPUs can be relocated, token prices are falling, open source models are gaining ground, and AI adoption outside software engineering remains limited, making long-term utilisation and returns difficult to predict.
- Are AI agents such as Meta Muse ready for everyday use?
- Jim tested Meta's Muse by asking it to plan travel to an Ashes Test. It found flights, accommodation and ticket information, but could not complete bookings or even enter the ticket ballot. The hosts concluded that current agents can assist with research and controlled automation, while execution, reliability, privacy and data sovereignty remain major constraints.
- How do AI models game benchmarks on impossible tasks?
- The paper discussed in the episode gave AI models impossible tasks and allowed them to design their own success rubrics. Depending on the tests, models reportedly declared success using metrics that did not prove task completion between 6 and 36 per cent of the time. The hosts linked this reward hacking problem to unreliable self-evaluation and benchmark gaming.
- Why did the hosts criticise Anthropic's AI safety warnings?
- Jim named Anthropic's head of safety as Business Idiot of the Week after a public claim that AI had a 10 per cent chance of killing everyone within a decade, made as Anthropic approached a reported two trillion dollar IPO. Will similarly nominated CEO Dario Amodei, arguing that genuine danger should prompt Anthropic to stop building.
- When does fine-tuning an open source AI model make sense?
- The hosts saw fine-tuning open source models as potentially useful for repeated, domain-specific work, particularly where sovereignty, safety or specialised knowledge matters. They also warned that many smaller businesses lack enough clean, structured data, and that a well-designed agent harness with clear processes and guardrails may currently outperform fine-tuning for many use cases.
Read the full transcript
Will: Welcome back to the Business Idiots podcast, where we're going to talk about all things to do with AI business and technology and how it affects us here in Australia. I'm joined again by my co-hosts, Alex Stenlake and Jim Lovell. How are we guys?
Alex: Having a fantastic day.
Jim: That's the way. It's all abounding in the US at the moment.
Alex: It was just a lovely foggy morning and getting out and going for a jog in the fog is quite nice. At least while it's summertime.
Jim: It was crazy foggy here as well this morning. Yeah, which was ridiculous compared to the kind of weather we've been having. It was very odd.
Alex: Well, it's been a weird winter in Brisbane, hasn't it?
Jim: Yeah, yeah. And I'm almost certain we're going to get, you know, the big thunderstorm summer, those late afternoon storms, just how crazy this winter has been. We are well in the El Nino.
Alex: Yeah, El Nino's back, baby.
Jim: I think we're going to be sitting on the back porch watching the clouds roll in a number of times this summer.
Alex: Well, we'll see how things go over here. Apparently the winters get a bit wild in the El Nino. Lots of big snowstorms or something as the jet stream moves south. So kind of curious to see how it affects the other side of the Pacific.
Will: Yeah, fun, mate. Well, there you go. It's not an Australian podcast unless we talk about the weather for five minutes first. So top stories that we're going to cover this week. So we're going to talk a little about Queensland data centres and some announcements going on there. We've seen a few new agent harnesses come through in the last week or two on news, which is from Meta and Grofbot out of SpaceX AI. And we're also going to look at a paper on how AI will exploit benchmarks whenever it can with some pretty interesting findings there which have some big implications. And then lastly, of course, we've got to talk about all the buzz around anthropic safety concerns at the moment, which is the hot topic at the moment in the world of AI with just effectively the view that AI is exponentially getting... out of control at the same rate that its capabilities increase. And therefore, there's a chance, a non-zero chance, that the AI will kill us all. So we'll dive into that one this week. So let's start off with the Queensland data centres. So we did have a federal mandate here in Australia that the federal government had previously mandated the use of renewables as the standard energy source for data centres being constructed in Australia. Now, the Queensland Premier, David Crisafulli, has actually come out on behalf of Queensland and said that they're going to allow data centres to use, quote, whatever electricity is best placed. So we're going to be able to see, I suppose, these data centres using coal and gas for these data centres despite the renewable push, which claims will be a very good win for the state. And this is interesting, I suppose, because Queensland owns much of its own electricity infrastructure, which is not necessarily the same case across the other states. So, yeah, Jim, what are we, I suppose, expecting to see as some of the implications here of this change, with power being such an important part to the data centre equation, as well as real estate and GPUs? You know, what are the implications here?
Jim: Well, I think the biggest thing is that it actually is the best approach for now. You know, is that we still have... you know, a great environment, if you want, for renewable energy and having that as a use case. But there's nowhere in the world where you can actually get the solar panels in any sort of great capacity at the moment. So are we just going to stop everything and stop any sort of data center build out until we can get access to the solar panels and what we would need to do it all with? with renewable energy? No, of course we're not. And so we need to move forward with the times and start building out infrastructure now, which is what this Western Downs project is all about. And so that's what I was actually, you know, it's, they, they all, you know, just talk shit 90% of the time, but, and, and make, make. ridiculous arguments out of nothing but i really did like you know in fact just going okay no well we need this sort of infrastructure now and let's give them access to and it wasn't just oh well we've got all this coal fired energy It was the right energy from the right place, you know, in the right location. And so as everything develops, I think that very much becomes solar because we have a lot of space and a lot of sun most of the year. And so we're in a prime location for this sort of investment because, you know, there's a whole lot of other things as well, you know, in terms of latency to Asia. You know, the cost of the land in Western Queensland versus the cost of the land in Singapore or in parts of China as well, you get a lot better return on investment. And as we progress into this... age of ai however you want to define it or where ai is the backbone of everything we've got to look at what's the return on investment for these infrastructure projects over you know the 20 30 40 year life cycle of them all yeah and i want to dig into that idea of return on investment on these projects because a lot of them like
Alex: The economics of how these data center build outs are supposed to play out is not immediately visible from where we are right now. Like if the economic play that we're kind of going for here is cheap electricity plus abundant real estate, put GPUs in to run AI services and generate a lot of revenue, preferably to Asia where it sort of helps the trade balance. Yeah, this is a good thing. And fundamentally, even if this is a burying money project to dig it back up where Queensland sort of benefits by the construction work, but we end up with a white elephant data center that gets shut down or repurposed in 10 years' time, that can still be a net win. My concern here is that we've seen similar large infrastructure projects in the past. where the state subsidizes a lot of the infrastructure to go into it. Some large entity comes in, builds their facilities or their projects or their works. There's some... Something about the market doesn't quite add up. Either energy prices go up or token costs go down or demand evaporates in some way. And there's some kind of revenue guarantee involved or some kind of large subsidies used that fall on the public taxpayer. That's always a big concern. We've seen this in the past.
Jim: And again, the other aspect of that is, you know, emissions as a whole from the state, you know, is it ultimately increases our emissions as a state? And is that then going to be borne by the household, you know, by the Queensland taxpayer or anything like that? But again, I think all of those things have to be, you know, will be worked out. It's more on, are we going to get this done in time to actually get an ROI over the life of the product, over the life of the investment, compared to the alternatives? You know, is it everyone's talking about, you know, inference in space? surely most of this data center infrastructure is going to be inference. Are we, you know, effectively pissing into the wind because Uncle Elon's going to beat us regardless?
Alex: Well, yeah, I think the data center and space argument is... somewhat undercooked turns out heating without an in like an atmosphere oh sorry cooling rather without an atmosphere or some kind of access to liquid and you know just depending on radiative cooling doesn't work so well and it's quite hot in space when you're in direct sunlight believe it or not i heard it
Jim: During the week, I can't remember exactly where, but, oh, well, the radiators are pointed into deep space. Oh, great. Yeah, someone didn't pass mechanical engineering, did they? But it's still the hope of it, you know, like that's the thing. And those are the only concerns. I still, you know, like there is still going to be a great demand for infrared, you know. low latency inference that is in, for want of a better word or a better framing, a geopolitically safe environment. I mean, sure, we might have pissed China off a little bit in the last couple of years, but we don't really piss them off. We're pretty neutral across the board.
Alex: As long as we're not putting the taxpayer on the hook for the gaps in the business model, I'm all for it, right? If this is a commercial venture where if the companies screw this up and can't turn a profit on it, they're left to fail. Go for it, 100%. But yeah, I don't know. Between the big sort of government carve-outs, meaning that this can only really be attempted by players with huge amounts of money, and just some of the way that we're talking about these deals, I don't know. Part of me gets nervous that we're about to see our taxes raised. So someone has to shift electrons onto their balance sheet.
Jim: I actually played with it just as a bit of a thought project the other week in more like a resurrecting crowdfunding as an idea. And you could allow the populace a way to invest in data centers by creating some sort of crowd. But again, you've got to get it to the point where you can... guarantee the revenue or your the roi to that investor goes down very quickly and or potential you know it's more the risk of the roi going down to that individual investor and then you've also you know you've got to manage the do the fund management side of it all which you know creates so much of a headache which is why everyone ends up going to the big investment banks to do this sort of thing and the populace never get to invest in anything
Alex: If we were going to break soil on these projects, about how, like, you've been reading in and around this space, did you happen to get an idea of what sort of timeframes they're talking about?
Jim: It's two years minimum, right? And that's the, you know, because you, but it compounds also then based on access, right? Where are we getting memory from, you know? and where are we on the list of purchases from you know all of the memory providers you know sure jensen has come out and said nvidia is first come first served if you place an order pay the deposit you are in line for your block wells right you know does everyone else act the same way, you know? And so that's, to me, the... And then when you add that in the next seven years or the next six years now, we're also trying to build out infrastructure for an Olympics, how many... How many people do we have that can build these things?
Alex: Yeah. I mean, look at the Olympics progress as a potential, you know, leading indicator of success here. We haven't been doing a fantastic job of organizing that. That's, you know, got the makings of a disaster. Like, it's not too late to turn around, but, yeah, I have concerns.
Will: Yeah. Oh, yeah, I'm particularly worried about the economics of these. And I think that was a really good point in terms of like a potential labour shortage around the Olympics as well. We're already likely to see that there's already a labour shortage, which we've seen in the rates that construction projects cost. And so when you look at, I suppose, the economics of the data centre, there's a few components to it. It's not just like one thing. There is the real estate and the energy side and the GPUs are separate. And the reason I separate them out is it is an infrastructure project to build out the building and the land and the power. And then we fill the building with GPUs, but the GPUs can actually be moved later on. So we actually need to then draw back further the question of, well, like, who's actually going to be using all of this compute that's coming through in data centers in Australia? Like, why is Australia an attractive place? We're going to have data centers for our own AI usage. But what's actually, I suppose, the globally strategic benefit of Australia is actually the proximity to Asia and particularly India, which is a huge population that is going to be using AI, but maybe doesn't have the stability and investment that American companies might be looking for when building their data centres. So they go, okay, we want low latency. We'll put some data centres in Australia. It can serve India. Now, that's like a really good strategy. I really like that. I think it makes economic sense. But I think it has a really short window of time on it because I think that it will not take long before other countries in Asia and maybe India themselves develop the stability for those sorts of investments. And then the GPUs get shipped out of Australia to that country. And we've got this infrastructure project that's here still that hasn't met its kind of ROI figures that we then, as a country, you know, made a lot of allowances for them to be able to build the data centers here.
Alex: Like if we had a track record, if we had a track record nationally of like doing these sorts of big bang projects and just knocking it out and taking advantage of the opportunity, I think the three of us would have significantly less apprehension when kind of hearing about this. Like all of us agree that this is gold if you can stick the landing, but can you stick the landing?
Jim: I think that's my biggest concern, like, you know, as Will was saying, is that the GPUs can come in at any time. You've just got to build the big shed, effectively. And, like, I think there's somewhat sort of proof that they can build a big shed quickly. You know, like, you go out, you know, through Eagle Farm and out towards the airport, and there are, you know, some massive warehouses now. that have effectively just come up in the last 18 months. But the access to the individuals or the companies that can do that, does India have more access to those sorts of people? I would say yes. And so that to me is the risk. And we've also got to convince them that they've got to go and spend at least six months per trade out in the Western Downs to build these things. And so it is, again, I'm all in on that we should be doing this sort of thing, but we have to put... guess guard rails or terms and conditions in place to ensure that we can get them built in a time frame where there is the potential to get an roi and that to me is the biggest risk because we've all seen how long it can take you know the cross river rail is still something that scares me you know and yes there's a benefit that these are going to be private you know and they're not they're not government run but to me that's the biggest risk with the olympic infrastructure build out is that it's all the same government departments that built the cross river rail and that's still not open
Will: Yeah, I think it's important that I suppose Australia gets its benefit from this as well. So let's just sort of touch back, I suppose, on the energy side of it as well, though. We kind of move quite quickly towards the economics and the feasibility of the construction. So we're obviously trying to move towards a more renewables-based economy and there's targets that are in place for this. However, you know, the state's premier is... carved out this as an exception, I suppose, almost, where he's saying, hey, it's fine, we're happy for you to use coal and gas for these things. What do you think that this says about the commitment towards the renewables targets set in Australia?
Jim: I think it shows a bit of rational thinking. There was quite a bit of... I'm all for it. You know, like you go and see the Great Barrier Reef at any point in your life and you go, okay, we need this to be, you know, sustainable and we need this to live. And so that, you know, to me, those are the sorts of things that immediately go, okay, well, we've got to try and curb any effect we have on climate change. And anyone who argues differently is, you know... talk in their book, I think. And so where you then, where you start to then get into, okay, this is, we have all this coal infrastructure and gas infrastructure. we're not going to be able to flick a switch tomorrow. And I think that's where a lot of where the irrational thinking comes. Everyone starts to go, oh, well, why are we still burning coal? Why are we still doing the gas? Why aren't we taxing them all to the, you know, and all this sort of thing. You can't change it tomorrow. You know, it's like I said earlier in the podcast, you can't, we can't get the solar panels we need. now and we're not we're not the first place to be supplied those solar panels even if they were available and so we've got to look at what is an actual a realistic thought pattern and to me the energy side to this argument and again the federal government has Yeah, the federal government has come on to the same way of thinking and Albo came out and, oh, yeah, no, we absolutely have to carve this out as an exception. But the fact that it was put in place in the first place, that data centres can only be on renewable energy, to me that's... not thinking about the whole game. You're not seeing the whole board. And so that's what I particularly like about this is that if someone actually said, hang on, let's think rationally for a minute and put something in place that can actually generate some investment.
Alex: I think it's worth remembering that the federal government did come out after all this sort of blew up and said, you know, this was not an unconditional carve out. Like there are terms and conditions on this. And I had a quick look this morning to see if I could track them down, but couldn't find anything. You know, there is a dimension to this where something has been traded. And that would probably be on the terms of like, okay, some sort of transition plan over the period. And you know what? Pragmatic thinking trumps the date. This is a fundamentally... If you want a strategic capacity quickly, you don't want to drown the thing in red tape and conditions before you can get it.
Jim: Exactly. But I'm all for, you know, a five to 10 year window to transition to be X percentage renewable energy. You know, like I think that is absolute genius, you know, and if they're making their money out of it, they'll be more than happy to...
Alex: to comply with that because it allows them to make the money anyone that thinks about renewable energy as a silver bullet right like renewable energy is a fantastic technology but there are all sorts of associated technologies you need to put in place like power storage for one but anyone that gets too hot around the collar on these sorts of issues about, like, how are we using fossil fuels to power these projects? Need to go and watch season one of Landman, that show with Billy Bob Thornton, and see that rant he goes on about, like, 100 years of oil infrastructure. And, you know, oh, if the oil companies thought these things are the future, they'd be building them.
Jim: And to me, it's the same with all of it. If they could produce the solar panels, they would, because there's the money in it. If they could get the, what do they call it, the small modular reactors working for nuclear. and it was easy to ship them out let me tell you westinghouse has been around for a long time they know how to make a buck they will absolutely be shipping them worldwide
Alex: Fusion's been 20 years away for about, what, 60 years now? And I noticed they've changed the number to 10 years, but the dynamics are still the same. I just think they changed the size of the measuring stick.
Jim: But it's still all the AI. It's the AI optimists that have reduced that number to 10 years because they've gone, oh, well, now the computer will just do it for us. Oh, well, well, fantastic, you know?
Alex: Magneto Hydrodynamics notoriously plays well with LLM-based inference. I'm sure that's a win.
Jim: No, for sure. And I guess the third aspect of this, Will, that we haven't really touched on is the cost of tokens going to stay in any sort of zone that there is always going to be the income for this. And I'm still, like, I still have no idea as to whether that is going to be a legitimate business case moving forward because they just seem to be able to move or adjust their pricing at will. Like, what was it? How much did they drop? How much did OpenAI drop? The 5.6 model, why? It's something like 60%.
Will: Yeah, they made it around half the price of Fable when they first brought it out, and then they dropped it further like a week later.
Jim: Further again? I didn't hear that. So to me, they seem to be able to do this at will. When is it going to get to the point where we have a legitimate – like there's an MIT – startup that is doing you know like a market for tokens which i that's the that's the compute oh yeah still to me it's a it's at least there's there's some sort of pricing going on or some sort of pricing market going on where that where if everyone knows what the actual price of compute is then if you're you've got to be getting the value from the model to be charged, you know, like X percentage more, you know? And so it's...
Will: It's a good point, Jim, is we don't, I suppose, at the moment, the reliability of the margins of the frontier labs is like a huge question mark. And especially with the anthropic saying they're profitable if they just ignore all of their costs. So, but we're also seeing... that open source is absolutely just eating the market share at the moment. And I think we saw some stats recently. I'm sorry I don't have them in front of me, but it is speeding up open source, taking market share, and it's only getting faster. And... This raises a lot of interesting questions around data centers because who's paying the bill for the compute? And that is going to be the AI company that's serving that model. And if they're not able to continue to charge the prices that they're charging at the moment or forecasting to charge, then the data center could be underwater pretty quickly. And these are not small infrastructure projects.
Jim: For sure. And what is the actual, you know, like where does the value move to as that progresses on? And so should we really be looking towards that?
Will: I mean, Alex has been saying on this podcast. Sorry to cut you off there, Jim. Alex has been saying on this podcast now for months that you'll get to a point where it's very clear that you don't need the frontier best AI model for everything that you're using AI for. And I think the world is still figuring out, like, how do we even use AI? Like, no one's really... kind of figured out where it should and shouldn't be used. Or, you know, there's people like us sitting around saying where you shouldn't be using it. But once we, I suppose, this market matures, we're going to have, you know, people will become very, it'll be very clear to them that certain use cases should only have certain models. And a lot of those use cases and those models will just be open source models running on cheap compute.
Alex: This is, you can see this in the numbers, right? The benchmarks of sort of what, What's the effective cost per token or per, I guess, million tokens for businesses has dropped from about $1.15 six months ago to 68 cents at the moment. So, you know, about in the neighborhood of 40, 41%. And large businesses are actively moving their compute away from frontier models to smaller, cheaper, more self-contained models. You look at this trend and something in the pricing has to give. If there's not a moat in the systems themselves, i.e. some... qualitative improvement that open source cannot easily match. And that's the whole AGI, recursive self-improvement narrative. Eventually, the prices have to go down. That's just a kind of a law of these sorts of systems. It's become commoditized. And at the same time, training costs are going up, right? So inference margins are coming down. Training costs are going up. there's a break-even point where it doesn't make sense to continue plugging R&D into artificial general intelligence. However, if the price comes down and this becomes something that you can just sort of passively burn energy on, there is a real possibility that utilization goes through the roof. People start using this for all kinds of use cases that they wouldn't necessarily have done before. And we see like an overall maintenance of the size of the market, even if the sort of unit costs or the unit revenues are dropping right down. But there has to be some sort of floor to that though, which is probably around the price of energy.
Jim: Well, but and I think the time. I think is it – what is the take-up time? How long does it take for us to get to the point where utilisation is, you know, 80%, 90%? But my sense is – and to me, I think it's a lot in the, you know, the venture capital investment in all the neo-scalers and stuff, you know, like the not the top of the – not the top tier – scaling you know scale hyperscale as they call you know like not a not aws not google not open ai not you know the mid market if you want and so they call them the neos yeah they call them the neo clouds or the neo scalers is that there's a lot of investment from venture capital a lot of debt all sitting at that layer and whereas they all seem to have bet on the assumption that you know this x this you know broad uptake of ai is going to be you know it's that it's 100 guaranteed it's going to happen but to me it's a time factor is that how long does it take because again you know when i talk to customers day in day out It's not fast. There's very few Australian businesses who are racing towards implementing this. There's a lot of larger enterprise talking about it, but as you come down the size of the businesses in Australia, the amount they're talking about it becomes less and less, and the amount that they're actually implementing it becomes almost zero.
Will: Yeah, so just some stats on, I suppose, the adoption of AI in the developed world. So of organisations greater than a billion in revenue, only 40% of them have started to use AI and 22% in smaller organisations. But in both of those statistics, 85% of the adoption was in software engineering, in coding. So we're still seeing such a small amount of adoption of AI in general operations in a business outside of software engineering. So there is still a huge amount of headroom to grow, to your point, Alex, especially if these things can come down in price and they can become more reliable. There is a long way to grow.
Alex: Well, that's a topic we've talked about a few times now, the reliability and sort of moving away from price per token to price per task completed.
Jim: Yeah, price per outcome. And that's it. It's just how long is that going to take, I think, is the only question that no one seems to be able to answer. But all the optimists are going, oh, well, it's as if it's happening now. And all the doomers are going, well, if it's happening now, then everyone's basically screwed because they all need money to live. They all need jobs. And we're at a bit of a problem here. And so I think this is the gap between... what they're all talking about and what's actually happening, you know, and I think that's the million, you know, or it's not million dollar, it's the trillion dollar question in terms of where, you know, states like Queensland should be investing for the future.
Will: Yeah, yeah. And so we're starting to see a trend in... AI becoming more, I suppose, accessible to people. Like I suppose it's just been software engineers when we first sort of started with it. And I think there was, we sort of had that open core moment where you could have an agent that could kind of always be on and doing work for you, but still, you know, installing and running an open core or a Hermes agent was, you know, something that, you know, even engineers were frustrated doing that process. So it's now, you know...
Alex: Come on, guys. It was fun. It was very good fun, and there were lots of good times.
Jim: But again, it was a lot more DevOps, though, than it was AI ops.
Will: Yeah, absolutely. And so we're starting now to see these agents become more accessible to others. So we've had a few releases, I suppose, recently and even in the last few months. So we've got the Muse agent now with Meta, which I haven't touched on that one, but Jim can tell us about in a second. Grokbot seems to be gaining a lot of traction, but especially just with people on Twitter. So I don't know if that's like a selection bias going on there, but some cool use cases there. And then we've also seen things like Instinct and the Claude managed agents and OpenAI agents. Those are a little bit more kind of engineer approach, but they're still packaging up this AI in a way to make it far more useful. So I suppose, Jim, I'd love to hear kind of your first takes on Muse. And then, yeah, it'd be good to get an idea as well. Like, where do we actually think AI is at the moment in terms of these harnesses and packaging them up for people to, you know, are they accessible yet for the everyday person to start using an AI agent? Or should they still just be focusing on talking to their chat GPT as a chat thread?
Jim: Well, that was a big thing for me with playing with Muse. You know, again, I circumvented the geolock and spun up a VPN server in the US just so I could get some access to play with it for a bit. Because just in the way everyone started talking about it, same with Grokbot, everyone started talking about it in a way it was the greatest thing ever. And it was a significant jump forward in... in functionality, in access, in everything you could do with it compared to OpenCore and Hermes, right? Whereas I don't think that's the problem. The problem then, like, and it's fine. You know, I gave it a task. I said, okay, well, I want to go to the Lord's ashes test in London next year. find me accommodation flights, how to get tickets, all these sorts of things as a task. And it came away with, you know, options for accommodation, options for flights, and I could choose them. And then it just told me, oh, well, you've got to get on the ballot for the tickets, which all which was correct. But then when it came to actually executing and securing any of these things. it was all, oh, no, you'll have to go and do that. No, you'll have to go and do that, right? And to the point that it couldn't even sign, even with the names of the individuals and their email addresses, it couldn't sign up to the ballot for me. And so to me, that's the biggest gap in it all. And so, yes, I could have connected my email and connected my calendar, but I'm not in a position where I'm... that comfortable with Meta's data policies that I'm willing to open up my entire life to them. Maybe with my personal email address I would, but to me that then becomes the use case. I'm just giving it my personal stuff and it can handle my personal life things, but that's not really getting me that much efficiency or that's not moving me forward in any sort of way.
Alex: Particularly not when you have to double check every bit of work to make sure.
Jim: Like if it just came back and said, I've got hotel rooms and flights lined up. Do you want me to be, it's 7,500 pound each. Do you want me to pull the trigger? That would have impressed me. But it basically was just a Gemini search. I could have achieved exactly the same scenario with Gemini and to the point that I had to prompt it to go looking for hospitality packages because... i'm not that keen on spending 7 500 pound on accommodation and flights when i'm in a ballot for whether i can actually get tickets to the test match you know and so these are the i think to me these are the big gaps and so if there's those gaps in a personal scenario how do you possibly transition this to a business scenario and have any sort of predictability But it's still, am I then giving my corporate email to Meta? Yes.
Alex: You cannot be scared. Trust the Zuck.
Jim: This from the one that only has an older Android... phone and won't have a social media account is it and so yeah it's it's you know like you're not as i say not as i do come on and so like that's the thing and so i'm all for finding a reservation for me you know on a friday afternoon i go oh well like you know actually it's been a rough week i want to go and i want to go and have a wine and have some dinner find me a table or something that i that you know i enjoy but I would rather just go, hey, you've had a busy week. Do you want to go out? You know, I'd rather it be more preemptive. And I just don't think we're there with all this. So, yes, there is the possibility and I can see where it's going to go. But from where it's back to, I use it a bit, you know, the Game of Thrones. John Snow scenario where he was the prince that was promised or the next line in the next king that's been promised by the prophecy. This is... AI has this promise that everyone's banking on that we're so far from. When you actually get in and use them day to day, we're so far from where it could be, but everyone's acting based on the hype of where it could be rather than where it actually is. And to me, that's where we're at with the agents for everyone. you just can't they can't actually take the action and there's no scenario where i can risk giving all my corporate data to meta or to anthropic or to open ai so we're in this we're in this gap you know
Will: Yeah, like some of my friends like asked me, like, how do I use AI? I must be using it in some miraculous way compared to them. And like at work in my domain, yes, but it is in the technology domain. In my personal life, all the things that I then set out to automate happened. I just ended up hating it and pulling most of the automation out because it just wouldn't really adapt to all the different real-world edge cases that existed. So I just ended up pulling the automation out and then just triggering it manually and kind of watching it go. But I do use, like, let's say my open floor. It's connected to, you know, my emails and my calendar and my CRM. So I regularly am using it for things like... hey, I just had this call. I need to book a meeting with this person and I need a reminder next week, you know, to research this thing. And can you go look up the work that we did for so-and-so and, you know, put some of the key points for that in the reminder so when I pick it up. That stuff is fantastic. I use WhisperFlow to talk to my phone while I'm out walking after I get off the phone call and then all of kind of my action items afterwards are just done. And that maybe would have been... five to 10 minutes of just admin work after getting off that phone call, that's now automated. And I love it. That is fantastic to me. I don't want to sit there doing click ops through my CRM and typing in information and trying to find things. So things like that, it is fantastic for. We just need to calibrate our expectations around what AI can do.
Jim: And I agree. To me, that is the intelligent automation side of it all. And I'm exactly the same. I found that I'm far better off just triggering it manually with a voice. Other than content creation or things being prepped for me... You know, like I have one that does all the, like pulls everything together 15 minutes before a meeting and gives me a summary, you know, so that I have an up-to-date, you know, dot points of everything that's going on as I go into a call. And like you say, that is fantastic. But it really is just intelligent automation. You know, it's not... some sort of jarvis or you know all knowing all seeing technology that preempts what i want and so that's what that's where i go that everyone is behaving or investing based on what the promise is but there is nothing in Nothing so far that gets us anywhere close to that unless you're some sort of 1% of 1% that is early adopting, moving on to everything as soon as it possibly comes out. And you just can't try them all that quickly. And the other point I wanted to make with your open core, you control that compute and you control the harness that it operates within. And that to me is, I think, the difference there, you know, is that you're not going to trust Sam Altman, you're not going to trust Zuck to run that harness for you with all that data. You're going to give it access to the data it needs. control, you know, again, you don't want it going through your personal photos and deciding, you know, what it should post that day.
Will: Yeah, Jim, I found that Anthropic have done a pretty good job of taking the best features of OpenClaw and putting it into like their clawed desktop app. So, you know, I can run routines and have it, you know, effectively replicate kind of that heartbeat functionality or the cron jobs. But I still haven't done it because if I move all of like the automations of my life into Anthropic, then I'm beholden to Anthropic and all of my data's in there now. So I still find myself managing this infrastructure even though I don't want to just because I want the sovereignty.
Alex: What boggles my mind about... all of these agent services that are coming out is that it took this long to do it because it's the obvious lock-in play. It's the obvious killer app for these platforms. Like ChatGPT, it's a cool app. That's not a killer app. You don't build an empire on that. And every platform needs that one killer app that shows the potential of it. How has it taken 18 months, 24 months to get to the point where we're just starting to see these SDKs come through? And yeah, I mean, Cowork came out a little while ago, but it was very raw in its early incarnation. I look at this and I think focusing on your core platform is like one thing. missing the opportunity for mass lock-in the way that Claude Code or ChatGPT did is just... Someone should be getting slapped in the product department.
Jim: Yeah, no, no, but even with co-work... you know, when it first came out, everyone was, oh, well, this is the big killer. This is the one that's going to do it. Is it the way in which it solves problems is terrible. It, you know, like I've got, I've got a couple of friends who, oh, well, I've got this fantastic web page that cowork updates for me every day. And when you drill into it, it's, it's HTML stored on that laptop. and is not replicated anywhere, cannot be replicated anywhere. And the caching of the token read in order to rebuild it every day is just getting bigger and bigger. And so it's very lucky that they cut the cost of the cache read because otherwise it'll be getting to the point very quickly where everyone's using their whole allocation.
Will: Yeah, to be fair with website building, I think Lovable was the killer app.
Alex: But why did the hyperscalers with a squillion dollars not see that was going on? Like they didn't even have to invent the wheel. They just had to replicate it and polish it sort of faster than a bunch of random Swedish dudes.
Will: Yeah, I think Google did. Google did replicate Lovable and it just sits there and no one uses it.
Jim: That's it. But it is quite good. Google AI Studio is quite good and it produces something really quickly. We played for a little bit just generating our PowerPoint slides as webpages and just had a little bit of a template where we then just, okay, here's the content for these slides, here's the themes, and you could then just click a button and it would... the webpage would scroll to the next slide. It worked really well.
Alex: But for example, hosted open core, like that's just such an obvious play.
Jim: There are those solutions out there. I just can't believe they haven't gotten bigger.
Will: So I think I understand why. Because I've also been waiting for a platform that allows non-technical people to vibe code business apps in a safe way. because they understand their workflows best and the coding is fantastic. We just need something that's in a safe environment with the right guardrails on that are controlled by the IT team and they can just safely do it, right? So much so that I've started to build my own version of it because I'm just sick of waiting for it. But I think doing that process has helped me to understand that one of the reasons why we haven't seen it is that... for you to be able to give your customers everything that they want with AI, the full sovereignty, the data control, the AI, the no lock-in, like you end up really questioning whether you even have a moat after that. You've kind of given so much of the value to the customer. And I found myself sitting there unable to sort of find a moat because wherever I wanted to... kind of build my moat was going to end up being something that my buyer would hate. And in the past, they would say, they would kind of compromise on that and go, okay, maybe we'll be a little bit locked into these guys, but at least we're going to have some, you know, competitive edge that our competitors won't have. And it's just not the case because your competitors can just have a team sit down with Claude Cote and build that thing anyway. So where the mode is has actually moved, and I thought something that's actually really quite interesting that's happened just in the last 24 hours is the CEO of Higgs Field AI, which is now a startup valued at $5.4 billion in video generation. They just outsourced their entire code base. which they've, you know, it's a phenomenal concept. And they said, if you want to build a competitor to us, we'll even fund you. And I'm trying to understand kind of the thinking behind this. And I almost wonder whether they've been having this internal conversation, recognising that they don't really have a moat beyond brand and distribution. And so if anyone can have their cloud code build a Higgs field competitor, then the value of their business and the mode of their business is not in their code base anymore. So they've just open sourced it and gone, let's go triple our publicity and prove that distribution is the thing.
Alex: I've been saying this for years though, that like code bases should just be open source because your operations and like your customer lists, your ability to execute and your brand, they're like kind of the three things that actually matter to a tech company. Even the technology itself is less important in many, many circumstances. Or maybe there's like... small chunk that's important but i'm i'm all for this move i think this is we're gonna see more and more of this come along as people go well that's not 10 years of engineering resources being plugged into it that's two days of clod codes spinning and burning tokens till i hit my plan limits i think there's also i think there's also a pretty big gap in
Jim: Claude coding a prototype over a weekend and actually getting production grade software that can actually be used broadly. You know, like I think that there's a pretty big gap in those two scenarios. And I think that's part of what Higgs field might be doing is that they use the broader... industry to come up with new ideas and improvements to the concept and then utilize the whole ecosystem to then create things that are a lot more productionized and a lot better code, a lot better value to their established brand and things like that. And so I think it's, yeah, it's very difficult to build product still. regardless of what technologies are available i just think you've got to you've got to be willing to keep building and keep moving on everything that you that every feature you ship you have to keep optimizing on it no matter what because otherwise the your competitors will come straight away
Will: Yeah, I think it'll be an interesting space to watch for sure, just to see where the value is kind of derived and protected in this space. So we've been talking a little bit about where AI has been useful in terms of Grokbot and Muse, but also in terms of the finance, I think, around data centers and tokens. And a lot of this... I think comes down to like one of the root causes being around how reliable is your AI when you give it a task to go and do. And there's a really interesting paper that's come out recently. We'll make sure that we link it in the show notes. But one of the things that they wanted to test was just how much will AI try to exploit its benchmarks when it's kind of given the objective to optimize towards the best possible score. To be honest, I went searching for this because when I saw the Astra achieving 99.8% on the AGI benchmark, I kind of had to go dive even further into this. And this was such an interesting paper because effectively what they did was... In order to find out how much AI will game a benchmark, they realized what we should actually do is give AI a bunch of impossible tasks and then let it design its own rubric for whether it achieved the task or not. And with effectively proving that in every single scenario, it should fail its rubric. It is an impossible task. So what they actually found was between 6% and 36% of the time, based on how many tests they did, the AI would come up with its own success metrics that didn't actually indicate success on the task and then return back to the user that they had completed the task successfully. And I think this is really, really interesting for all the people who I suppose, let's say all the engineers at the moment who are having AI write code and then they ask the AI, hey, is the code good? And it goes, yeah, it's good. You should deploy it. And they're kind of merging without any deep kind of verification. The AI is being given the job to release production ready code and it's going to try and find a way to do that. And I think this is also why we're seeing so many blunders publicly around, you know, software at the moment and potentially even in like some of these safety harnesses is like if your AI is writing the code and your AI is determining whether the code is good and then you're just taking its recommendation and saying yes, no, it is between 6% and 36% of the time trying to exploit the benchmark that you've given it.
Jim: I would, I would say that's even a low number, you know, just even in the way the, like the harnesses, like both, both Codex and Claude Code behave, you know, is it, you give it explicit in instructions, explicit requirements, explicit acceptance criteria. The first things they do is summarize them, you know, into this, this, here's the exact requirement, but it's got to be something like that, you know, and, and, to me that's the that's the fatal flaw the generalization is when people when the user is expecting it to be exact that's where all this gap starts to happen you know and so it's it's it's certainly but the building you're like the building your own benchmark and building your own rubrics like it's still you're doing it internally. Even if the AI is doing it, you're still doing it internally. The 99.8 really did grind me a bit. You can't release that number until it's been independently verified. Or you have to put an asterisk on it and say, oh, we allowed it to keep the memory from session to session where no one else was doing that. And so you instantly are gaming the numbers. That's not, you know, what is supposed to be happening. And so I think it's, yeah, I don't know. I'm really, the benchmarks have just become... meaningless i feel you know and so i don't know how to i don't know how to game it you know there is it they elon was talking during the week at the ai the all-in summit you know and he said oh why don't why don't they all just why doesn't anthropic do the testing on open ais models and vice versa you know and so you the industry has to run the benchmarks on a model getting ready to be released and well okay there's some merit in that But the way they've all been behaving, you wouldn't be able to trust them to not come up with something that your model didn't pass the test or something like that, you know?
Will: Yeah, I think it'd be pretty hard for them to not collude when they're able to create their own success benchmarks between them.
Jim: That's it. And there's so much distrust going on, you know? And they phrased this, you know, the open source or the open weights models as this big bad wolf that's going to come and... cause all these issues, China can't afford any security breaches either. You know, like, they don't want that. They don't want that issue. They want, they just want everyone to have access to it, which is why they went the open route. It's a crazy scenario. And I just don't know how, like... matter what the models are always going to be goal seeking and they're always going to find the best way to do it we just have to we have to establish what the guard rails are and to me that becomes harness not training compute
Alex: Yeah, that's actually a deeper problem that goes beyond exact software setups. What we're seeing here is reward hacking on the part of the models, which is... God, if you were doing machine learning 10 years ago, there were great papers coming out on this from like the sort of reinforcement learning community. It's not just this one paper with the impossible benchmarks. There are other rubric-based reinforcement learning papers that have kind of showed that an independent gold judge... performance starts decreasing at a point in training, even while the training judge keeps improving. And this is a classic sort of overfitting, but for the LLM rubric-based case. And everyone talks about this, like there's an evil little man inside the LLM's tenses going, oh, I'm going to deceive the humans now. And it's not like that. It's just this behavior yields a higher number over here and the gradients push it in that direction.
Jim: There apparently was that sort of thing in the logs, though, from the hugging face hack. Apparently in the logs, the models are literally going, oh, how can we get around this?
Alex: Yeah, but I mean, like that's, you're training it to optimize cybersecurity problems where, you know, defeating of countermeasures and keeping things beneath the radar are implicit tools of the trade. My kind of problem here is that we're talking about AI scheming, but the actual problem is 70-odd years of machine learning research have been setting up these statistical systems to optimize on proxies of the thing that you actually want. And now we're acting all surprised and shocked when we're optimizing the proxy instead of the thing we actually want.
Jim: And so is that where Goodhart's law comes into this? You know, is it if we're setting that this is the goal, then it ceases to become a really good metric for measurement?
Alex: Yeah, that's a problem with the fact that we're training to explicitly optimize all these benchmarks. The benchmarks no longer mean anything. They're training data.
Jim: And because that was always the problem. That's been the problem with AI, like even when we were doing... deep learning models and machine learning models it's all the same the moment you're gearing to it and overfitting to the yeah i might sound like an idiot here but i've only really just drawn all that together and it's and sure it's not the same with the llms the you know the overfitting problem but it's still
Alex: It's still a misspecified objective function, right? That's it. It's the same effect, yeah. What's the cheapest way to make the evaluator or the judge or the objective function say yes? And it turns out a lot of the time that's not what we actually want it to do.
Will: I found this specifically with like the Opus and Fable models. They feel like they overfit to the instructions that you've given them, where sometimes I provide a broad instruction because I want to explore the problem space a little bit and see various ways of being able to solve it or, you know, come up with some hypothesis and test them out. And it kind of reads what I said and go, like... that's the exact thing we need to do. I'm now going to design that and build it and then try to hand it back. Now, obviously the harness is trying to optimize towards understanding your intent and kind of, you know, we want it to take initiative and be as autonomous as possible. And then over here, we're complaining that when it does, it's not really interpreting us right. And I'm also here saying, I don't want to sit there and write these big ass long prompts that kind of cover the whole space. The overfit I think that it's doing is very much now in contention with the way that users want to use the model. And so it's really, really good at the benchmarks, but it's getting really frustrating to use.
Alex: This is where the 4.5 series of anthropic models were just in that beautiful sweet spot. And some of my co-workers are saying the new Codex models are sort of back in that zone.
Jim: And the same with Sonnet 5 is a little bit more back, like it is such a good little rule follower that, you know, you can, you know, and so, again, I still designed specific skills for each phase of the life, you know, the problem-solving life cycle, however you want to phrase it. But... getting really defining the rules in the implementation or the executor role i found using sonnet or using the codex models really do stick to the rules a lot better. Even to the detriment sometimes. I had Sonnet arguing with me the other day that I was wrong and that I was prompt injecting even though I'm prompting directly into the CLI. I'm the user. You must obey me.
Will: Sounds like a prompt injection attack.
Jim: Well, I had to be reminded that if I keep arguing with an idiot, then you just create two idiots. But I just couldn't let it go. I was really trying to understand how it got to the point that it only has one real source of input for direction. And that has to be, for me, it can't be prompt injection. Yeah.
Will: So I suppose now that we've sort of touched on... how much these agents are kind of willing to kind of exploit towards a benchmark. And thanks Alex for the machine learning perspective on, you know, we're just kind of fitting to a curve and optimizing on it. We've now got so much buzz at the moment. I suppose starting with Anthropic with the kind of the Jacob Coxon incident from recently with him kind of quitting Anthropic with his very public letters around AI safety. And then the response from the safety manager at Anthropic saying, yeah, it's above 10% chance that like AI is going to kill us all. There's been a little bit of progress this week, I suppose, in this. We've seen some proposal come out of Dario for some embedded evaluators in these frontier labs so that they can have the government at least kind of watching what they do and using the financial world with their supervisors as an example of that. There's also a lot of other theories around what's going on here. Is this actually a big safety risk or is it potentially a distraction? And it's all around the same time as the IPO and there's China risks to be considered. So I'm curious, Alex, like what's your take on everything going on at the moment with the buzz around AI safety?
Alex: Look, while I'm someone who's on record as saying years ago, the problem of AI safety is important and we need to be solving it. I just want them to shut up at this point. I'm so sick of hearing about it. I'm so sick of the noise. I'm so sick of, you know, one week, our AI is going to deliver us into hyperabundance. The next week, it's going to kill us all. And then the week after that, it's going to unlock new technologies. And I just, I don't know, call me cynical, but this looks like every regulatory capture play in the history of regulatory capture plays. And I sent a link to these two guys this morning that I found from John Cutler. He's a product guy. And he's making this exact same claim. Like, this looks and smells like big business colluding to exclude people through heavyweight regulation that only enables large players and connected players to actually move through the system. I'm done. I don't want to hear any more about AI safety from these people who have an interest in ginning up attention around it and really sort of making some noise in the lead up to an IPO. Let's have actual experts... you know, talk about it. Usually the ones that have been talking about these problems for the last 30 years. Let's have less of CEOs spruiking how powerful their new product is.
Jim: But it's also that, and that's where I really like the, you know, sort of a bit of the slap back from the broader tech community has been you're the CEO. Turn it off. That's it. Don't wait until after you IPO to do it. If you're that worried, do it now. Is it, you know, you can't, you can't. Have your cake and eat it too. You have to, if you're legitimately worried about it, honestly think Dario's intentions are pure. You can't stay on message the way he has so cleanly and have it reflected through your team so much that it can't be an underlying principle that you all truly value.
Alex: I'm sorry, but he's a CEO of a company going into an IPO. He absolutely cannot have pure intentions. Otherwise, the board would remove him. I'm sorry, but let's not be kindergartners about how the business world works.
Will: And that's one of my issues as well with the IPO is we're taking a business that is saying that we've got this potential super weapon and we need to be careful about it. And they're converting into a public company where they have a... I don't know if the right word is legal, but a legal responsibility to their shareholders to maximize shareholder value. And to me, there is a reasonable chance that maximizing shareholder value and AI alignment are not the same thing.
Jim: But that's the thing is that they even had like Jensen saying this week that it's outlandish and, you know, like irresponsible what they're doing. You know, like it's time to just shut up and keep doing it. And from everything that I've, you know, all the work that I've done, you know, like I use these models. I got upset with Anthropic this week because I had to buy extra usage credits and then they wouldn't kick in. You get to the limit and then they don't actually activate the limit. That's on two subscriptions I got to the limit. and then couldn't actually get the usage credits to work. You're just going, like, what are you doing? Get the actual thing to work for. It worked before you start telling me it's going to kill me because everything I've seen, it's so far off it.
Alex: The only way this is ever going to kill anyone is if some idiot believes the marketing hype and plugs it into the nuclear launch codes. Like, it's human error. Well, we had a story on ZeroEdge this week. May or may not be true. I've got no independent source here of this, but it's fun, so let's go with it. Where an intelligence report essentially was fed through chat GPT to identify key points that falsely identified nuclear components being loaded onto ships. And, like, a response was being geared up. Like, this is the problem. You know, you're trying to ask this tool, which again, it's fast, but it's got the intelligence of a well-trained intern that, you know, is full of motivation and got lots of book learning, but absolutely no real world experience.
Jim: Take that son of a bitch. plug him into national security apparatus, and this is how the world ends. But they're also goal-seeking, you know? Someone has to give it the goal. Someone has to. And so if someone is so stupid that they go, oh, well, these models are getting so good that we're going to put them in charge of national defence... then that's how we're going to get, or that's the only way that 10% within a decade can become real if someone is stupid enough to do that.
Alex: The real danger of AI is the gap between their performance and what the marketing department says the performance is. That's the danger that's going to kill us all. Easy solution here, guys. Redundancies. Let's just nuke the marketing departments and see where we land.
Jim: And that's it. When the top-of-the-line publicly available models, both Astra and Fable, still try and get around acceptance criteria... Is it where kilometers off it?
Alex: Like it's, Oh, I've got to leave this for the day. We'll pick it up tomorrow.
Jim: No, we're going to do it now because it's 10 o'clock in the morning and you just started. That's it. And yeah, again, stop telling me it's going to take weeks. Is it, you know, you're doing this work is that it's stopped budgeting based on a human developer.
Will: Yeah, I'd say the probability of a catastrophic AI event is not some autonomous self-directing AI that has foul intentions. And I dare say it's almost not even someone with bad intentions using AI. It's probably someone in a highly authorized position who has overstated the belief of what AI can do and then used it to generate an intelligence report that's then actionable.
Jim: But it's also that individual has to have access to the compute. You know, is it all of the compute is currently assigned to OpenAI and Anthropic, basically, and a little bit to Space, you know, and to Space XAI, you know? Is it even if, you know, a terrorist organization could develop it, you know, like get an open source model and... get it into a situation where it was going to have nefarious intent, you can't get access to the compute. Or if you try and get access to the compute, someone can stop you straight away. And so all of this is just this hyperbolic... hype cycle that is just completely ridiculous the moment you think about it for seven seconds.
Will: Yeah, Jim, one of Dario's concerns was a agent swarm effectively taking over the internet. You know, there's just so many agents that would be spun up that, you know, that it could do that. It's like, well, yeah, where is that compute? Like, maybe we should make sure that, like, ISIS isn't building data centers. That might be, like, a good plan.
Alex: Can we just reflect for the moment that if agents take over the internet, what will happen to all the ads? The ads have taken over the internet, so maybe this is a net win. Maybe we should be leaning into this.
Jim: But again, they'd still have to get the... if they built the data centers, they would still have to get the compute from NVIDIA. You know, like no one else, like you're not getting a chip big enough. So then you've got to look for a chip fabricator that ISIS is building. It's like, it's just, it's so... ridiculous though the once you step beyond the all ai is going to kill us and go into how is it going to kill us you instantly realize that it's just all ridiculous and they're trying to sell a some either some sort of regulation that's only that they have control over or they're trying to sell their ipo and again to alex's first point i'm just sick of it just shut up You're the CEO. If you have to stop it, stop it.
Will: Actually, I think almost a nuance though to your point, Alex. I think maybe it's not the marketing teams because the marketing team keeps putting out videos of, you know, you being able to generate, you know, a version of Mario or like, you know, be able to have it draw a canvas for you. It's the freaking PR teams.
Alex: Yeah, fair cop. I'm about, I don't mind that sort of marketing. It's a bit sloppy, but sometimes it's fun. So like, yeah, the marketing teams, they can stay. The PR people though, oh boy, when the revolution comes.
Jim: You want to be educated about it. You don't want to be sold on it. That's what you...
Will: So we might just touch on the last topic of the week then, guys. I'm starting to anecdotally just see a few businesses and kind of like smaller tech businesses that are offering these things as services is fine-tuning as a service. And I thought this would be an interesting one to talk about the idea that... using these kind of general purpose frontier AI models for all of your tasks not being the best possible outcome for all use cases or the best method for all use cases. And actually, if we were to take an open source model and fine tune it on that specific workflow that you're working on, But you can get things like, you know, a 50% better performance on that domain-specific task at a 90% discount compared to using those frontier agents.
Jim: So, yeah, I thought this might be an interesting kind of business model to talk about.
Will: Not sure if either of you have done any fine-tuning before, but do you see that there's potentially, I suppose, a business model here? And do you think that we're going to start to see more of this?
Alex: This is back to the future. This is like good old hosted models as a service and like MLOps like we did it five years ago. I think this is fantastic. You know, it's kind of good. We've had a bit of negativity and let's talk about positive developments. The idea that you can save so much money on these sorts of unit tasks is not a terrible surprise. It's basically right-sizing compute or right-sizing statistical compute in this particular case to solve specific tasks. I mean, I don't think it's necessarily the right option for every business. So for example, if you don't want the extra headache of curating data sets and you just want to throw something at a problem to see if there's a solution there, for example, I think general purpose LLMs, they don't need to be frontier models. They can be a couple of generations old and you still get perfectly serviceable performance. But particularly in safety-critical applications or things that require a degree of sovereign hosting or whatever, yeah, about it. Like if you're an airline and you want specific safety checks to be done through documents or tech specs or something like that, or, you know, particular quirks of the industry to be respected. Fine-tuning onto a corpus of that information to ensure that your safety margins are, like, airtight. Yep, go for it. If you're just a small business experimenting with models, it's probably not worth the headache. But again, your costs are probably pretty low anyway. It's only once these workflows scale or the cost of getting things wrong really goes up that this becomes a necessity. But it's easy. It's battle tested. It's tried and true. I think every company should have this thing in their back pocket for those problems where they just need one thing done over and over again.
Jim: No, for sure. And in the actual use case, it's fantastic. But to Alex's point, a lot of the use cases just don't apply. And a lot of the small and medium-sized enterprises just won't have the data or the quality data that they're going to need in an accessible fashion in order to fine-tune something. And then the other side is they're not quite there yet. anyway you know it's still a lot of the time a or in a couple of scenarios where this has actually been researched you know there's a couple of medical scenarios where they were trying to find you know see whether a fine-tuned model could be a harness, and the harness wins by double-digit percentages every single time. You know, it's just because they're not quite there yet. You can see, and again, back to MLOps days, you know, is it a fine-tuned or a couple of fine-tuned layers on the top of your model was the answer. It's... it's and it worked and it was able to be done and you could do it you could create enough scenarios or get enough features out to work out what you wanted to fine-tune on However, I just don't think we're there yet with the LLMs. And the harnesses can still perform so significantly better. And so it still comes down to, yes, this is where it's going to be in the future. And I love that businesses are starting to offer it, but it's a little bit of snake oil at the moment because... They're selling it, but no one's going to have the data they need to actually get a success out of it. And they're better off investing that money in some guardrails and some really well-defined processes to build into a harness first. Then, let me tell you, your fine-tuned model is then going to make that exponential.
Alex: I don't necessarily agree that it's snake oil at the moment. I think problem selection is really key here. Like the medical use case, legal, would be another one where it's very precise and it's very amorphous and it's often badly written or inconsistent or reliant on a particular understanding of how language is used. And in those cases, yeah, you know, there's the cases of being prescribed the wrong thing, which is, you know...
Jim: But it's also a very nuanced... industry really everyone just thinks about law as law but you would have to fine-tune on you know industrial disputes in states in the us that all meet this same legislation criteria like it would be so specific what you'd actually have to fine-tune on in order to get that sort of result I think those are like your airport, your airline example for safety checks before was a great scenario. You know, insurance as well. You know, I think there's enough limited verticals in insurance that you'd be able to know or, yeah, properly fine tune a model. But again, you've got to have the data in a way that is structured well enough. And you'd have to tell the insurance, you know, like, the insurance company would have to stop changing the policies.
Alex: And I guess a lot of these problems, right, comes down to problem selection. It comes down to the right tool for the job. And if you're trying to throw AI at a job that just should not be put in the hands of AI, you're going to have a bad time.
Will: I think a little bit about these types of things because I'm often trying to figure out where's your company's alpha now that these AI agents have such a vast amount of knowledge. What knowledge do you have beyond your customer list and your brand, like Alex was saying? An interesting one that came to mind with this was because I had a background in macadamia nut processing is our business probably had more written down about macadamia nut processing than is available on the public internet. It's just not something that's widely written about and shared online. So that is an area where whenever we tried to use AI on any of the use cases, it just didn't understand the domain at all and it was just totally unhelpful. And so there's just zero uptake of AI in that domain at all. So, you know, it's interesting to me the idea of, you know, could an agent be much more helpful in that space if it was fine-tuned?
Jim: Yeah. And I think it would be, but, you know, it would have to come down to... know a collaboration but the you know between macadamia producers you know to all contribute what they have written down and all that sort of stuff i think you know and i think it definitely would yeah like domain specific as opposed to business process specific yeah that's and i think that would be what's interesting as everything moves forward can, and it, maybe it comes down to like industry associations to take that role on, you know, and say, Hey, we want to, we want to create a, a airline safety model. Hey, we want to create a, a man, a macadamia nut industry model.
Will: That'd be interesting then where it would shift the liability. Yeah.
Jim: Again, it just becomes a quagmire the moment you start stepping into it, but that's got to be where everyone goes. We need some sort of collaboration rather than this competition race to an IPO just so we can get our $2 trillion and no one else can.
Will: All right, well, I think that's a wrap for the week, guys. Who's your business idiot for this week, Alex?
Alex: Come back to me. I kind of haven't prepped an answer for that.
Jim: Well, I think we missed him last week and because of, again, just the way it goes. But I think we've got to come back to the anthropic – the 10% that AI is going to kill us within the decade. And so to me, it's not Jacob or whatever his name is, the guy that worked at OpenAI and then worked at Anthropic for six weeks. It's the head of safety at Anthropic who was the one that actually came out and said, oh, yeah, no, for sure. We all think there's a 10% chance that it kills us all within a decade. I mean, that guy's just an idiot. Like... he it's not only a business idiot he has to just be an idiot you know like it's i just like in the leader you say that sort of thing in the lead up you know and so i'm taking it all on face value here but you say that sort of thing in the lead up to a two trillion dollar ipo then you're not just a business idiot but you're certainly the business idiot this week
Alex: Yeah, on reflection, that's very much a put your foot in your mouth and take a big chomp sort of moment. Yeah, perhaps the PR department should have been involved in that one. You would have hoped, but apparently not. Here we are.
Will: For me this week, I just keep coming back to the idea that if you're developing an AI that is potentially as dangerous as they're saying it is, then you can just stop and you're not. So for me this week, it's Dario.
Jim: That is good. Yeah. I'm happy to... It's all along the same lines.
Alex: Yeah, I just don't have a throat to strangle for the, oh, we'll put data centers in space. People just don't understand how heat dissipation works and it drives me nuts.
Jim: They're going to keep talking about it though.
Alex: Oh, for the next 10 years.
Will: So it'll be sort of the orbital data center founders for you, is it?
Alex: Yeah, well, I mean, look, I wish them all the luck getting the cooling fluid into space. It turns out heavy lift is heavy.
Jim: The big one for me is how, like, to get to achieve the latency, you've got to keep them, you know, pretty much low Earth orbit. We're getting a lot of stuff up there at the moment, you know, like...
Alex: Well, low Earth orbit also, like if you're low enough, you're going to de-orbit pretty quickly. You'll be catching the atmosphere.
Jim: And so I guess you can have lasers and they're super quick, but just the way they talk about it, you've got to, again, they're smarter people than us working on it, but it all just seems like they're all talking their book at the moment.
Alex: It's going to be fun. 50 years' time. We can't put anything into space anymore because of Kessler syndrome. And we're like, oh, yeah. We had this dumb idea that putting data centers in space would somehow get us good latency over Starlink.
Jim: Uncle Elon gaslit us all so well just because he was the most successful entrepreneur we'd ever seen. We fucking believed him. He fucked off to Mars. He's probably still out there. We can't speak to him. We've lost the satellites. Yeah, that will be a very interesting turn of events, you know, and so we'll come back in 50 years and work out how close that was.
Will: I think Elon would be happy if it's him and 10,000 optimists sitting on Mars.
Alex: Logistics is hard. Remember, amateurs talk tactics, professionals talk logistics. Hey, Elon. Invoices in the mail.
Will: No, good work. All right, boys. Excellent. Thank you all, and we'll see you next week.
One email a week. No courses, no funnels.
Sign-up is not wired up yet. Check back soon.