Closed models doesn't mean monopoly, cause they compete. And if the best thing for humanity is for good open models to exist, even those evidently required good closed models to copy from.
We have an entire generation of technologists convinced that you need a monopoly for a business to be viable, so every tech company is either trying to subsidize its way into a monopoly or running a deadweight loss.
Unsurprisingly, the resulting economy in aggregate is non-competitive, as it relies entirely on supply chains from competitive marketplaces.
> The state of the art models are going to get better and more expensive and smaller models are going to get cheaper.
Why do you think this will be true?
Right now I see the major US labs betting on gaining an advantage from having way more compute, and I see Chinese labs competing with one another in a resource-scarce environment, so they place much more emphasis on compute-efficiency.
But the supply chains that feed into the massive data center growth in the US are strained; there are energy, memory, and logistical bottlenecks to name a few.
In the medium-long run, compute capacity will not grow exponentially forever. Somehow it has for decades, but there can be no infinite exponential growth, and that point may be when the planet really starts to cook itself.
Maybe the US labs will become more compute-constrained, and then have to compete on efficiency.
Or maybe things change fundamentally in some other way I'm not thinking of.
The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly. They can't burn trillions forever. Nvidia's 75% profit margins are not sustainable forever.
>The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly.
By all accounts the AI capex boom is justified up by actual usage, rather than some nefarious plan for "locking up the supply of computing infrastructure". Just look at people complaining about claude availability and anthropic adding various load-shedding measures a few months ago.
Right but that could be more evenly distributed. There is a circular trade right now giving these few players near infinite resources that is blocking that from happening.
Commoditize your complement - I expect to see this most in consumer AI (after that starts actually working...)
It will be important for Apple to have good enough, cheap local LLM models that run on-device.
If the barrier to performance shifts from fundamental model capability to context collection and management I would expect to see folks focused on that problem continuing to drive open-weight LLM model development in some shape or form.
Probably. There are a lot of countries, especially third world ones, with very lax legal systems, not to mention the multitude of countries where law basically doesn't exist.
In the English language, "America" refers to a country. It is synonymous with "The United States of America". I say this as someone who lives in the same continent as that country, but not in the country itself.
Maybe you're thinking of "North America", "South America", or "the Americas".
I've got news for you there. The Biden administration tried to take the first real antitrust action in decades and then suddenly a bunch of tech oligarchs switched from supporting Democrats to supporting a convicted felon
To me the biggest (but not only) issue is that blindly connecting sensitive tools to 3rd party services has been normalized. Every time I hear the word "claw" I cringe...
Pretty much every major American inference provider claims to make a profit on API-based inference. Consumer plans might be subsidized overall, but it's hard to say since they're a black box and some consumers don't fully use their plans
Selling inference is not fundamentally different from selling compute - you amortize the lifetime cost of owning and operating the GPUs and then turn that into a per-token price. The risk of loss would be if there is low demand (and thus your facilities run underutilized), but I doubt inference providers are suffering from this.
Where the long-term payoff still seems speculative, is for companies doing training rather than just inference.
There’s a lot of debate over what the useful lifespan of the hardware is though. A number that seems very vibes based determines if these datacenters are a good investment or disastrous.
I specifically remember this debate coming up when the H100 was the only player on the table and AMD came out with a card that was almost as fast in at least benchmarks but like half the cost. I haven't seen a follow up with real world use though and as a home labber I know that in the last three weeks the support for AMD stuff at least has gotten impressively useful covering even cuda if you enjoy pain and suffering.
What I'm curious about are what about the other stuff out there such as the ARM and tensor chips.
All of them. It's simply impossible to sell tokens by usage at a loss now. You'll be arbitraged to death in a few days. It only makes sense to subsidize cost if you're selling a subscription.
If they were they would show evidence because they'd pull in more investment. I don't believe their claim that they make profits on inference, especially not with reports like this coming out.
I'm not sure how everyone in the US forgot that monopolies are bad
reply