Hacker Newsnew | past | comments | ask | show | jobs | submit | greenmilk's commentslogin

The Chinese gov't and also the rest of humanity...

I'm not sure how everyone in the US forgot that monopolies are bad


Closed models doesn't mean monopoly, cause they compete. And if the best thing for humanity is for good open models to exist, even those evidently required good closed models to copy from.

Who's got a monopoly though lol?

We have an entire generation of technologists convinced that you need a monopoly for a business to be viable, so every tech company is either trying to subsidize its way into a monopoly or running a deadweight loss.

Unsurprisingly, the resulting economy in aggregate is non-competitive, as it relies entirely on supply chains from competitive marketplaces.


> The state of the art models are going to get better and more expensive and smaller models are going to get cheaper.

Why do you think this will be true?

Right now I see the major US labs betting on gaining an advantage from having way more compute, and I see Chinese labs competing with one another in a resource-scarce environment, so they place much more emphasis on compute-efficiency.

But the supply chains that feed into the massive data center growth in the US are strained; there are energy, memory, and logistical bottlenecks to name a few.

In the medium-long run, compute capacity will not grow exponentially forever. Somehow it has for decades, but there can be no infinite exponential growth, and that point may be when the planet really starts to cook itself.

Maybe the US labs will become more compute-constrained, and then have to compete on efficiency.

Or maybe things change fundamentally in some other way I'm not thinking of.


The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly. They can't burn trillions forever. Nvidia's 75% profit margins are not sustainable forever.

Things will normalize, but it will take time.


>The labs have a perverse incentive to make things as expensive compute wise as possible. The only thing keeping this somewhat in check is competition, but it's intentionally being gatekept by locking up the supply of computing infrastructure. With 3 players it's pretty easy to collude even if indirectly.

By all accounts the AI capex boom is justified up by actual usage, rather than some nefarious plan for "locking up the supply of computing infrastructure". Just look at people complaining about claude availability and anthropic adding various load-shedding measures a few months ago.


Right but that could be more evenly distributed. There is a circular trade right now giving these few players near infinite resources that is blocking that from happening.


Commoditize your complement - I expect to see this most in consumer AI (after that starts actually working...)

It will be important for Apple to have good enough, cheap local LLM models that run on-device.

If the barrier to performance shifts from fundamental model capability to context collection and management I would expect to see folks focused on that problem continuing to drive open-weight LLM model development in some shape or form.


>so they place much more emphasis on compute-efficiency.

Maybe on training, but on inference they use more tokens than comparable western models.

https://artificialanalysis.ai/?output-tokens=intelligence-vs...


I believe this is the process that defines Julia Evans' writing (and cartooning)!


Is there any country where it is?


Probably. There are a lot of countries, especially third world ones, with very lax legal systems, not to mention the multitude of countries where law basically doesn't exist.


Haiti comes to mind.


At least in Germany in B2B contracts that might be possible.

For b2c, no chance


America


America is a continent. Maybe you were referring to the US


In the English language, "America" refers to a country. It is synonymous with "The United States of America". I say this as someone who lives in the same continent as that country, but not in the country itself.

Maybe you're thinking of "North America", "South America", or "the Americas".


I've got news for you there. The Biden administration tried to take the first real antitrust action in decades and then suddenly a bunch of tech oligarchs switched from supporting Democrats to supporting a convicted felon


To me the biggest (but not only) issue is that blindly connecting sensitive tools to 3rd party services has been normalized. Every time I hear the word "claw" I cringe...


Are any inference providers currently making profit (on inference, I know google makes money)?


Pretty much every major American inference provider claims to make a profit on API-based inference. Consumer plans might be subsidized overall, but it's hard to say since they're a black box and some consumers don't fully use their plans


Third parties selling open-weight inference on OpenRouter are surely selling on a profit. Zero reason to subsidize it.


They could be VC-backed and selling at a loss to grow market share


Selling inference is not fundamentally different from selling compute - you amortize the lifetime cost of owning and operating the GPUs and then turn that into a per-token price. The risk of loss would be if there is low demand (and thus your facilities run underutilized), but I doubt inference providers are suffering from this.

Where the long-term payoff still seems speculative, is for companies doing training rather than just inference.


There’s a lot of debate over what the useful lifespan of the hardware is though. A number that seems very vibes based determines if these datacenters are a good investment or disastrous.


I specifically remember this debate coming up when the H100 was the only player on the table and AMD came out with a card that was almost as fast in at least benchmarks but like half the cost. I haven't seen a follow up with real world use though and as a home labber I know that in the last three weeks the support for AMD stuff at least has gotten impressively useful covering even cuda if you enjoy pain and suffering.

What I'm curious about are what about the other stuff out there such as the ARM and tensor chips.


All of them. It's simply impossible to sell tokens by usage at a loss now. You'll be arbitraged to death in a few days. It only makes sense to subsidize cost if you're selling a subscription.


How do you arbitrage closed weight models? Who would buy from a middle man at increased price? Who is offering Priceline but for tokens?


Google definitely makes money in other areas. Do they make money on inference?


If they were they would show evidence because they'd pull in more investment. I don't believe their claim that they make profits on inference, especially not with reports like this coming out.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: