Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
The current balance of power in open models (interconnects.ai)
120 points by gmays 1 day ago | hide | past | favorite | 54 comments
 help



It's becoming more clear that for big enterprises to really adopt AI, they need to use open models. Especially if they want to own their own intelligence, which they should.

I spent the last 2 days building basic AI agents to automate some mundane supply chain workflows for a large company. Those seemingly boring workflows had bank statements, supplier IDs and other sensitive information.

For me it was all alarm bells, there is no way they can afford to give closed models access to this data. I was compelled to figure out an open model based solution for them, which made me realize that this is probably the only way for enterprises going forward.


Everybody including companies need their own models, especially considering that llm providers like openAi have no problem siphoning off your data and intelligence by inspecting metadata. and claiming any resulting value as their own.

I don't get this. It's not like the models are running on their own GPUs.

So if you're running open models on AWS GPUs, you might as well run Claude (which AWS supports, and doesn't share any data with Anthropic).

Same with Azure/OpenAI.


> (which AWS supports, and doesn't share any data with Anthropic)

I want to believe (pinky promises from terms of service don't count)


If you don't believe AWS follows their ToS, you shouldn't use any cloud provider for CPU/data storage/anything else either. That's not a mainstream position in the industry.

> which AWS supports, and doesn't share any data with Anthropic

Ah... oh.

Well, it's a nice thought.


Sufficiently large enterprises could get their own GPUs to run the models though.

You really don't need to be large. $100k can buy you a lot of compute and it's less than hiring an engineer. With that kind of money you can build an LLM server for a dozen people.

One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.

Open models on-prem is the future, not a single doubt in my mind.


As long as you don't have "realtime" workloads, owning the GPUs quickly becomes the economical option. The main cost problems is in e.g. chat applications where the workload is spikey, and users expect an near-instant response, for which you need to scale the GPUs to the highest spikes of the workload.

Why does this necessitate using closed models? I don't see the difference between putting your data on a cloud DB and using a cloud model.

Because the big labs are incredibly untrustworthy. Their leaders are publicly lying constantly.

I think that's true for Sam, not sure about the others.

https://www.openaifiles.org/ceo-integrity


Same here. We are at a stage where we are evaluating cost/benefit for my enterprise. I am heavily pushing for in house hardware with open models.

If you don't mind me asking: what is your goto for this? "Building Basic AI Agents to automate some mundane supply chain workflow".

The Enterprise plans with OpenAI and Anthropic include clauses that they won't use your data for training (unlike the 'civilian' plans).

I guess it depends if we believe that or not.


I guess it depends if we believe that or not.

I see the flaw in your clever plan...


I've actually flipped on this the last few days because of liability.

The big labs are going to be on the hook for rogue behavior by Claude or Sol. Customers will be able to sue for damages and deflect regulators if their customer data is abused or their agents attack external services.

If you use a Chinese OSS model and it goes rogue? Yeah good luck with that, your shop is 100% on the hook.


People - usually suit-wearers - have been making this spurious claim for decades, but it doesn't hold water.

The largest of the finest print reminding you that it's 'sold as is' (or more encompassing variants that might continue '... with no warranty for fitness of purpose') means that liability remains in the lap of the purchaser / consumer / operator.

(This has been a source of immense frustration over my career - where such people have assured me that they have 'recourse' (it's always vaguely described) by spending money on proprietary products & services, rather than opting for functionally equivalent or superior free options.)

I think your third paragraph is implying a distinction (or conflating the difference?) between LLMaaS's and self-hosting publicly available models.

If it's just where it's hosted that provides the legal insulation then things like OpenRouter would give you that. (But again, I suggest that it would not.)


This all depends on the SLA that gets signed.

If a frontier lab is willing to draft an SLA that assumes liability, corporate will pay for it as long as the cost/benefit is in favor of it over insourcing.

Right?


Sure, but that's quite a fanciful universe you're imagining there - the feasibility of a corporation obtaining insurance to cover that offer of liability ownership has got to be close to zero.

> The big labs are going to be on the hook for rogue behavior

They haven't so far.


I think the minute a big lab is found to be liable, the whole edifice along with trillions of dollars of investment and VC comes tumbling down. I think that is part of the reason the labs are pushing for more regulation. They can say "We're not liable, we complied with all of the regulations". The actions of multi-billion parameter models trained on data harvested from millions of Internet users over the years can never really be understood - if a business is found to be liable for that, then nobody would ever operate in that space.

I heard a similar theory on CNBC's Squawk Box.

Sorkin presented a theory that Nvidia bought HuggingFace to protect OpenAI from legal consequences. That a lawsuit determining accountability of actions by LLMs could threaten the AI financial network.

https://youtu.be/4qV5WWgFTS8?t=323


They are pushing for regulation because they are human being and don't want all human beings to die.

I'm sorry you are so jaded you can't recognize honesty when you see it, but that is what everyone deep in the AI space, including the non-executive researchers, are worried about.


Enterprises have barely deployed empowered agents yet. The models capable of doing this have only been available for months. Give it a little time, it's coming.

... when are large companies on the hook for anything, ever?

I mean, hypothetically, yes, but class-action lawsuits get settled out-of-court, the lawyers get paid in Ferrari-multiples, the plaintiffs get paid in McDonalds coupons that expire in two weeks.

Slaps-on-the-wrist are written into the laws; a million-dollar fine is existential for a small company, and likely not even a line-item at Anthropic.


I know this is the risk management answer, but when you're the featured story on the news because of a data breach, noone hears "butbutbut it's Anthropic/OpenAI/whoevers fault...". So it's a balance between "there's someone we can sue" and "what's our reputation worth".

It actually does help a lot to be able to say your OpenAI agent was the fault. People recognize the name. The press doesn't want to write about Better Home Life Insurance agents running loose on the internet, nobody cares.

Impressive you know. They are a few months behind but still the world uses their models. Bet you because the users suspect that Americans will pull the rug from under them. Few months behind versus being left behind forever. Would have been nice to know the differences in spending between the models but I suspect the data is impossible to collect.

It's not a suspicion, it has actually happened. US government briefly banned Fable for non-americans, and the american corporations pick and choose who has access to their unrestricted models which in practice excludes even american citizens, to say nothing of foreigners like myself.

They are literally segmenting the world into haves and have-nots, just like nuclear power. Thank god China is out there and constantly undermining them with their non-stop open weights model releases.

The best situation for us mere mortals is one where they struggle against each other endlessly without any hope of victory. The second either the US or China wins, it's pretty much over for us and unimaginable oppression will quickly follow.


They are good enough and vastly cheaper. It’s not fear of the future. It’s economic reality of the present.

That’s certainly part of it. I think people underestimate the value of getting to ‘good enough’. Once that’s reached, the economics start to be more of a factor in decision making.

However, I think geopolitics has also become a factor as the very existence of this article indicates.


> However, I think geopolitics has also become a factor

The world powers have only said things like "bigger and more important than the Manhattan project”.

So what gives you that idea?


Really important work here. Really glad it's being shared a) with US Leadership and b) with the community.

DeepSeek and Qwen set a floor American open-weight models can't match without subsidy or a different cost structure, and that price/performance ratio drives adoption more than benchmarks.

It's worth noting that Chinese AI labs aren't profitable, so in some sense they're not doing it without subsidy either.

Anthropic is profitable now. They have massive debt looming over them, but they're at least making more than they spend each quarter.

I’m super uncomfortable with the idea of a single company or country holding all the cards when it comes to AI. So from that perspective I welcome the Chinese models.

Additionally, the big labs’ formulation of AI as a US-vs-China national security issue is very convenient for keeping those pesky regulations at bay, and for ensuring the AI financial bubble doesn’t pop before everyone can unload in an IPO first. Both of these offend my sense of fair play.

The Chinese labs aren’t just putting models out there, they’re publishing and sharing their research and innovation too. That’s just a better way of doing science and contributing to the development of the whole field.


Meta's Llama strategy really shifted things, democratizing access and accelerating innovation. The gap is closing way faster than expected.

The open-weight vs. closed model is wrong comparison, I suggest we call them local-installable models and cloud-only models.

Some open-weight models aren't so open in their license.


This is a prepared statement for congress by Nathan Lambert, author of https://rlhfbook.com/ and AllenAi employee (one of the few who have shared their models as open source)

Unfortunately Nathan recently left AI2, along with some of the leadership most involved in their truly open models. https://www.geekwire.com/2026/allen-institute-for-ai-ceo-ali...

I did leave ai2, but I'm still in the open-source, non profit sector :)

Thanks for what you've done and continue to do for the ecosystem!

So many paragraphs of Claude explaining why open models will be frontier soon, which is the daily propaganda. Beneficiaries are startups selling the "open works" narrative and Big AI for the "we are not an oligopoly" narrative.

He says he is briefing Congress members. As presented, not a single one will understand anything, which is perhaps desired. It is very hard to tell what game is being played here unless Substack released bulk subscription numbers.


This reads like a Markov chain, those are neat.

Data points from my own usage, for whatever they're worth[1]. I'm just an individual contributor and hardly represent enterprise or large orgs:

I've been on Kimi, with a little DeepSeek-V4-Pro, GLM 5.2/5.3, and MiMo thrown in, for probably about a year now. It's great here!

1) For DeepSeek, I recommend their Reasonix harness strongly, due to its alignment to DeepSeek's prefix cache. It means mostly (95%+) cache hit input tokens, so very cheap large-scale code analyses and things that require mega context windows (at the cost of some attentional drift, yes). Reasonix does require that you send data to China. This is fine. I mostly use this for big, expansive ingestion of open-source codebases to figure out how something really works, usually something that documentation doesn't quite reach.

The economy of doing it this way versus American frontier model companies' token pricing cannot be overstated. I think I topped up $10 in June (2.5 months ago) and have still not burned through it, despite cycling untold tens of millions of tokens through it.

My biggest annoyance is that DeepSeek does seem to be considerably rate-limited of late, at least during working hours in Beijing, which is a range that I gather to be quite expansive there. I'm not blasting it with anything, I'm just noting that the agent takes 10-20 minutes to do stuff that takes much less time if I'm willing to pay the OpenRouter premium.

2) For most everyday stuff outside of where Reasonix + DeepSeek just makes overwhelming sense, I use OpenCode/Maki/Pi/whatever harness I feel like using today with Kimi K3, via OpenRouter. This does not require sending data to China.

I also use Kimi K3 in Zed via OpenRouter quite a bit, but sometimes like to mix it up with the other models.

3) Because I have the most experience with it, I can say with confidence that I would generally consider the SWE capabilities of Kimi to be on par with Claude, at least for the bottom 99% of purposes--and certainly, any routine business programming.

I think this has been true for a long time, well before K3. I've been using Kimi since K2.5.

4) For local hardware experiments on my MacBook Pro (M4 Max, 128 GB unified memory), Qwen3.6-35B-A3B (speed) and Qwen3.8-27B (intelligence, but slow). As has been widely noted, this amount of unified memory isn't as useful as it seems, due to memory bandwidth and decoding constraints, lack of tensor cores (on the M4 Max, anyway), etc.

A giant bag of memory isn't fast, but it'll let you load some impressively big models.

The future M5 Studio Macs will continue in this general vein, but will of course be somewhat faster, particularly due to the apparition of tensor cores in the M5 -- excuse me, "Neural Accelerators".

Still, if you really want to cook, get a real GPU. Real GPU running quantisations is still a lot better than a big slow bag of unified memory.

5) Overall, the Chinese models are simply excellent, and cater to lots of use-cases and tastes. However, I'll still tend to use Claude ($20/mo subscription) for general Q&A, whether of a technical nature or otherwise, particularly where web research and worldly knowledge is required.

[1] Disclaimer: this comment is an elaboration of https://news.ycombinator.com/item?id=49809605, which is not something I'd normally do. However, it seems a lot more relevant here than where I had originally posted it, in an article about gauging MiMo Pro v2.6 capabilities.


As I understand it, all real world use cases requires provider data centers and the local efforts are experimental.

So "open models" buy you little and the independence from the oligopoly is not achieved.


Yeah, that's broadly accurate, although details matter and local models are surprisingly capable. Local models are viable for a lot of small use-cases, but no, there's no deluding oneself that frontier quality doesn't require frontier size. However, everyone who thinks they need frontier "intelligence" for internal CRUD type tasks should look in the mirror and ask themselves if that's really true.

Still, even if you need other people's industrial-class hardware, open models offer a lot more freedom and options. They allow you to use GPU capacity from entities who are not themselves building or training models and are ostensibly disinterested.

There's a big range of possibilities here in terms of data sovereignty and so forth.

1) You can use OpenRouter to route your open model requests to US-based inference providers with ZDR (zero data retention), as far as you can believe anything in this world. If you look at who actually serves open Chinese models on OpenRouter, you'll see a lot of folks like Digital Ocean, etc. I suppose I can't vouch for their purity, no, but I'd much rather send data there than send it to Dario.

2) Or, you can rent GPUs from companies like Runpod or Vast.ai and serve some very sizable models to yourself (e.g. using their pre-built vLLM images). You can't serve a model like Kimi K3 to yourself that way, at least not in any economically reasonable way. However, a private H100SXM or B200 can go a long way. You could serve the big Qwens, or DeepSeek-V4-Flash--you could do a lot if you're willing to spend on a rented GPU with sizable VRAM.

3) Finally, if you have and want to spend $750K-$1MM+ (I suspect I'm low-balling at this point), and if you can get them in the current demand climate, you can absolutely buy 16 x H200SXMs, with the appropriate boards to take them, pay for 20-25 kW of cooling, etc., and run one of these models yourself, on your kitchen floor if you like. You simply cannot do that with Anthropic or OpenAI.


Thank you for sharing this perspective. I agree with the conclusions outlined. However, I would have appreciated more insights into the OpenAI OSS family of models like gpt-oss.

[flagged]


Not a fan of overregulation killing business, but Europe is definitely an important check against the surveillance powers of big tech companies in US. European regulations become a reference point for the rest of the world, especially poorer countries who dont have power to regulate US tech companies by themselves.

Zero Mistrals in the game eh

As an aside, it’s always funny to me that anybody outside of Europe listens to anything the Europeans say about AI. The Europeans literally have zero skin in the game. Their entire raison d’être is just to pass regulations that they then attempt to extraterritorially enforce, but they contribute almost nothing to the area of interest. It’s legit insane to me. The contributions they make vs the amount of inconvenience they are vis-a-vis compliance with their own esoteric rules probably fit some Pareto maximum of 20% of users requiring 80% of the work for compliance. The rest of us (ie, the non-European world) should honestly just point and laugh and then move on when these people try to influence things.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: