Usage is actually Claude now because of Opus 5.5 since it a better model that Astra. I maxed out my 200$ Claude plan with 10b token on Opus 5 and 5.5 is cheaper. I maxed out two Codex accounts with like not even 5b tokens.
With other software, devs convince their managers of the importance of using open source stuff in their stack. With AI, it's usually managers choosing what models to use for the devs. The US labs don't need to give a damn how much devs like open source
This isn't about liking open source. This is about the labs just being cool and doing cool shit instead of the opposite which is Anthropic where all they talking about is killing everyone and taking everyone's job.
Well maybe they should because companies that do cool shit tend to attract people wanting to do cool shit and those are the people you want to be working at your company.
Unless you want your company to become Google which use to do cool shit but now they dont and now they don't have a single person at the company that can build cool shit so they just release a bunch of lame shit.
I was dev, and now I am Senior level manager.
Open Weight models are current main focus for many companies with full alignment with top management for very simple reasons:
- stable and predictable performance (no pre-launch models degradation)
- ability to tune them for specific business cases (though still rare tbh)
- better (at least 60% Opus vs Kimi (real,3rd party)) and more competitive pricing
- flat pricing if tokenusage is big enough to justify renting GPU
- decent quality
- much higher guarantees that data will not be sent somewhere (assuming 3rd party inference providers)
- and cherry on top: flat and minimal pricing with absolute confidentiality using Alibaba Apsara stack of recently released AMD Instinct Coder box[1]
> The US labs don't need to give a damn how much devs like open source
In the short term, true.
In the long term, unknown but typically when you hold progress that way while other countries don't you at best end up becoming siloed while the rest of the world continues on without you.
You mean all of the frontier models that the Chinese distillation clones are copying? Yeah kinda cool imo. If a dashboard showing training for a model that doesn't even come close to anything us labs have released in 6 months is "cool", then you're a loser
Anthropic and OpenAI literally stole from every human in history and youre out here complaining that the Chinese are distilling models and releasing them to the public?
Because without those labs to distill from the pathetic Chinese labs wouldn't have anything. Im not impressed by them copying US labs not sure why you are. But go off ccp bot
Maybe the Chinese app is different, but the Kimi app shows reasoning traces so how can they route the request to Claude models which don't show reasoning trace and still provide a reasoning trace on the UI?
So I think they are in fact just lying because it doesn't make sense why they would route requests to Claude over their own model.
Claude's reasoning traces are encrypted, but there was a design flaw that made it possible to extract them: https://stolen-thoughts.com/
And it makes perfect sense for them to route some requests to Claude, as it lets them do competitor research on realistic data. I suspect Anthropic does similar competitor research on Kimi, though presumably hosted on their own infrastructure and maybe without serving the results to customers.
that can be achieved by storing user traces and then running evals on own model vs competitor model. you don't need to route live customer requests to a competitor for this.
For agentic tasks where the model outputs tool calls that run on the customer's computer, you can't just store and eval later, because then the execution environment is no longer available.
> I think it's very clear that DeepSeek is obviously the best AI lab in the world.
It's pretty clear they're the best at what they're optimizing for - which does seem aligned with what a lot of people on HN want from models - but not everyone...
Considering the fact that Google/Anthropic/OpenAI have WAY more compute and the race is this close, it's obvious that DeepSeek/GLM/Qwen teams are better or we're approaching a wall in terms of progress.
Basically, the nice folks at OpenAI or Anthropic saying: "You distilled from our model which is built on the stolen data that we ourselves suctioned up from the entire internet without regard to copyright law! Only we get to vacuum up the whole internet. That's our special prerogative.".
"So you think in 3 years AI is going to solve longstanding math problems because it was used to write some coherent sentences?" — people with the same amount of foresight in 2023
Are you saying at anything that can solve longstanding math problems necessarily has the means, motive, and capability to kill 8 billion people in just 3 years?
No dude, it was an analogy. Sorry to pick on you, you're just another ignorant uninformed take in this thread, but come on. This technology is progressing at an insane rate, the shit it's doing now is science fiction from just a year or two ago, and people just keep on moving the goal posts, it's maddening. The technology is unpredictable and that in itself has risks. You must realize that the window of possibility is widening the further into the future we go! Can you pick people with epistemology you trust and see if there's anything you can learn in this moment? Can you try to challenge your own ideas instead of retreating into comfortable certitude? There are real risks here, they are worth taking seriously, and serious people are doing so!
I think they will be, Nvidia has released open source model that also included the data it was trained on. I think they are the only ones that have done that.
The government should give incentives to do things that benefit the future of the country. All auto manufactures had the same incentives but instead of building actual good vehicles they instead shipped more production to Mexico and continue to build the absolute most dog shit vehicles in the world.
Not a single US auto manufacturer, outside of Tesla, has a good EV, nor FSD. Waymo is buying thousands of Chinese cars because of how shit US car manufacturer are.
They are the worst of the worst. They shouldn't even be in business they are that bad at building cars which they literally invented.
reply