Absolutely! Chinese models are both cheaper and more capable in many cases, compared to the American models and their makers continuously fumbling or reducing model capability with each update. Deepseek decreased costs when they released Flash 4.1 you would not see any American company do this, in reverse they would try charge you more.
OpenAI decreased prices with the 5.6 model family.
And later they further cut Sol and Terra pricing by 20% (maybe only in the API) and Luna by 80%.
In fact Luna still outperformed DeepSeek Flash 4.1 in cost per task on Artificial Analysis when I last checked.
However, Luna is slightly less intelligent. I have a feeling that it's pretty dumb and prone to hallucination unless running at xhigh or max effort, where it somehow manages to work quite well.
I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.
The competition is great, and I hope Chinese models will continue to force leading US labs to offer models at a low price point.
That said, I don't think the Chinese labs have anything over OpenAI and Anthropic when it comes to capability or efficiency - I have no reason not to believe the US labs have even lower cost to serve the models.
OpenAI had to cut costs because of Anthropic. I also do not trust the benchmarks when it comes to models anymore. I have tried both Claude and OpenAI models and while it is true that the 5.6 series is smarter than Deepseek (at the time i tested it against 4.0) at that price it is still not worth it and sometimes randomly refuses to do tasks or stops midway etc.
Do also remember China is this far in the AI race despite all chip restrictions from America. If they were in equal standards I truly think Chinese models would have long surpassed American ones. Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling meanwhile their own models claimed to be Qwen¹ and their stance against open models is negative² and they still keep blaming China for it.
> Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling
Why wouldn't he? If there really was 25,000 accounts breaking ToS any CEO would at minimum be upset. Evidence of Claude distilling qwen would be damning but that a) makes no sense b) doesn't exist afaik.
They have the logs, created a report and sent a letter to Congress. Whether you believe it is entirely up to you. Given Chinese firms record on IP theft, it's entirely believable. I don't have any doubts, but I might question how they attribute it to a specific firm.
And they took multiple measures presumably to stop "distillation", such as hiding reasoning steps.
Chinese models kept improving in capability regardless, and are in some ways more impressive than Claude/ChatGPT.
So yeah, I think they are bulshitters. The can create reports and send letter to congress simply because they know if allowed to compete freely the Chinese models will eventually prevail.
Also, very rich of you to mention Chinese firms record on IP theft when Anthropic and OpenAI are companies entirely built on large scale IP theft.
So first it’s “Chinese companies cut costs, and you’d never see American companies do that”, and then when it’s pointed out that one of the leading American labs literally just did that, it’s “yeah, but they had to because of competition”.
What do you think is motivating the Chinese labs, benevolence?
> If they were in equal standards I truly think Chinese models would have long surpassed American ones.
Limitations often lead to creativity to overcome them. The Chinese AI labs have had to focus much more on efficiency so they got good at it. Meanwhile breaking new ground is often harder than replicating it. So even if they had matching compute it's not a given they'd be better.
> I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.
So you don't have much perspective on things, it seems. Let me introduce you to the GLM 5.2 and then 5.3/5.3 flash series of... "oh, wow, I should have bought some RTX PRO 6000's while they were 'cheap'" stage of progression.
As someone carrying multiple max subscriptions to both claude and codex - primary workhorse is glm 5.3 flash running on rented GPUs for less than a latte/hr.
I also found qwen 3.6 27B nearly useless for my own needs. DS4 flash 0731 and then 4.1 have been nearly as eye opening as glm 5.3 flash, but have their own warts.
Have you tried Qwen 3.8 Flash Next? You can run it on one spark with reasonable context sizes at about 30 tps, and it's as good as DS Flash 0731. Maybe even a tie with GLM 5.3, though like everything it depends on the use case.
I’m having the same issue. Hold max subscriptions on both frontier labs but I’ve been forced to use open source models because token limits are not what they used to be. So I end up using Astra and Fable for reviewing, and open source models for implementing.
> OpenAI reduced prices and Anthropic increased weekly usage limits.
As a Max x20 and Pro x20 subscriber, can tell you that it doesn't matter since they continually move the baseline of token use So in practice you feel that you're continually getting less from your subscription.
While it never happened to me in the past, i reached my weekly limit within 3 days using Opus 5. And the Open AI weekly limit essentially is a Claude Max x20 5-hour limit. Not even talking about the baseline in intelligence : on release day Astra was so good that it lead me to move to Pro x20. Now it's dumb af and token use is insane.
Deepseek 4.1 Flash has been a lifeboat for me, finally able to work without being constrained/distracted by limits and with what is in my view even better intelligence than Opus 5 for a fraction of the costs. DS is not messing up my brain with load-bearing pseudo jargon in every sentence. It respects coding guidelines, and completes even the most complex tasks most of the time in one shot.
DS 4.1 had been able to add complex features to my repo without breaking a sweat (330k lines of F# + 4M circa lines of an Angular frontend). Writes very idiomatic F# and respects our guidelines and style perfectly. Just completed an extensive UI/UX research and implementation work.
I am ditching both x20 subs and will only keep a Pro x5 because wife does a lot of design work and needs solid image generation capabilities.
Absolutely! DeepSeek-V4-Flash-0731 has become my daily driver. It's pretty amazing what it can do for what it costs at deepinfra.com (I don't use deepseek as a provider since they train on your data [at least their honest about it]). GLM-5.1 was my daily driver before that and Kimi K2.5 before that.
My primary use is AI coding agent. Its vastly cheaper than Kimi K3 and I haven't found a scenario where I really need Kimi K3 versus smaller models. GLM-5.3 Flash is good but there is series of bugs in the vllm middleware that prevent GLM models from getting all of their reasoning content returned to them that impairs inference quality. A lot of inference providers use vllm which makes it hard to find a good provider for GLM. I've been using friendli.ai but using GLM-5.3 Flash from them is more expensive then using DS V4 Flash from deepinfra.com simply because deepinfra.com is so cheap. The DS V4 Flash cost at together.ai is similar to the GLM-5.3 Flash from friendli.ai or at least that's what I found in my benchmarks a week ago: https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-r...
4.1 consistently surprises me in capability for the price. And I don't think I'm the only one. It's been dominating the leaderboard at OpenRouter, and I just got an email today from Fireworks saying they were _raising_ the price by about 30%. I'll probably switch, because their infra doesn't support being the highest-cost, but it's still telling.
I tried it a few times and liked the speed, but often found it ended up looping, i.e. repeating the same token sequence (e.g. the same sequence of 5 paragraphs) over and over again until it hit the max output limit. This doesn't end up happening every session, but does every now and then.
My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results
I was using a vibe coded bare bones harness.
I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.
I've had that looping issue with open models too. But never 4.1. I wonder if it's a model + harness combo? But yeah, one loop issue and I'm done with a model forever.
Yeah, that's what I'd been leaning towards.
No mcp, but I'll see if I can reproduce and debug it, since other people don't seem to have that problem as badly as I've experienced it (and the idea of having a nasty bug like that bothers me).
No mcp support.
I'll try copying deepseek harness's basic tool call formats as a starting point.
I haven't tried 4.1 flash as I'm assuming its a preview. I did not get good results from the preview version of 4.0 flash (i.e. the one that did not include the month and date of release in its name).
Affordability is derivative of control which is really what I care about.
I'm just not going to build long term infra that depends on something that another person can and will - objectively based on experience - take away from me at some unknown point in the future.
The biggest benefit of open models is they keep all the other players honest. The extent to which they feel they can dictate terms is directly set by the threshold where they feel people will take the trade to run open models instead.
American models are on the frontier of capability. Chinese models are on the frontier of efficiency. The problem for American labs is that Chinese models are more than capable enough for the vast majority of applications that people care about at this point, so efficiency is more interesting for people.
it's really weird to me at the moment because both OpenAI and Anthropic seem to be competing in an extreme benchmaxxing contest on super intelligence that actually nobody cares about. I haven't really cared about model intelligence since about Opus 4.8. It is by far not my biggest problem. I don't need to replace or support Einstein in my production workflow. I just need basic intelligence that can equal a routine office worker - safely and reliably. What they doing - chasing super-intelligence but dramatically escalating risk - is actively what I don't need.
I really think they have drunk too much of their own kool aid and become completely detached from what the market wants.
The cost issue is obviously of prime importance, but I'd also argue that the transparency of the innovations creates a tremendous cross-pollination, and not only within the Chinese communities but in the US/Europe as well. How many of us are learning the practical aspects of actually running and building AI based primarily on open models? As an example - how far would the work of vLLM or SGLang or even NVIDIA itself (all random examples) be without these models and the challenges they pose?
The analogy makes little sense. The USA was not in front of the USSR and Sputnik merely showed that. It is at this point that the Americans woke up, put a lot of effort and finally were able to surpass the Soviets during the Apollo missions.
China was never ahead of the USA in AI. So perhaps a more proper analogy is the Moon landing. In real history the side that lost the race never got its mojo back...
I was simply saying that when (not if) Chinese AI models will pass Americans, it will probably be game over and Americans will never catch up, let alone become leaders again.
Check the names of the researchers in the DeepSeek's latest paper. Full of Chinese names. Check the list of names in Google's paper. A very similar view. Anecdotal, but quite thought-provoking...
I did try to use Chinese open models, but for my production work they simply couldn't cope at all; both GLM 5.3 and Deepseek v4 went into infinite loop and wasted my tokens until my OpenRouter wallet reached 0; good thing I didn't enable the auto topup. US models, by contrast, breezed past them.
Even for simpler tasks, Chinese models took long time to complete, and I needed to supervise closely. The price , in the end, didn't come cheap, mainly because too much time wasted on thinking.
So maybe one day Chinese models will squeeze out the American ones, but today is not that day.
So no, I am not excited about Chinese models ( just because its open weight and not American).
Not too be "that guy" (e.g. "you're using it wrong"), I just want to humbly ask — have you tried blacklisting "underperformers" in OpenRouter config?
Here on HN was a post few days ago titled like "so you want to use openrouter", there was a benchmark in capabilities between providers which showed some aggressively quantize and basically break models and tool calling.
I am in no way a professional power user, but I frequently suffered from "call fails" (e.g. unclosed tags, broken agent loop, broken thinking blocks), so I had to babysit agent on it's loop. After I blacklisted like 20 providers (I think most broken were Nebius and DigitalOcean) these issues completely went away. I had several agents work on my small tasks for 18+ hours with no issues.
Yep, I'm trending in that direction, and I'm someone with Claude stickers all over my laptop. My main app dev work is still going to Claude, but everything else is going to China even at API rates now.
One simple task: I needed an LLM to go through and clean up a few thousand page descriptions and titles in my personal search engine index, where the human web page authors had put in no effort sigh. I did a shoot out between Claude, Luna, GLM 5.3 Flash and Deepseek. Despite the high cost, Claude's descriptions were terrible, and even Opus warned me that the descriptions coming back from Haiku were "generalized, not accurate". I expected I would choose Luna because of price, and occasionally it did have wonderful descriptions (one captured emotion in a way no other model did). But in the end, the GLM 5.3 Flash descriptions were the easiest to read, they flow well while also being accurate & including necessary keywords, and being highly affordable. So it won out. It's a task that is nowhere near frontier, but a task where somehow China is better than frontier.
API rates still aren’t quite competitive with the OpenAI x20 accounts, but they are definitely getting close with deepseek 4.1 flash. I spent a few days with only 4.1 and was very impressed.
Also, frankly, as a fellow Canadian it's pretty clear that the biggest "rival" the US has right now is itself. Just passed out in the corner puking on itself shouting about all the foreigners who won't talk to it.
I'm from Europe and I hate America way more than China now. Used to be about equal but then Trump started extorting Ukraine, threatening their own allies and sending billions to Israel to help with a genocide. I think that exposed America for what it really is.
China is enabling russia way more than trump, China doesn't care too much about 'morals' either. Chinese companies have been quite important in the construction sector of the WB settlements. Even though I'm not a great fan of Trump I don't see a reason at all to prefer the chinese.
As a New Zealander, I would agree - no reason to prefer the Chinese. But Trump's America is not an attractive option either and there's no reason to prefer it. And given the choice between two ugly options, the rational choice is the cheaper one, surely.
>China is enabling russia way more than trump, China doesn't care too much about 'morals' either
The difference is China has a good reason to. China doesn't look appealing because they're more moral than anyone else, but what they have going for them is that they still behave like a rational actor. At least their behavior is intelligible in terms of their own interests. The world can deal with a long term selfish superpower but not an unhinged one
I don't think there's a person in China that has as much of a seething hatred for America's 'allies' in Europe as J.D. Vance or half of the American techbro commentariat does
China is a somewhat neutral player, supplying both Russians and Ukrainians. Their attitude and action is way less one-sided than Trump's; especially in the first year of his latest presidency.
With trump his actions being one sided you mean one sided towards ukraine? They still get lots of Intel from Americans and Americans hardly but anything from Russia. But you're right that china supplies both I wouldn't exactly call that neutral as much as just in their self interest.
Because at present the pedophile US president is making it his mission to molest my country. China, for all its faults (including espionage, which the US is also guilty of) is mostly focused on conducting trade.
Half of Canada now uses the word 'enemy' when asked for an adjective to describe America or China. We're equivalent in their eyes now because we elected Trump a second time and all that he has said and done in 2.0
It's closer to a cousin you used to be close with despite some moral failings, but who has now has a substance abuse problem and is lashing out at family and friends.
Eh, that's pretty misleading: in that poll, Canadians weren't asked for an adjective to describe America, they were given a list of three options “Ally/Neutral/Enemy”.
Yes, someone can still blow up a pipe and they look the other way. On the other hand, you can also draw parallels to themselves becoming increasingly reliant on US vs UK in the past.
I was going to start by stating the difference is “a” is unspecific, “the” implies the subject is already known.
However as I thought more about it I realized it gets a little more complicated. I’m not academically studied in English, though I am a native speaker.
For example if you say, “Our town has a castle” - that’s a specific castle but you wouldn’t say “our town has the castle.”
Why? It has to do with what the listener knows. “Definite” means the subject is known to the listener. If I’m introducing a castle into the conversation that you don’t know, I’m talking about _a_ castle. Now that it’s introduced, I can tell you about _the_ castle.
It's because the town has one out of the set of all castles.
You can also say "Our town has the castle", which... now there it gets weird. I think it means it doesn't have much, but it has that. Well I guess it's specific in that it is one of the few (specific) things it has going for it.
I would take "Our town has the castle" to imply that there's a group of towns that collectively only have one castle, and your town is the one that has it.
Compare: "Sacramento is the capital." You could say Sacramento is a capital if the reference group includes Albany, Austin, Tallahassee, etc. Or you could say Sacramento is the capital if the reference group includes Los Angeles, San Diego, San Francisco, etc.
> I think it means it doesn't have much, but it has that.
Yeah. I think what happens is that the implied group of things in which the castle belongs changes. "The town has the castle" could be understood as “there is only one castle and the town has it” (the group is that of all castles). But it’s nonsense. Your understanding seems right to me, but I am not a native English speaker. It is something along the lines of “among the things the town has, the only notable one is the castle” (it’s the only castle in the set of things the town has).
What is tricky is that the group in which the castle is unique is implicit.
FWIW all that is barely different from German (I'm a German native speaker). These implied groups seem like a nightmare if you have to learn how they work. Since I "just know", I've never had to think about it much.
Native English speakers, please correct me if I'm wrong about the "the castle" example.
"Our town has the castle". Is reasonable for meaning that it has the best/oldest/most historical/largest one. Or in general that the specific castle that has something which separates it from the rest.
"Our town has the castle" would mean that of (all the gin joints in) all the towns in all the world, only yours has a castle. Also not academically qualified, also a native speaker.
You use either “a” or “the” when you are specifying a class of things. “Our town has a castle” could be any castle. “Our town has the castle” would be a bit more rare and sounds weird as a general statement, but is valid in certain circumstances. “The castle” here means that the castle is something, or connected to something we have already been discussing and I’m referring to it, so the “our town” must be the new information. Like maybe we are discussing what we should do when you next visit, and I suggest seeing some particular famous tapestries. You say “wait, isn’t that castle [in which the tapestry is displayed] in Edinburgh?” And I reply “no, our town has the castle.”
You don’t use any article at all when what you’re talking about is unique. For example, “our town has Neuschwanstein Castle.”
Edit: or there is something which follows, e.g. “our town has the castle where they filmed Braveheart.”
Fair. I was envisioning the statement "our town has the castle" outside of any other qualifying context about the castle. Like a couple of schoolchildren bragging about their towns. "Our town has a racetrack." "Well our town has a castle!"
I couldn't disagree more. The frontier labs are already working with the military to kill people. What you're going to get (even more than we already have) is a two tier system where those with weapons and a proven desire to use them will have the best AI and us regular citizens will have the hobbled AI. A complete reverse of what would keep us safe.
There is no "right". Alignment is shorthand for ideological alignment. There's always people judging whether an answer was right and the answer for that will be different in Silicon Valley than it'll be in China or in Europe.
Consider for example the question "What caused the French Revolution?" Many different answers could be given, all technically correct. What gets emphasized is where the ideology lives.
Also there's no "alignment" for cybersec. The line between blue and red is really a perspective issue. If you go over the "tokenkiddie" problem, when you get to the real security issues, your model either detects them and you can secure your systems, or it refuses and then attackers will use abliterated models to find them.
It does indeed mean ideological alignment. But we don't get to leave the answer blank. They have to pick an ideology to put in there, and whatever they pick will have huge consequences.
My point is maybe we don't let some of the world's richest people with some very... interesting ideas pick which ideology we digitalize? Maybe we figure out a way to give people a say in this?
Probably has something to do with the Alignment Lead at Anthropic agreeing with him, saying there’s a >10% chance that AI kills humanity in the next decade.
It would be nice if one of these models would produce a novel theory or advance the field in a positive direction.
Most (all?) of the big discoveries have been counterexamples, which is just sort of a systematic tearing down human ingenuity. I know that counterexamples are an important part of progress and discovery, but it just feels bad to me.
But I'm not a mathematician, maybe I'm totally misreading the vibe.
But yeah, Terry Tao considered this exact situation in advance and is on record that this exact outcome (rushing to priority before an explanation) would be the worst possible result. https://mathstodon.xyz/@tao/117207849921390904
We will have to see whether any other millennium problems fall. I guess that in a year the scope of AI math will be much clearer, for now it's still a bunch of incidents of unclear pattern.
Since he wrote this five days ago, when these efforts were already underway, if he was not Terence Tao I would suspect he had inside access. But since he said he did not and was speaking hypothetically, and he seems to be an honest person as far as I can judge, I guess some people are just on another level.
this isn't really true anymore. First, a number of the big results are constructions, not counterexamples. For example the existence of a non-sofic group. It was widely believed that non-sofic groups existed (so it wasn't a "counterexample" to a widely believed conjecture), but no constructions were known.
There are other examples though. For example, NP hardness of n^{1/400}-approx CVP. Like any NP hardness proof, this shows you can faithfully encode a hard problem (3SAT here iirc) in terms of another candidate hard problem. Not really a counterexample at all.
reply