Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It also depends on how many tokens it needs to burn through to accomplish something.

At this point, I always look at things like Artificial Analysis' total cost to run their tests. It'll take into consideration the cost of tokens, how many tokens it burns through, and how effectively it uses caching (and the price of that caching).

If a model "costs the same" but its reasoning ends up going through a ton more tokens, it doesn't really cost the same in real world usage.



Precisely. GLM 5.2 Thinking is pretty damn good - but it regularly does something nonsensical. Or even spits out what looks like a fragment of its memory cache. Or returns a bunch of Chinese.

I find myself having to resubmit a query very often...so it being a third of the cost of other AIs isn't really relevant.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: