Hey there, yep we found that on Apple devices specifically running on CPU is fast enough that Metal support is not needed. Thanks for flagging this though, and if usecases that would benefit from Metal support come up we will be adding it to the binaries.
I tried this today for labelling - and for that task it was very bad MNLI was better - so you are going to need to match the use case for this pretty exactly. (at 29MB params one would expect that!) I'm obviously not saying labelling is a good use case :-) just adding a data point.
Jev has put the cat amongst the pigeons so suddenly everyone is looking at classifiers and encoder only models again.
My ideal model would be a general purpose LLM API that can answer classification questions and as it does so distils to an encoder only model so that the more classifications I do the cheaper it gets (i.e. the more it offloads to the classifier). If anyone ever wants to do this as a service do let me know, because it's just another piece of code to manage in each new project that needs classification.
Also a model that could do this internally would be nice :-)
Hey! Yeah I think for labelling the model would need to have much better world knowledge than its current size allows. Jev really is a very good model, I think it has a very strong place in the upcoming tech stacks. Really good suggestion to make a continuously distilled model, we are going to have to look into that one :)
Good luck with this model/product, in the excitement of LLMs people seem to forget applicability. I very much like to see innovation in this space, so well done!
Please have a 'readable version' option so I don't have to exhaust myself parsing the sites layout. I get that it's unique but most of us just want to work out what you're offering in 5-10 seconds of our time.
I'd bet that no one, not even the site's [human] creators, has ever read that homepage end to end. At best, it might have been handed over to a swarm of reviewer agents.
Why on earth would we want such a lock-in at this stage when there is no clear winner. This is an area in which I would encourage everyone to build their own (using OSS) on top of existing cloud infrastructure.
There are no shortage of good harnesses out there, and you can re-use coding ones like pi, opencode, dsh even claude and codex. For most use-cases that is already the loop you need. If you're looking for something less like a coding agent then Vercel's Eve is okay too.
reply