Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Ah ok, I think we made different assumptions about whether the model was specific to the particular dataset so each one would need a new model — a dictionary is specific to the particular dataset being compressed, right? I was thinking the LLM would be a general-purpose text compression model.


Not particularly. You could make a dictionary from "the English web", with common character sequences found on those sites you use as input.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: