Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Lag is a real killer for LLMs.

I'm curious to hear more about this. My experience has been that inference speeds are the #1 cause of delay by orders of magnitude, and I'd assume those won't go down substantially on edge devices because the cloud will be getting faster at approximately the same rate.

Have people outside the US benchmarked OpenAI's response times and found network lag to be a substantial contributor to slowness?



of course not, it's just text. people here are just spitballing.

groq is way better for inference speeds btw




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: