Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem some form of continual learning (model weights update like dreaming)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: