Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Parent is likely thinking of sparse attention which allows a significantly longer context to fit in memory


My comment was harsher than it needed to be and I'm sorry, I think I should have gotten my point across in a better way.

With that out of the way, parent was wondering why compaction is necessary arguing that "context window is not some physical barrier but rather the attention just getting saturated". We're trying to explain that 3+2=2+3 and you people are sitting in the back going "well, actually, not all groups are abelian".




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: