Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Expert offloading significantly helps with the VRAM capacity.

Most MoE architectures have a few experts that are always running; this, the router, KV, and whatever else you have space for can stay in fast VRAM; and the remaining experts can be offloaded.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: