Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I run Gemma 4 26B-A4B with 256k context (maximum) on Radeon 9070XT 16GB VRAM + 64GB RAM with partial GPU offload (with recommended LMStudio settings) at very reasonable 35 tokens per second, this model is similiar in size so I expect similar performance.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: