Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Ok, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request.
 help



Even without ninfer I would get over 80 on LM studio with default settings, so it should be noticeably more on 6000. You might want to try different a different inference engine or settings.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: