Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
lowbloodsugar
6 days ago
|
parent
|
context
|
favorite
| on:
Qwen 3.8 27B available on Cerebras at 1500 tokens/...
Ok, I need to try that. I'm getting 45tok/s with vLLM on my 6000. >600tok/s concurrent, but 45tok/s single request.
help
pllbnk
6 days ago
[–]
Even without ninfer I would get over 80 on LM studio with default settings, so it should be noticeably more on 6000. You might want to try different a different inference engine or settings.
reply
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: