Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That’s very unlikely, tokenization is really simple and usually quite fast (scales with input size). Unexpected slowness with a larger context window might point to a non-existent or unoptimized KV cache.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: