Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Humanity's Last Exam (w/ tools)

This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.

 help



Do we even know what the true ceiling of hle is? I'm pretty sure some of their public sample questions are wrong (ambiguous/nonsensical with most logical interpretation trivial)



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: