A benchmark has to have a gradient to be effective.
- A benchmark has to have a gradient to be effective.
- If all of the test runs are terrible or all too excellent, the eval doesn't give you a gradient.
- Getting an eval that is calibrated to the dynamic range of the current capabilities of the system under test is hard.
- Once everyone scores 100%, the benchmark is fully saturated, and you need a new one.
- Auto-improving benchmarks can be self-steering north-stars.