intermediate

The Benchmark That Only Wins for Ten Seconds

Read the counters before the options. Nothing here is labelled with the answer.

The report

Our new vectorized encoder benchmarks 35% faster than the old one. In production it is about 4% faster, and on the busiest nodes it is slower. The benchmark runs the same input on the same instance type. We have run it dozens of times and it is consistent.

CountersSIMULATED
benchmark duration8 s per runEach benchmark run completes in a few seconds.
core frequency, first 5 s of benchmark≈ 4.6 GHzThe core runs near its maximum clock at the start of a run.
core frequency, after 60 s sustained load≈ 2.9 GHzUnder continuous load the clock settles substantially lower.
frequency during wide-vector sections≈ 2.6 GHzThe clock is lower again while wide vector instructions are executing.
instructions retirednew version 0.62× the oldThe new version retires substantially fewer instructions for the same work.
package powerat the configured limit under sustained loadThe processor is drawing as much power as it is permitted to draw.
benchmark: cores active1 of 32The benchmark exercises a single core.
production: cores active28–32 of 32Production runs the workload on nearly every core at once.
What is the hardware doing?