Can it run?
Fit the model and runtime on the target hardware.Start from the real task
The benchmark only matters if it represents the interaction, context size and sustained behavior the product actually needs.
Record the exact candidate
Model family, parameter scale and quantization remain part of the evidence rather than disappearing behind one score.
Keep the execution layer explicit
Runtime and backend choices can materially change latency, throughput and resource behavior.
Measure where the decision lives
Representative hardware is mandatory when memory, thermals or battery determine whether local execution is viable.
Make the run reproducible
Context, concurrency and other relevant settings travel with the result so comparisons remain interpretable.