"Which model should we use?" is the wrong question. The right one is "which model should handle this call?" โ and the answer changes with the workload, the latency budget, and what you're willing to pay to be wrong less often.
These posts are about making that call deliberately: small-first escalation, benchmark skepticism, and treating model choice as a routing decision instead of a one-time procurement.