I tested 3 local LLMs on my actual work — and each model won at something different

I stopped defaulting to one model