I stopped running the biggest local LLM that could fit, and a 2B model handles 90% of what I need

Smaller doesn't mean lesser