After self-hosting LLMs for a year, I realized that models are not the real bottleneck

I stopped upgrading models and fixed my prompting instead.