I've been running some of the biggest open-weight LLMs for free on Nvidia's cloud

You can use big models for free, though there aren't any promises on speed.