
Right — and even if your machine doesn't have 16 GB of memory, it's now very easy to find compute providers that specialise in open models like Gemma — for example I just switched to Baseten and it's working really well. There are many: Wafer, Friendli, and so on. Their thing is they're not trying to lock you in; the only money they want to make is computing more tokens for you per unit time. And they know OpenRouter is benchmarking up front, so if their compute offering isn't as good as someone else's, people switch immediately.