For evaluating subs you have to put the money and try by yourself to get your own measurement, since the subs quotas and new models are always changing and overall info is unreliable.
20$ of Codex/Google/Chinese gives you some fair amount of usage to test them, and Opencode Go for 10$ lets you try a good amount of models with good quota. I don't use Openrouter because it gets more expensive than the subs, but for testing, swapping and being completely independent, a proxy service is the best solution.
About benchmarks, I usually agree with DeepSWE. Looking at usage rankings of models in Openrouter is also a good signal.
For changing models locally I just use pi (used also opencode in the past) or if the sub does not allow login in external harnesses I just use whatever cli they have, last year there were differences but today they are all good enough for my needs and have basically the same features.
I would start by using AI to port the tangled mess into as much a tidy IaC as possible, with git backed declarative config tools, to prepare the safe ground where agents have context, can get loose and easily review, lint/check or test in throwaway VM before applying it.
A fleet of Openclaws (or equivalent) with control of machines for 100% hands off management seems like a russian roulette to me as of today, but a 90% handoff to AI for reviewing current config of X in our system, reading docs and configuring a deterministic tool is achievable and somewhat safe as you can easily review and trace what has been done.
I had the project already starred, I lurk around the sandbox space from time to time, and the project it's a few months old and was already discoverable and in a usable state before today.
reply