Running out of your subscription limit tokens is a real issue that I have been dealing with more frequently lately, so let's chat about current strategies.
The first time Fable came out, I noticed that I started running out of tokens more frequently with Opus as well. Before that, I ran my regular work and really any question I had about the projects and codebases on Opus 4.8 max thinking effort, rarely hitting limits. My guess is that this update bumped pricing in general, but it could also be that my usage increased. Anyway, since then I have to find ways to deal with the situation.
First of all, during the workday, I'm constantly watching the codexBar:

Currently, I'm on Claude team plan, which I was told is higher in limits than the $100 5x subscription but lower than the $200 20x one. It's supposed to be somewhere in the middle.
I have to be strategic about model and effort levels. I think Opus 4.8 max level is great for any serious codebase question or coding task, but I have now Opus 4.8 high effort mode as my default and manually switch it to higher levels where I want it to do better work. For very complex multi-system multi-step tasks I turn on ULTRACODE or Fable if the session window allows it. I also have some extra usage enabled by the team plan that I would drain from if I run into any limits. It turns out, you can't turn off tapping into that extra token money, even if you don't trust Fable to do a good job with your task and you know it will eat up all your tokens with a particular task.
When running into Claude limits, I usually start using my ChatGPT Plus subscription, which currently is enough to get me through.
I'm using Conductor, so switching models or effort levels is just a shortcut hit away.
So to sum it up, to stay on the loop in the current token economy and still use frontier models and thinking levels. You will need to choose wisely between them for different tasks until tokens are solved.
("tokens are solved" - I had a buddy recently telling me that it's funny to him people are freaking out about tokens and the environment etc, because soon we'll have token stations orbiting the Earth, generating tokens through the sun ā š)
[AIN#27] AI Nuggets Week#27
šŗ YouTube links
Our Agent Negotiated a Vendor Renewal, Became a CFO and a Better SDR .. but has too many guardrails (SaaStr AI) - "21 agents, 3 humans and a dog", and here I am trying to build mini harnesses and babysit one agent at a time to crush these gigantic codebases.
LLM Observability, Evaluation, Experimentation Platform (Dat Ngo, Arize) - Promo talk, but the premise of having your agent usage visualized and traced seemed interesting.
š¦ X links
Fable works way better than everything (Michael Koper, @michaelkoper) - "I don't know what it is, but Fable works so well for me. Way better than all the other models combined." - I don't know what to think, I just let Fable do a review on a big PR and it cost me the entire session window limit plus extra usage š I feel irresponsible using it so far more than anything else, but it does generate some cool stuff on lighter tasks. Feels like Fable works similar to Opus 4.8 in ULTRACODE sometimes which in turn "feels like" it consumes less tokens on average.
š From the wild
Turning code into images to cut model costs
At this point, it's whatever works for ya š š«
An AI agent topped our ClickFunnels commit leaderboard
The fun twist: most of the human runner-up's commits were agent-written too. The leaderboard is agents all the way down now.
Claude Tag: Anthropic's ambient agent for the masses
Claude goes OpenClaw?
Fable 5 is back, and so is the pricing tier math
A few days of half-cost Sonnet 5 bliss, then the top tier returns at twice Opus prices. Model routing rules everywhere got rewritten twice in one week.
Claude Sonnet 5
The important corner of the Internet seems unimpressed:
The verdict? The community is largely unimpressed and confused about Sonnet 5's value proposition for high-effort tasks. The consensus is to stick with Opus 4.8 for anything important and only consider Sonnet 5 for low/medium-effort, cost-sensitive, or speed-dependent jobs. - reddit
Users say it can complete complex tasks, but often generates convoluted code needing senior refactoring while using more tokens than expected, reducing its apparent cost efficiency. - hn
I haven't tried it yet since there was no crazy need to save tokens in my lighter weight workflows - maxing out on Opus and co, still.
Member discussion: