×
Evaluate Your Team's AI Readiness

Take our official assessment and explore training paths to build secure, compliant AI workflows.

Access Portal

Business

I Burned Two Billion Tokens for $125. That Number Is a Business Risk.

5 minute read

It rained all weekend, so I built things.

Three projects. About 20 hours across two days. Up to 10 simultaneous agents running in Codex, kicked off as separate work trees so they wouldn’t collide with each other. When I looked at the usage numbers afterward, I had burned two billion tokens.

My bill was roughly $125, because I was on a consumer plan and OpenAI kept resetting my usage. By Saturday night I was down to 13% and figured that was the end of the road. Woke up Sunday and I was back at 83%. Went out for a while, came back to 10 or 11%, and it bounced to 100%. Ran out again Monday and got another reset.

So I looked up what that weekend would have actually cost. It’s complicated input tokens price differently than output tokens but I ran almost entirely on GPT 5.6 high, which isn’t the ultra model. It’s the recommended workhorse. Somewhere north of $8,000 to $10,000 at retail.

The Gap Isn’t a Curiosity. It’s Your Business Case.

Here’s the part that I don’t think a lot of mid-market and smaller organizations understand. When you have a $200-a-month Max plan or a comparable GPT plan, you get a huge pool of included tokens. There are guardrails hour limits on certain models, usage caps but the economics are wildly favorable.

If you’re on an enterprise plan, you fall out of that. Your team building in Codex is paying an entirely different rate. It can be a negotiated rate, which makes it genuinely hard to even pin the number down. But it isn’t $125 for two billion tokens.

And the subsidy is visible. Anthropic has been running a 50% discount on tokens for its users, and that’s scheduled to come off in mid-September. That’s one shoe. There will be others.

So the honest gut check: would I have written a $10,000 check to build what I built over that weekend? Some of it is sales fronting, some of it is operations, all of it will be valuable to my organization. And still probably not. That’s a hard bet for an entrepreneur.

Which is exactly why I spent the weekend coding. Not discipline. FOMO. I want the foundational pieces built while the cost is still in alignment, so that if pricing moves I’m tweaking an existing system rather than deciding whether to start one.

Matt Raised the Better Version of the Question

On the episode, Matt pointed out that I was worried about the wrong number.

His team built something internally they think of as an AI brain. It captures every call, every Slack message, every email tied to a client project. It segments them by project. It flags when something’s going off the rails, and it surfaces topics worth covering in a quarterly business review based on that quarter’s conversations. It runs pretty affordably right now.

His point: the build cost is the risk I was talking about. But if your entire business is built around a system like that, the exposure isn’t what it cost to make. It’s what happens when running it gets three times more expensive. Build cost is a one-time decision you can defer. Run cost is a line item that scales with how essential the thing has become.

That’s the sequence that should worry people. Cheap tokens make the pilot look great. The pilot becomes production-critical. Then the pricing normalizes, and you’re not evaluating an experiment anymore you’re renegotiating the cost of something your operations depend on.

Why I’m Not Fully Convinced Prices Will Spike

There’s a real counterweight, and it’s worth understanding if you’re planning around this.

There is no switching cost. If you’re building the right way, your code lives in GitHub. Moving from Claude Code to Codex meant connecting the repository on the other side and asking what we’re doing next. The main mechanical difference was renaming my CLAUDE.md files to AGENTS.md. That’s it. Grok has gotten meaningfully better at coding. Kimi is there. Even if you’re not coding, you can ask a model to summarize everything it knows about you and import that context somewhere else in ten minutes.

As long as that transferability holds and there are credible alternatives, these labs are in a prisoner’s dilemma. The resets I got all weekend were competitive warfare, not generosity. Getting away with big price increases in that environment is hard.

The way it stops being hard is regulatory. Shut off access to the Chinese models and there are national security arguments for that and the field narrows fast.

So I’m not predicting a cliff. I’m saying the number is unknowable enough that you should know what your system costs to run today, and what it would cost at two or three times that, before it becomes something you can’t turn off. If you don’t know that number, you don’t know what you’ve built.

Book a conversation if that’s a number you’d rather work out with someone.

×
AI Solutions Graphic
Free Resource
Get Started Now