this post was submitted on 03 Jul 2026
606 points (99.2% liked)
Technology
87279 readers
3158 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related news or articles.
- Be excellent to each other!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
- Check for duplicates before posting, duplicates may be removed
- Accounts 7 days and younger will have their posts automatically removed.
Approved Bots
founded 3 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Something I’ve been noticing recently is that while the cost per token on specific models hasn’t gone up, the provided interfaces for using those models are starting to chew up significantly larger numbers of tokens for the same tasks that used fewer tokens with older versions of the interface software just a few months ago. Likely the interfaces are applying more expensive guardrail prompts and charging the end user for those tokens — but the end result is that it costs 4x as much to get the same work done.
"Tokens" are just made up.
These "tokens" that are used to "measure" how much you use, they are not a real dimension that can be measured. Just an artificial counter that goes up when they decide that it should go up.
They can change the "size" of a "token" every day, and every second, and every microsecond....
It's not like that. Tokens are an inherent computational property of how a model calculates the probabilities and such to generate text.
Having said that, what a token means in terms of computation varies wildly between models and is not directly comparable. So attributing a money value to tokens in general, independently of the model, is weird by nature.
And even within a model, the number of tokens needed to generate a response is very variable too, depending of the model itself and the parameters with which it has been configured (thinking mode, temperature, etc.).
So yeah, companies can pretty much set any price they want and there's not much anyone can do about it.
It does make sense for the provider as those for a specific model provide a good measure for computational effort, for that doecific model. That doesn't mean that token rate comparison between models give you a good picture.