A Beginner's Guide to AI

The Next AI Crisis Won’t Be Hallucinations. It Will Be Costs

9 min
Jul 18, 2026about 2 months ago
Listen to Episode
Summary

The episode explores the emerging crisis of rising AI token costs for businesses using large language models and AI agents. The host shares a real-world example from his university startup where a single scientist's project generated $180 in API costs in one week, highlighting the need for businesses to implement cost monitoring, caps, and governance around AI token usage before it spirals out of control.

Insights
  • Token costs are becoming a critical business concern as AI agents and complex workflows consume exponentially more tokens than simple chatbot interactions
  • Businesses need to shift from fixed-price AI models to usage-based monitoring systems with clear caps and dashboards to prevent runaway costs
  • AI implementation requires balancing adoption with cost control—too little usage suggests poor adoption, but uncapped usage can create financial liability
  • Without proper governance, a single user or agent loop can generate thousands of dollars in unexpected costs within days
  • Business leaders must implement measurement systems and usage dashboards before deploying AI agents to teams with API access
Trends
Rising LLM API costs creating new financial risk category for enterprisesShift from fixed-price to usage-based AI pricing models forcing business model recalculationsAI agents and multi-step workflows driving significantly higher token consumption than traditional chatbot usageEmergence of cost governance and FinOps practices for AI spendingPotential for uncontrolled AI agent loops to generate massive unexpected costsGrowing need for AI cost monitoring dashboards and usage analytics toolsToken pricing becoming a key factor in AI adoption ROI calculations
Topics
AI token cost managementLLM API pricing and billingAI agent cost governanceBusiness model pricing strategies for AI servicesCost monitoring dashboards for AI usageAI spending caps and controlsQualitative research with AI agentsDocumentary method researchAI adoption ROI calculationUncontrolled AI agent loopsFixed-price vs usage-based AI pricingInternal AI cost allocationAI financial risk management
Companies
OpenAI
Referenced as example of organization where employees might generate $150,000+ in token costs
Anthropic
Referenced alongside OpenAI as organization where high token cost scenarios are possible
University of the Armed Forces Munich
Host's university startup developing AI tools for qualitative research and interview analysis
People
Dietmar
Host sharing personal experience managing AI costs at his university startup project
Quotes
"What happens if everybody that has access to the app pays 24 euros a month produces over one week and 180 dollars in costs"
Dietmar~5:30
"The cost of LLMs rise. And you as a business leader, you have to make a decision and you have to see how you can cap this whole thing because it can get out of control."
Dietmar~8:00
"If you don't program them right, then they might run into a loop and do things over and over again until someone stops them and each loop costs tokens."
Dietmar~15:30
"There's those cases with people using up $150,000 in token. This is more like the people working at OpenAI or Anthropic, but that is possible."
Dietmar~16:00
"We have to be in between not using AI using AI too much we have to see how this develops but keep an eye on this it really important can really get out of control"
Dietmar~13:00
Full Transcript
Token crisis, spending too much for tokens. It's the first time it hit us, so I wanted to make a quick episode on token prices. Welcome to another episode of the Beginner's Guide to AI. It's my Friday night opinion episode again. Don't forget to go to beginnersguide.nl and get the newsletter. Also go to AI for the 99%, which is my podcast for small and medium enterprises, where I give you tips and tricks about how to be successful with AI. Just go to your podcast app and type AI for the 99%. But before I talk too much, let's jump right into the episode. I work in this university project. It's a startup that comes from the University of the Armed Forces in Munich. And I am the head of this firm. And I obviously am also responsible for checking bills that come in. And we have some people that are scientists that develop great tools for analyzing interviews, some talks, and whatever people say in a qualitative way. AI can help you a lot and AI agents can help you a lot. You can not only put in your, let's say, interview with the person you had, but you also can reference all the interviews you had. You upload them in an ARAG. And then you can also develop a whole process of how to go through this. One of the scientists is just developing this. It's his PhD project and it's great. He has a series of, I don't know, 70 prompts there, prompt files, or he puts in Cloud Cowork, which is connected to our app, our program. And yeah then there was this bill thing We had entropic bills coming every day with it was like at a certain point he changed to higher amounts of bills and in the end it okay he spent 180 dollars in one project one thing and we were like okay that is is it worth it yes it's for science it's worth that we can finance this and it's also establishing a new way of doing this kind of research it's called documentary method and it's good to have ai in it or a perspective on ai in this and will be interesting how the people react but then i obviously thought as head of this firm what happens if everybody that has access to the app pays 24 euros a month produces over one week and 180 dollars in costs and yeah obviously there are things he had to try out but still the cost the what happens what we realize now internally is we have a fixed price before we had a model where the people had a fixed price and had to buy credits. That's a term for tokens. We changed the model because we thought, okay, yeah, let's go to make a fixed price. It's much easier for the people for us. And now it changes that the models get more and more expensive. So we have to think, what do we do? Do we raise the prices for the fixed price? We have a mixed calculation, obviously, because some people don't use many tokens. Others use a lot. Does it work? so in the back end where we make our calculations things get much more complicated with those rising token prices and this is actually a thing that i take and this is actually a thing i take from it and this is what they want to want to mirror this for you to you because if you are in a leading position and you have people working for you and up until now it was okay they used some tokens they basically did some jgbt stuff in a style of they typed something in a field And yeah that was the tokens you paid And now they use agents They use lots of information that gets sent to the model and lots of information that gets back. This is a high token count. And then like iterations of this. And so tokens, the price of token rises and the price of token, the number of token rises. and so the price or the cost of the LLMs rise. And you as a business leader, you have to make a decision and you have to see how you can cap this whole thing because it can get out of control. With us, it's like this guy is responsible. Actually, he's not just a scientist. He's one of the owners of the firm. So if he spends too much, it's also his own money. He has an incentive to calculate and to control the thing. But what if the people don't care for it and they send out an agent that does stuff and comes back with $1,500 token costs the next day and the firm has to pay it. So think about how you can cap those. But first of all, think about how you can measure it. Do you have a measure system? Do you have a dashboard where you can see how many tokens get used? Who uses those token switch models? and if you are still in those normal contracts where everything is included, it's okay. So it's only important for the people who actually use APIs and have access to more tokens. If you have still a normal fixed price model where you get a cap on those expensive models, it's not a problem. But if you don't have this cap, be careful in the sense of see what your people do. also if the people don't do anything you also have to question there must be something wrong they don't use what you pay for and they don't use ai and that's also a disadvantage so that's a that's a yeah it's it's we have to we have to be in between not using ai using ai too much we have to see how this develops but keep an eye on this it really important can really get out of control so introduce caps once you know that it could be a problem for your firm and was just a small example and like i said it's at the moment not a big problem for us we introduced an agent people don't use it there's still not much that they can do without knowing how it works so we are not yet there that it gets dangerous for our business model as you probably don't offer ai services it's for you more the internal side it's like looking at who uses how much of which model and is this a danger in a financial sense so keep an eye on it and see that you keep that under control i don't want to tell you i don't want to scare you but there's those cases with people using up $150,000 in token. This is more like the people working at OpenAI or Anthropic, but that is possible. So be really careful if there's no cap and people spin up 10 agents that do enormously much work, not always the right way of work, not much. I mean, if you don't program them right, then they might run into a loop and do things over and over again until someone stops them and each loop costs tokens. So be really careful about that. And that's it for today. Just a cautious episode. Be careful of the token usage in your firm. Yeah, I hope you liked this personal insight. Thank you for staying at the end of the episode. So don't forget, as always, to go to beginnersguide.nl and get the newsletter for all the new episodes. It's Dietmar from Argo Berlin signing off.