Falling AI token prices drive higher enterprise LLM bills

Token costs plunged in 2024, yet enterprise LLM spending tripled and 93% of firms exceeded token budgets.

Enterprise bills for large language models rose even as per-token costs fell sharply in 2024. Token prices dropped from about $20 per million to roughly $0.07, and industry data show enterprise LLM spend tripled over the following year. A survey of firms found 93% exceeded their token budgets and about 20% cut back AI use because costs spiked.

Tokens are the units models use to process text and other inputs. Providers meter both input tokens (prompts, documents, conversation history) and output tokens (model responses), and they often charge separately for each. Some pricing examples put input-token rates at roughly $0.20 per million for mid-tier models, while other models are priced at about $2 per million input tokens and $10 per million output tokens.

Volume adds up quickly when firms use agents and automation across many workflows. During testing at Advyzon Investment Management, a developer reported a single planning task consumed about 50,000 tokens, according to Kevin Hughes, president of financial planning at Advyzon. Firms that add AI to meeting prep, research and client-facing workflows saw routine tasks drive high token counts. About one in five firms in the survey said they scaled back AI use after costs rose.

Some investors and managers compare the pattern to the early cloud era, when cheaper unit prices led to higher total bills as more workloads moved online. Ray Wu, founding managing partner at Alumni Ventures, noted that falling unit costs shifted the challenge to cost management once usage scaled.

Measuring return on token spend is uneven across firms. Todd Ahlsten, chief investment officer at Parnassus Investments, pointed to wide variation in how tokens are used: some support research, others run trading ideas, and some power HR and payroll. That variation makes it hard for finance teams to set a consistent value per token.

Firms seeking to control spending are matching model choice to task, auditing integrations before adding AI layers, and defining whether budgets are fixed or variable before rollout. Hughes used an arcade-quotes analogy: “You put your quarters in it, and you keep playing until you’re officially done,” describing how teams may keep triggering models without stopping.

Some cost controls include setting model tiers by workflow, tracking token use by application, and measuring outcomes such as time saved or client impact rather than raw token counts. Finance teams are treating token spend as an operational line item to monitor and manage as AI use expands across firms.

Articles by this author