While the adoption of generative AI continues to advance, concerns are rising that the increasing volume of token consumption associated with model usage may strain corporate AI budgets. According to research by Gartner, many cases have been reported where unexpected cost increases become a challenge in budget management surrounding AI implementation.

From the perspective of AI cost management, there is a movement to apply FinOps (a methodology for optimizing the costs of IT infrastructure, such as the cloud) to AI operations. Additionally, benchmarks to evaluate the balance between model performance and cost are being increasingly emphasized. According to "lang-token-bench" conducted by Deep Insider, it has been shown that there are significant differences in token efficiency depending on the model. For example, even among the latest models such as GPT-5.5 and Claude Opus 4.7, there are disparities in the number of tokens and the cost required to obtain equivalent output.

In development environments, the balance between the quality of AI-generated code and token consumption has become a subject of debate for specific tasks such as code generation. There is a greater need than ever for an ROI (Return on Investment) perspective—specifically, whether AI-driven coding can demonstrate efficiency enough to reduce labor costs, or whether token consumption might offset those economic benefits.


Source: