Matching Complexity to Capability
This past month the AI conversation pivoted hard from capability to cost. Tokens are burning through budgets, and the believers are blinking. Uber blew its entire 2026 AI coding budget in four months. Microsoft started canceling internal licenses for an AI coding tool because the bills outran the value shipped.
The numbers explain why. The same task can cost 25 dollars or 180 dollars per million tokens, depending only on which model you point it at. A routing study showed you can send each request to the cheapest model that can handle it, cut your bill by up to 85%, and keep 95% of the quality. The logic is old. Route simple work to cheap models. Reserve frontier compute for the hard 20%.
The edge now is an orchestration layer, a gateway or agentic layer that matches cost to complexity on every single request.
Sitting with this, my thought was simple. This optimization is not for AI alone.
Senior, expensive judgment gets spent on work a junior could do. Juniors get handed calls they have not earned yet. In my Ask GeKo conversations, employees routinely tell me 20 to 30% of their time goes to work that should sit elsewhere. The misrouting stays invisible because the old org was never built to be modular. AI now prices every single request. Your org chart still prices bodies not tasks.
The coming of AI should be a big enabler for change. However, Deloitte’s 2026 survey found that when companies respond to AI, 53% train their people and only about a third redesign who actually does what. We are upgrading the worker and ignoring the routing.
The AI world will keep optimizing token routing. Human misrouting stays hidden in the swamp of legacy culture, until leaders name the mismatch and act. Take the boxes on the org chart, break them into tasks, and reallocate each to where it belongs.
Originally published on LinkedIn