AI coding assistants have quietly become a real line item in engineering budgets - and most teams have no strategy for it, just a bill. The uncomfortable truth: you don't pay for tokens, you pay for tokens that produce no value.
This talk is a practical playbook built on two levers. First, model routing: matching model tiers to task phases - frontier models for planning and review, mid-tier for implementation, the cheapest for mechanical work. Second, context management: why agents that explore large codebases burn premium tokens on searching instead of thinking, and the architectural pattern that fixes it - a code intelligence layer where an expensive model asks questions and a cheap, locally-run system that already knows the codebase answers them.
This talk has been presented at AI Coding Summit NYC, check out the latest edition of this Tech Conference.




















