Most of the advice on observing AI systems is being invented from scratch right now, and a good deal of it was solved a decade ago at hyperscale. I have spent years running observability on large distributed systems, and this talk brings that hard-won discipline to AI and agentic workloads. Classic observability assumes deterministic behaviour, clear pass or fail, and an error when something breaks. AI breaks all three, and the worst failures are silent, with quality degrading while the dashboard stays green. I cover which distributed-systems patterns transfer directly, what has to be rethought (new signals like tokens, cost and tool-call success, and tracing across agent and tool boundaries), and evaluation as the new form of testing, turning eval scores into SLIs and alerting on quality drift rather than just errors.
This talk has been presented at AI Coding Summit NYC, check out the latest edition of this Tech Conference.


















