Getting a coding agent to produce a diff is easy. Deciding when that diff deserves a pull request is the engineering problem.
This talk follows one issue through a real coding-agent workflow. We inspect the trace at the first meaningful divergence, then add only the controls that the observed failure justifies: agent-legible context, a checkable completion contract, independent verification, and action boundaries. The comparison becomes a practical framework for sizing the harness to the task: when instructions and tests are enough, when a separate evaluator or sandbox earns its cost, and when a human checkpoint should remain. Attendees will leave with a repeatable method for turning agent failures into systematic improvements instead of longer prompts and more manual review.
This talk has been presented at AI Coding Summit NYC, check out the latest edition of this Tech Conference.





















