Browser-based AI inference is stuck between two bad options: follow a standard and get portability at the cost of speed, or hand-tune your own kernels and burn weeks of deep GPU expertise per model. For Transformers.js, we tried something that probably shouldn't have worked as well as it did: handing that hand-tuning work to an AI coding agent, with way less supervision than seemed responsible. It found multiple-x speedups across LLMs, audio transcription and text-embedding models running in the browser, on hardware we already had.
This talk is a case study, not a pitch. I'll walk through the loop that turned autonomous kernel optimization from chaotic into trustworthy, the one instruction that changed the agent's behavior more than any amount of prompt engineering, and the counterintuitive lessons from real experiments, including why our biggest wins did not come from clever algorithms and why you should never let your agent gaslight you into thinking it's done.
This talk has been presented at AI Coding Summit Berlin, check out the latest edition of this Tech Conference.
















