Most of my work is failure handling.
Latency, interrupted sentences, tool calls that half-fail, a model that answers the wrong question with total confidence. At Scale AI I shipped a multi-agent system that works through up to 20K denied claims a week, and owned the eval framework that moved audit accuracy from 64% to 88%. Almost all of that came out of the failure cases.









