Semantic trace summaries
Can summaries of state-changing tool calls help predict whether an agent succeeds? My study of public τ-bench trajectories found no reliable overall gain when both the task and agent were held out.
Read the study
Software engineer working on agent evaluation and systems
I build evaluation and systems infrastructure for tool-using agents. At Google Cloud AI, I work on agent framework efficiency and quality, with a focus on making systems faster, more reliable, and less expensive.
Can summaries of state-changing tool calls help predict whether an agent succeeds? My study of public τ-bench trajectories found no reliable overall gain when both the task and agent were held out.
Read the study
I founded Strust to measure code-conversion performance and conformance in legacy modernization.
Read the project case studyA request can accumulate queueing delay that average CPU utilization does not reveal, so measure its path before increasing the worker count.

I'm especially interested in agent evaluation, ML inference, GPU systems, compilers, and runtimes. I also publish reproducible studies on reliable software.
Outside of work I play cello, backpack, write, and release jazz and blues recordings as Copperbelt. I also share photographs from my travels.