July 30, 2026
We lifted task completion from 37% to 50% on τ-Banking, the hardest domain in τ-bench, the leading benchmark for customer-service agents.
May 8, 2026
Why booking agents fail mid-conversation, and what a specialized small model fixes.
May 22, 2026
What a year of running eval infrastructure taught us about what matters in deployments.