Submitted by Haoran Zhang 86 π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Simplified Reasoning 31 3