arxiv:2602.07276
Pengrui Han
barryhpr
AI & ML interests
None yet
Recent Activity
upvoted a paper about 12 hours ago
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents upvoted a paper about 1 month ago
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior upvoted a paper 2 months ago
Interactive Evaluation Requires a Design Science