OSGuard: A Benchmark for Safety in Computer-Use Agents
arXiv:2606.15034v1 Announce Type: new Abstract: Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks. However, task success alone can miss failures in which an agent re…
这条 研究论文 信号说明,来自 arXiv cs.AI 的信息已经不只是单点新闻,而是值得放进产品、研究和行业判断里的趋势线索。