Accomplish_ai和NYU研究人员开源BoundaryBench基准测试,用于评估企业安全策略下代理的真实性能,揭示通用排行榜的误导性
AI 摘要
Accomplish_ai和NYU研究人员开源BoundaryBench基准测试,用于评估企业安全策略下代理的真实性能,揭示通用排行榜的误导性。
推荐理由常规快讯,保留列表
原文
RT Or Hiltch<br>Today we are open-sourcing @boundarybench, a new paper, benchmark, and GitHub repo that enterprises can use to find out the REAL performance of their agents.<br><br>Boundary-Bench was developed by researchers from @Accomplish_ai and NYU, where we tested 12 frontier agents across roughly 10,000 runs, with realistic enterprise policies, simulating environments with EDR, SASE, and DLP security tools enforcing those policies.<br><br>We did this because generic leaderboard scores are being generated under conditions no security team would ever allow, which means orgs are making deployment and risk decisions based on numbers that don't hold up.<br><br>The results are surprising >><br><video width="2048" height="1152" src="https://video.twimg.com/amplify_video/2085032993891254272/vid/avc1/3840x2160/j_cFwcVgE7faYj2P.mp4?tag=29" controls="controls" poster="https://pbs.twimg.com/amplify_video_thumb/2085032993891254272/img/5NeKI5xtypkBJ6wh.jpg"></video>
讨论
暂无评论。