论文 WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
笔记 精读:AgentLAB — LLM Agent 长期攻击的系统性基准
2026-07-11
论文 COMA: A Compositional Misleading Attack Class on Security-RAG, and a Causal Counterfactual Defense
论文 🔴 保留 · `Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Benchmarking`
论文 Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
论文 StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents