SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 15 days ago • 157
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 19 days ago • 281
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement Paper • 2608.11951 • Published 21 days ago • 9
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Paper • 2608.12743 • Published 20 days ago • 44
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 263
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads Paper • 2608.04570 • Published 28 days ago • 41
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills Paper • 2608.03700 • Published 29 days ago • 9