Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Paper • 2608.18940 • Published 12 days ago • 34
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 17 days ago • 280
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published 18 days ago • 46
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published about 1 month ago • 263
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published Jul 31 • 24