FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 25 days ago • 149
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 26 days ago • 207
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published Aug 11 • 70
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome Paper • 2603.28407 • Published Mar 30 • 70
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification Paper • 2603.15726 • Published Mar 16 • 187