Abstract
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades. Our code is available at: https://github.com/dongyuanjushi/Opera.
Community
Opera is a verbal critic for long-horizon coding agents that treats each correction as a persistent note and follows it until the diagnosed problem is actually resolved, rather than stopping once feedback is delivered. It decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before sending it, and tracks the agent's subsequent actions to separate real fixes from mere compliance. As a test-time critic, Opera improves resolve rates by up to 12.4, 15.0, and 8.9 points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1 across four policy models, outperforms competitive critic baselines on all three benchmarks, and also helps when a policy critiques itself. Its guided rollouts also serve as near on-policy training data: fine-tuning Qwen3.5-9B on them boosts held-out SWE-Bench Pro performance by 10.2 points without a critic at inference, matching fine-tuning on a stronger model's rollouts while, unlike that approach, preserving performance when switching harnesses from OpenHands to Terminus-2.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents (2026)
- CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents (2026)
- Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents (2026)
- FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents (2026)
- Learning When and How to Intervene: A Hindsight-Distilled Sentinel for Coding Agents (2026)
- Cross-Benchmark Transfer from RL on Agentic Coding Tasks (2026)
- DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.33987 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper