AREX: Towards a Recursively Self-Improving Agent for Deep Research
Abstract
Deep research requires agents to find answers that jointly satisfy multiple constraints.
Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks.
This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement.
We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents.
AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research.
To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model.
We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning.
To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected.
We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model.
Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요