SearchMaster: Grounded and Regulated Self-Play for Search Agents
Abstract
Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipelines often depend on human-written tasks, expert demonstrations, or stronger teacher models. We present SearchMaster, a self-play framework that trains a single LLM from search tasks it generates, solves, and verifies in a local search environment. The key challenge is that self-generated tasks and rollouts can yield misleading signals: pseudo multi-hop questions, success-rate difficulty estimates that ignore search depth, and rollouts with excessive opening but little targeted evidence acquisition. SearchMaster addresses these failure modes with three controls. An Evidence-Chain Generator (ECG) grounds task generation in explicit cross-document evidence chains to reduce pseudo multi-hop questions. A Search-Depth Reward (SDR) scores task difficulty by the search depth of successful rollouts rather than success rate alone, keeping retained tasks search-intensive. An Over-Opening Penalty (OOP) regulates tool use by discouraging excessive document opening, avoiding long but shallow browsing. Verified Proposer and Solver rollouts are then jointly optimized with GRPO. Across six deep-search benchmarks, SearchMaster improves a Qwen3.5-9B backbone from 38.19% to 51.52% average accuracy, with a 30.1-point gain on BrowseComp-Plus. These results show that grounded and regulated self-play can provide effective search-agent training data without human-labeled QA pairs or expert demonstrations. The code is available at https://github.com/WentaoTan/SearchMaster.
Introduction
LLM-based search agents improve knowledge-intensive question answering by retrieving evidence from the web [22, 47, 14, 43]. Recent work has improved this capability either through reinforcement learning for tool use [10, 33, 53, 41, 31] or by constructing multi-turn search data for supervised training [54, 15, 6, 13]. Despite different training objectives, both rely on high-quality search data: questions that require cross-document retrieval and trajectories that acquire evidence through targeted search. Obtaining such data is expensive, often requiring human annotation or expert demonstrations.
Self-play appears to be a natural solution: the model can generate search tasks and learn by solving them. Existing search self-play methods have explored this direction by constructing questions from predefined answers [18] or through open-ended exploration [49]. While these approaches reduce reliance on externally written tasks, search-agent self-play has a distinct failure mode: self-generated tasks and rollouts can look useful while providing weak or misleading supervision. For example, the model may browse multiple documents yet still write a question answerable from one of them alone. Moreover, prior search self-play methods typically proxy task difficulty by rollout accuracy, often favoring tasks with moderate success rates; however, this signal may fail to distinguish shallow lookup tasks from those requiring multi-step search. Finally, document opening often exposes useful evidence, but the model may gradually over-rely on it, revisiting already-seen documents instead of acquiring targeted evidence. Thus, the central challenge is not merely generating more data, but stabilizing the self-play loop so that its training signals reinforce evidence-grounded tasks, multi-step search, and efficient tool use.
To address this challenge, we propose SearchMaster, a grounded and regulated self-play framework. SearchMaster uses a shared trainable policy for two roles: a Proposer, which generates multi-hop search tasks within a local search environment, and a Solver, which answers these tasks through browser-tool rollouts. A frozen Verifier evaluates task quality and solution accuracy, turning verified rollouts into rewards for GRPO training [30].
SearchMaster operationalizes this grounded-and-regulated principle through three mechanisms that jointly ground self-generated tasks and regulate the search behavior they induce (Figure 1). The Evidence-Chain Generator (ECG) requires the Proposer to incrementally build an explicit cross-document evidence chain during exploration and derive the question from the full chain. The Search-Depth Reward (SDR) scores a task by the search depth of its successful rollouts, encouraging the Proposer to generate tasks that require deeper retrieval. The Over-Opening Penalty (OOP) constrains the open-to-search ratio, discouraging redundant opening without penalizing tasks that legitimately require inspecting multiple documents.

Figure 1. SearchMaster grounds tasks and regulates behavior. (a) ECG builds explicit evidence chains to reduce pseudo multi-hop questions. (b) SDR rewards deeper successful rollouts, favoring multi-step search. (c) OOP penalizes excessive document opening to reduce long but shallow browsing.
Overall, our main contribution is to formulate search-agent self-play as a problem of grounding and regulating self-generated supervision. This perspective turns a potentially noisy self-play loop into one that promotes cross-document task grounding, search-depth-aware task selection, and balanced tool use. Our experiments show that SearchMaster substantially improves the Qwen3.5-9B backbone (Qwen Team [28]) across six deep-search benchmarks, raising average accuracy from 38.19% to 51.52% (+13.3 points), with an especially large 30.1-point gain on BrowseComp-Plus (Chen et al. [5]) (30.12% to 60.24%). Notably, this improvement relies solely on a local search environment, without human-labeled data or expert supervision. We will release the code, trained checkpoints, and self-generated training data.
Related Work
Search Agent Training
Early search agents primarily relied on prompting or human demonstrations for web tool use (Guu et al. [8]; Lewis et al. [11]; Schick et al. [29]). For instance, WebGPT (Nakano et al. [22]) trained language models to browse and cite web evidence using human demonstrations, while ReAct (Yao et al. [47]) used prompting to interleave reasoning with tool usage. These approaches show that external evidence improves answer reliability; however, achieving stable long-horizon search behavior demands further training.
Recent efforts to improve search-agent training broadly fall into two categories. The first focuses on optimizing search behavior through reinforcement learning and reward design on existing QA tasks (Tan et al. [35]; Mei et al. [20]). Search-R1 (Jin et al. [10]) and R1-Searcher (Song et al. [33]) use outcome-based rewards to encourage search during reasoning. Recognizing the limitations of final-answer rewards in guiding long trajectories, subsequent methods incorporate richer process-level signals: R-Search (Zhao et al. [53]) introduces multiple rewards for search and reasoning quality, StepSearch (Wang et al. [41]) applies step-wise feedback to improve retrieval and reduce redundancy, AutoRefine (Shi et al. [31]) iteratively refines retrieved knowledge across search calls, and AutoSearch (Sun et al. [34]) adapts search depth to balance accuracy and efficiency.
The second line of work aims to construct higher-quality training data by generating multi-turn search trajectories and challenging tasks. OpenResearcher (Li et al. [15]) builds an offline search environment and leverages a teacher model to produce multi-turn search-and-browse trajectories. WebSailor (Li et al. [13]) emphasizes task difficulty by creating high-uncertainty web tasks through structured sampling and information obfuscation. Extensions include synthetic data generation, multi-agent collaboration, and staged post-training (Chu et al. [6]; Liu et al. [17]; Yao et al. [48]). Despite these advances, acquiring search-worthy tasks and trajectories remains costly, often relying on human annotation, expert models, or complex generation pipelines.
Overall, while these approaches advance search-agent training, obtaining scalable multi-hop search data remains a challenge. This motivates self-play, where models generate and learn from their own search experiences.
Self-Play for Search Agent Training
Self-play offers a promising avenue to reduce reliance on external labeled data. In reasoning and tool-use domains, prior work shows that models can bootstrap learning by generating tasks and solutions themselves when validity and correctness can be verified (Huang et al. [9]; Zhao et al. [52]; Acikgoz et al. [1]). However, self-play for search agents is more challenging, since generated tasks must be grounded in evidence, require genuine multi-hop search, and induce meaningful rather than shallow tool-use trajectories.
Search Self-Play (Lu et al. [18]) is an early attempt at self-play for search-agent training: it samples a target answer from a predefined set, and the Proposer searches for supporting evidence to formulate a question pointing to it. While this enables self-play, task generation remains constrained by predefined answers rather than open-ended exploration. Dr. Zero (Yue et al. [49]) instead lets the Proposer generate questions directly through external search, rewarding task difficulty by the Solver’s success rate. This moves toward more autonomous task generation, but success rate remains a coarse signal: it reflects whether the Solver answers correctly, not whether the task requires nontrivial search beyond a shallow lookup. Beyond this difficulty-estimation issue, pseudo multi-hop questions and over-opening drift remain underexplored. SearchMaster addresses these search-specific issues by making evidence construction, task difficulty, and tool-use balance explicit parts of the self-play loop, enabling more reliable self-evolution of search agents.
Method
Problem Setup and Framework Overview
Figure 2 illustrates the SearchMaster framework. It trains a single policy within a search environment constructed over document collection . The policy interacts with through browser tools , where search returns ranked documents for a query, open reveals a document’s content, and find locates exact matches within opened documents.

Figure 2. Overview of the SearchMaster self-play loop. A shared policy acts as both Proposer and Solver in a local search environment. The Proposer builds evidence-chain tasks (ECG), filters remove invalid candidates, and the Solver answers passing tasks through independent rollouts. A frozen Verifier, SDR, and OOP provide rewards, and verified Proposer/Solver rollouts are optimized jointly with GRPO.
During training, the same policy acts as both the Proposer and the Solver . The Verifier is a frozen copy of the initial model and remains fixed.
Each self-play iteration samples unlabeled seed documents from . For each seed, the Proposer samples task-generation rollouts, each maintaining an ECG evidence chain and producing a candidate task . Quality filters remove invalid candidates before Solver rollouts. For each passing task, the Solver receives only and produces answer rollouts. The frozen Verifier judges each answer, producing correctness scores , and we records the search depth of each rollout as the number of unique search queries. These signals define rewards for both roles: rewards Solver correctness, successful-rollout depths estimate task difficulty for the Proposer through SDR, and OOP is subtracted from both rewards to discourage excessive opening. The resulting samples form a Proposer GRPO group over the task-generation rollouts and a Solver GRPO group over the answer rollouts of the highest-reward passing task. Since one seed can produce multiple passing tasks and up to Solver rollouts, using all Solver rollouts would make Solver samples dominate the update; we therefore keep only the Solver rollouts of the highest-reward passing task.
Evidence-Chain Task Generation
Generating genuine multi-hop search tasks is challenging because pseudo multi-hop tasks can arise even when the Proposer is explicitly instructed to generate questions requiring cross-document evidence. In one failure mode, the Proposer stays close to the seed, searching for surface-level facts and asking a question answerable from the seed alone. In another, it reaches related documents but does not encode the cross-document links into the question, leaving the final answer recoverable from just one visited document. Examples of both cases are provided in the Supplementary Material.
To address this, the Evidence-Chain Generator (ECG) turns task proposal into explicit chain construction. Starting from a seed document , the Proposer repeatedly follows entity or fact links across documents: it selects a salient entity, searches for a connected document, identifies a new entity or fact, and extends the chain. The question is then generated with explicit requirements for full-chain necessity and no single-document answer.
Formally, for each , the Proposer samples independent rollouts:
where is the Proposer prompt (Supplementary Material), the tool-use trace, the evidence chain of entities or facts, the question generated from , and its reference answer. By deriving from this constrained chain rather than a single document, ECG biases task generation toward questions that require combining evidence across documents.
Before launching Solver rollouts, SearchMaster applies four filters to each candidate task:
Tool-use filter: treats candidates with too few search or open actions as shallow.
Format filter: checks that are present and that contains at least two evidence items.
Validity filter: asks the Verifier whether is valid, is supported by , and the chain is logically sound.
Parametric-knowledge filter: removes questions the Verifier can answer without tools.
A candidate that fails any filter receives a low Proposer reward, whereas those passing all four are sent to the Solver.
Solver Rollouts and Search Depth
After a task passes the filters, the Solver receives only the question and produces independent answer rollouts
where is the Solver prompt (Supplementary Material), is the tool-use trace, and is the final answer. The frozen Verifier compares with the reference answer and outputs a correctness score . We also compute the search depth as the number of unique search queries in , and let denote the number of successful rollouts. These signals play different roles: rewards Solver correctness, while the search depths of successful rollouts estimate how much search the proposed task requires for the Proposer reward.
Reward Design and Over-Opening Penalty
The Proposer reward favors tasks that are solvable but still challenging for the current Solver. Candidates that fail the quality filters receive low rewards. Passing tasks with or receive only a low base reward , since they are respectively too hard or too easy. If all tasks generated from a seed receive rewards that do not exceed , we discard all data generated from that seed.
For the remaining tasks with , success rate alone does not distinguish shallow lookup from deeper search. We therefore measure task difficulty by the easiest successful search depth,
rather than by the mean or maximum search depth. If any successful rollout solves the task with shallow search, the task admits a shortcut and should not receive a high difficulty reward. We then compute the Search-Depth Reward (SDR):
where is the target depth at which the reward saturates. This cap prevents SDR from rewarding arbitrarily many searches, while still favoring tasks whose easiest successful solution requires nontrivial search depth.
We now summarize the base Proposer reward as follows, where the first three cases correspond to the tool-use, format, and validity/parametric filters introduced above:
For the Solver, the base reward encourages answering correctly with tool-supported evidence:
As training proceeds, the model can drift toward excessive open calls. Early in training, opening documents often provides useful information that helps the model answer correctly, so such behavior is reinforced by answer rewards. Over time, however, the policy may over-rely on open: many opens revisit already-seen documents rather than acquiring new evidence, producing long but shallow trajectories. A concrete over-opening case is provided in the Supplementary Material. Accordingly, we introduce the Over-Opening Penalty (OOP) to regularize this behavior in both roles, since the Proposer and Solver share the same policy . OOP penalizes the ratio of open to search actions rather than the raw number of opened documents, so it discourages redundant opening without suppressing legitimate multi-document exploration. Specifically, let and count open and search calls in a trajectory . For , define:
where and are the open-to-search ratios at which the penalty starts and reaches its maximum, so rises linearly from at to at . The final rewards then subtract OOP with coefficient :
Training and Optimization
After rewards are assigned, each retained seed document contributes two GRPO groups: a Proposer group containing the task-generation rollouts, and a Solver group containing the answer rollouts of the highest-reward passing task. Seeds whose highest task reward does not exceed are discarded. Because of this seed-level gate, the final training data matches our high-quality search-data goal: retained tasks are grounded for cross-document retrieval, successful rollouts exhibit nontrivial search depth, and OOP regularizes tool use toward targeted search.
Within each retained group, rewards are normalized into advantages by subtracting the group mean and dividing by the group standard deviation with a small numerical constant. The same advantage is assigned to every generated token in the sample. We then update the shared policy with a token-level clipped objective and a KL regularizer,
where the expectation is over unmasked generated tokens , is the per-token importance ratio, and is the asymmetric clip range. Tool observations remain in the context but are masked out of the loss, so gradients flow only through model-generated tokens. Proposer and Solver samples are optimized jointly under the shared policy . Pseudocode for the full self-play iteration is provided in the Supplementary Material.
Experiments
Experimental Setup
Implementation Details. The local search database is the offline corpus released by OpenResearcher [15] (about 15M documents, B tokens), indexed with a Qwen3-Embedding-8B [46] FAISS retriever. We use Qwen3.5-9B (Qwen Team 2026) as the default backbone and train it with GRPO at a constant learning rate of , clip range , and KL weight . Each iteration randomly samples seed documents from and, after the four filters and the seed-level quality gate, retains 64 qualified seed documents for training. For each seed document, the Proposer samples rollouts, and each passing candidate launches independent Solver rollouts for verification and reward computation, with up to 200 tool calls per rollout. Although Solver rollouts are generated for all passing candidates to compute rewards, only the Solver group of the highest-reward passing task is used for optimization, so each retained seed contributes training samples. Each iteration therefore contributes a global optimization batch size of . For reward design, we set , , and OOP parameters , , and . Rollouts use a 256K-token context and decoding temperature 1.0. We train for 20 iterations, so SearchMaster learns from qualified seed documents in total. Prompt templates, filter thresholds, and additional training hyperparameters are provided in the Supplementary Material.
Benchmarks. We evaluate SearchMaster on both offline and online search benchmarks. The offline benchmark is BrowseComp-Plus [5], with 830 browsing questions over an offline retrieval corpus drawn from a different source than our OpenResearcher training corpus. For transfer to live web search, we use five online search benchmarks: BrowseComp [42] (1,266 hard-to-find factual questions), GAIA [21] (103 text-only tool-use assistant tasks), SEAL-0 [27] (111 search-intensive long-tail questions), WebWalkerQA [45] (680 multi-page navigation questions), and XBench-DeepSearch [4] (100 web-search questions). During evaluation, we do not use any context-management method; the context length is 256K tokens, decoding temperature is 1.0, and outputs are scored by GPT-5 [32] against the ground-truth answers, with the judging prompt provided in the Supplementary Material.
Main Results
Table 1 reports the main results. On BrowseComp-Plus, which supports multiple retrievers, we report all methods under its Qwen3-Embedding-8B [46] retriever setting for fair comparison.
| Method | BrowseComp-Plus | BrowseComp | GAIA | SEAL-0 | WebWalkerQA | XBench | Avg. |
| Proprietary models and systems | |||||||
| OpenAI o3 (OpenAI 2025c) | 63.49 | 49.70 | 70.50 | 15.30 | 71.70 | 67.00 | – |
| OpenAI o4-mini (OpenAI 2025c) | – | 28.30 | 60.00 | – | – | – | – |
| GPT-5-high (Singh et al. 2025) | 70.12 | 54.90 | 76.40 | 43.20 | – | 77.80 | – |
| GPT-4.1 (OpenAI 2025b) | 35.42 | – | – | – | – | – | – |
| Claude Sonnet 4 (Anthropic 2025) | 36.75 | 12.20 | 68.30 | – | 61.70 | 65.00 | – |
| Claude Opus 4 (Anthropic 2025) | 36.14 | – | – | – | – | – | – |
| Gemini 2.5 Pro (Gemini Team 2025) | 28.67 | – | – | – | – | – | – |
| Gemini 2.5 Flash (Gemini Team 2025) | 33.01 | – | – | – | – | – | – |
| OpenAI DeepResearch (OpenAI 2025a) | – | 51.50 | 67.40 | – | – | – | – |
| Open-source models and systems | |||||||
| GLM-4.5 (Zeng et al. 2025) | – | 26.40 | 66.00 | – | 65.60 | 70.00 | – |
| Kimi K2 (Team et al. 2025a) | – | 14.10 | 57.70 | – | 63.00 | 50.00 | – |
| DeepSeek-V3.1 (Liu et al. 2024) | – | 30.00 | 63.10 | – | 61.20 | 71.00 | – |
| Qwen3-32B (Yang et al. 2025) | 10.36 | – | – | – | – | – | – |
| SearchR1-32B (Jin et al. 2025) | 10.36 | – | – | – | – | – | – |
| gpt-oss-120B-high (Agarwal et al. 2025) | 42.89 | – | – | – | – | – | – |
| WebThinker-32B (Li et al. 2026a) | – | 2.80 | 48.50 | – | 46.50 | 24.00 | – |
| WebExplorer-8B (Liu et al. 2025) | – | 15.70 | 50.00 | – | 62.70 | 53.70 | – |
| WebDancer-QwQ (Wu et al. 2026) | – | 3.80 | 51.50 | – | 47.90 | – | – |
| WebShaper-72B (Tao et al. 2025b) | – | – | 60.10 | – | 52.20 | – | – |
| DeepDive-32B (Lu et al. 2025b) | – | 15.30 | – | 25.50 | – | 51.80 | – |
| OffSeeker-8B (Zhou et al. 2026) | – | 12.80 | 51.50 | – | 61.70 | 49.00 | – |
| WebSailor-72B (Li et al. 2025b) | – | 12.00 | 55.40 | – | – | 55.00 | – |
| WebSailor-V2-30B (Li et al. 2025a) | – | 35.30 | 74.10 | – | – | 73.70 | – |
| WebLeaper-C (Tao et al. 2025a) | – | 38.80 | 73.20 | 48.60 | – | 72.00 | – |
| BrowseMaster (Pang et al. 2025) | – | 30.00 | 68.00 | – | 62.10 | 66.00 | – |
| OpenResearcher-30B (Li et al. 2026b) | 54.80 | 26.30 | 64.10 | – | – | 65.00 | – |
| REDSearcher-30B (Chu et al. 2026) | – | 42.10 | 80.10 | – | – | – | – |
| Tongyi-DR-30B (Team et al. 2025c) | – | 43.40 | 70.90 | – | 72.20 | 75.00 | – |
| MiroThinker-72B (Team et al. 2025b) | – | 47.10 | 81.90 | 51.00 | 62.10 | 77.80 | – |
| Ours (Qwen3.5-9B backbone) | |||||||
| Qwen3.5-9B (Qwen Team 2026) | 30.12 | 20.93 | 50.49 | 26.13 | 41.47 | 60.00 | 38.19 |
| SearchMaster | 60.24 | 28.75 | 57.28 | 35.14 | 59.71 | 68.00 | 51.52 |
Table 1. Main evaluation results. All columns report accuracy (%), and Avg. is the arithmetic mean over the six benchmarks. External baselines are reported from their source papers when available.
Self-play yields substantial gains on BrowseComp-Plus. SearchMaster raises the Qwen3.5-9B backbone from 30.12% to 60.24%, an absolute gain of 30.1 points. This lifts the 9B model above much larger open-source systems such as gpt-oss-120B-high (42.89%) and proprietary models such as GPT-4.1 (35.42%) and Claude Opus 4 (36.14%), while narrowing the gap to stronger reasoning systems such as OpenAI o3 (63.49%). Notably, this is achieved purely through self-play over a local search environment, without human-written questions or advanced teacher models. These results show that a single model can bootstrap strong deep-search ability entirely from its own generated tasks. The Supplementary Material examines the source of this gain: the base model already issues valid tool calls but searches ineffectively, whereas SearchMaster learns to search more deeply and follow evidence across documents.
The learned search behavior generalizes to live web search. Although SearchMaster is trained entirely in an offline search environment, it also performs strongly on online benchmarks that query the live, open web. It improves the backbone on all five online search benchmarks, with gains ranging from +6.8 on GAIA to +18.2 on WebWalkerQA. This transfer from offline training to live web search shows that SearchMaster equips the model with search behavior that generalizes beyond its training environment.
Ablation Study
In this section, we study the contribution of ECG, SDR, and OOP. All variants use the same training budget and filtering pipeline unless otherwise noted. We first train a naive self-play baseline that keeps the same task filters and Solver verification as SearchMaster, but uses rollout accuracy as the task-difficulty signal, via a success-rate term , where is the number of successful Solver rollouts among attempts, so that the reward peaks when half of the rollouts succeed. Starting from this baseline, ECG changes the Proposer prompt into evidence-chain construction, SDR replaces the success-rate difficulty signal with the search-depth reward, and OOP adds the over-opening penalty.
Table 2 reports the results on BrowseComp-Plus. Naive self-play alone lifts the backbone from 30.12% to 45.18%, showing that self-play over a local corpus already provides a useful signal. Adding ECG, SDR, and OOP improves accuracy monotonically, and the full SearchMaster reaches 60.24%. These mechanisms improve complementary parts of the self-play loop: chain-based task construction (ECG), search-depth-aware task selection (SDR), and tool-use regularization (OOP). The rest of this section examines whether each mechanism produces its intended behavioral effect, beyond the final accuracy.
| Variant | ECG | SDR | OOP | BrowseComp-Plus |
| Qwen3.5-9B | - | - | - | 30.12 |
| naive self-play | 45.18 | |||
| + ECG | ✓ | 53.25 | ||
| + SDR | ✓ | 52.89 | ||
| + ECG + SDR | ✓ | ✓ | 57.71 | |
| SearchMaster | ✓ | ✓ | ✓ | 60.24 |
Table 2. Ablation on BrowseComp-Plus. ECG, SDR, and OOP each contribute a further gain over naive self-play.
Does ECG Reduce Pseudo Multi-Hop Tasks? ECG makes evidence-chain construction explicit before question generation, so we directly test whether it reduces pseudo multi-hop tasks. We compare three settings on the same 1,000 seed documents: Qwen3.5-9B with the naive self-play prompt, Qwen3.5-9B with the ECG prompt, and the trained SearchMaster with ECG. Each setting generates 10 tasks per seed document. A GLM-5 judge labels each task as True Multi-Hop if it is answerable only by combining evidence across several documents, Pseudo Multi-Hop if it superficially involves multiple entities but is solvable by a single lookup, and Invalid if it leaks the answer or is unsupported by the retrieved evidence.
Table 3 shows that adding ECG to the base model increases true multi-hop tasks from 24.2% to 45.6% and reduces invalid tasks from 28.7% to 15.8%. After SearchMaster training, the true multi-hop rate further rises to 78.6%, while pseudo multi-hop drops to 15.0%. The Supplementary Material further compares Qwen3.5-9B with the naive and ECG prompts on the same seeds. In one case, the naive prompt searches around seed-level facts and asks a seed-answerable question; in another, it reaches related documents but leaves their links unused in the final question. These examples suggest that ECG helps in two ways: it drives the model to explore beyond seed-level facts, and it requires the explored evidence to be organized into an explicit chain before question generation. Together, these effects turn loosely related browsing into task-relevant cross-document dependencies.
| Method | True Multi-Hop ↑ | Pseudo ↓ | Invalid ↓ |
| Qwen3.5-9B + naive | 24.2 | 47.1 | 28.7 |
| Qwen3.5-9B + ECG | 45.6 | 38.6 | 15.8 |
| SearchMaster + ECG | 78.6 | 15.0 | 6.4 |
Table 3. Task-quality comparison across three settings. Each setting generates 10 tasks for each of the same 1,000 seed documents (10,000 per setting), judged by GLM-5 [50].
Does SDR Keep the Training Tasks Search-Intensive?
SDR scores a task by , the minimum search depth among successful Solver rollouts (Eq. 4). Because reflects the easiest successful solution, a large value indicates that the task cannot be solved by a shallow lookup. To test whether SDR keeps retained tasks search-intensive during training, we compare the naive self-play baseline, which uses a success-rate difficulty signal, with the +SDR setting from Table 2. We track across iterations in Figure 3, discarding a few tasks with extremely large for visualization.

Figure 3. Search depth of the training tasks over self-play. Lines show the mean and bands show the range across tasks.
Early in training, both settings produce tasks with increasing , suggesting that self-play initially discovers tasks requiring more search. Later, the two runs diverge. Under the success-rate signal, the Proposer has no direct incentive to preserve search depth: as the Solver improves, retained tasks become solvable with fewer searches, and settles around 2–3. SDR instead rewards tasks whose easiest successful solution still requires substantial search, so remains in the 8–10 range even as the Solver strengthens. This shows that SDR prevents the training tasks from decaying into shallow lookup tasks and maintains a multi-step search signal.
Does OOP Control Open Overuse?
To evaluate whether OOP regulates tool use, we compare two settings from Table 2: +ECG+SDR without OOP and the full SearchMaster with OOP. Since the Proposer and Solver share the same policy, we track the average open-to-search ratio of both roles across training iterations (Figure 4).

Figure 4. Average open/search ratio of the Proposer and the Solver over training. Without OOP, both roles drift toward higher open/search ratios; with OOP, the average ratio stays low for both roles.
Without OOP, the open-to-search ratio of both roles increases steadily, indicating a drift toward opening documents rather than issuing targeted searches. To inspect this behavior, the Supplementary Material analyzes late-training Solver rollouts with at least twice as many opens as searches. In these over-opening rollouts, 56.3% of the extra open calls simply re-open already-seen documents rather than acquiring new evidence. With OOP, the ratio of both roles remains low, showing that the penalty curbs over-opening while still allowing the searches and opens needed for genuine multi-document tasks.
Conclusion
We introduce SearchMaster, a self-play framework for training search agents from tasks discovered in a local search environment. The central challenge is not merely question generation, but transforming autonomous exploration into useful training data for search behavior. SearchMaster addresses this challenge through ECG for evidence-grounded task construction, SDR for search-depth difficulty calibration, and OOP for preventing over-opening in tool-use trajectories. Across six deep-search benchmarks, SearchMaster raises the same 9B backbone’s average accuracy from 38.19% to 51.52%, with a 30.1-point gain on BrowseComp-Plus.
While SearchMaster reduces dependence on annotated QA data, it does not eliminate all costs. Training still requires a searchable environment, multiple Solver rollouts per proposed task, and Verifier calls for grounding, parametric filtering, and correctness. The current implementation operates in a local search environment, which improves reproducibility, though it may limit coverage of the open web’s full diversity and volatility. Future work can extend SearchMaster beyond local corpora to more diverse and dynamic web environments, while reducing the rollout and verification cost of self-play training.
References
- [1]Acikgoz, E. C.; Qian, C.; Hübotter, J.; Ji, H.; Hakkani-Tür, D.; and Tur, G. 2026. Tool-ro: Self-evolving llm agents for tool-learning from zero data. arXiv preprint arXiv:2602.21320.
- [4]Chen, K.; Ren, Y.; Liu, Y.; Hu, X.; Tian, H.; Xie, T.; Liu, F.; Zhang, H.; Liu, H.; Gong, Y.; et al. 2025. xbench: Tracking agents productivity scaling with profession-aligned real-world evaluations. arXiv preprint arXiv:2506.13651.arxiv.org/abs/2506.13651
- [5]Chen, Z.; Ma, X.; Zhuang, S.; Nie, P.; Zou, K.; Sharify-moghaddam, S.; Liu, A.; Green, J.; Patel, K.; Meng, R.; et al. 2026. BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 22349–22370.
- [6]Chu, Z.; Wang, X.; Hong, J.; Fan, H.; Huang, Y.; Yang, Y.; Xu, G.; Zhao, C.; Xiang, C.; Hu, S.; et al. 2026. Redsearch: A scalable and cost-efficient framework for long-horizon search agents. arXiv preprint arXiv:2602.14234.arxiv.org/abs/2602.14234
- [8]Guu, K.; Lee, K.; Tung, Z.; Pasupat, P.; and Chang, M. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, 3929–3938. PMLR.
- [9]Huang, C.; Yu, W.; Wang, X.; Zhang, H.; Li, Z.; Li, R.; Huang, J.; Mi, H.; and Yu, D. 2025. R-zero: Self-evolving reasoning llm from zero data. arXiv preprint arXiv:2508.05004.arxiv.org/abs/2508.05004
- [10]Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; and Han, J. 2025. Search-rl: Training llms to reason and leverage search engines with reinforcement learning. arXiv preprint arXiv:2503.09516.
- [11]Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33: 9459–9474.
- [14]Li, X.; Jin, J.; Dong, G.; Qian, H.; Wu, Y.; Wen, J.-R.; Zhu, Y.; and Dou, Z. 2026a. Webthinker: Empowering large reasoning models with deep research capability. Advances in Neural Information Processing Systems, 38: 120091–120131.arxiv.org/abs/2504.21776
- [15]Li, Z.; Jiang, D.; Ma, X.; Zhang, H.; Nie, P.; Zhang, Y.; Zou, K.; Xie, J.; Zhang, Y.; and Chen, W. 2026b. Opensearcher: A fully open pipeline for long-horizon deep research trajectory synthesis. arXiv preprint arXiv:2603.20278.arxiv.org/abs/2603.20278
- [17]Liu, J.; Li, Y.; Zhang, C.; Li, J.; Chen, A.; Ji, K.; Cheng, W.; Wu, Z.; Du, C.; Xu, Q.; et al. 2025. Webexplorer: Explore and evolve for training long-horizon web agents. arXiv preprint arXiv:2509.06501.arxiv.org/abs/2509.06501
- [18]Lu, H.; Wen, Y.; Cheng, P.; Ding, R.; Guo, J.; Xu, H.; Wang, C.; Chen, H.; Jiang, X.; and Jiang, G. 2025a. Search self-play: Pushing the frontier of agent capability without supervision. arXiv preprint arXiv:2510.18821.arxiv.org/abs/2510.18821
- [20]Mei, J.; Hu, T.; Fu, D.; Wen, L.; Yang, X.; Wu, R.; Cai, P.; Cai, X.; Gao, X.; Yang, Y.; et al. 2025. O²-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering. arXiv preprint arXiv:2505.16582.arxiv.org/abs/2505.16582
- [21]Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; and Scialom, T. 2024. Gaia: a benchmark for general ai assistants. In International Conference on Learning Representations, volume 2024, 9025–9049.arxiv.org/abs/2311.12983
- [22]Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332.arxiv.org/abs/2112.09332
- [27]Pham, T.; Nguyen, N.; Zunjare, P.; Chen, W.; Tseng, Y.-M.; and Vu, T. 2025. SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models. arXiv preprint arXiv:2506.01062.arxiv.org/abs/2506.01062
- [28]Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents.
- [29]Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; and Scialom, T. 2023. Toolformer: Language models can teach themselves to use tools. Advances in neural information processing systems, 36: 68539–68551.
- [30]Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.arxiv.org/abs/2402.03300
- [31]Shi, Y.; Li, S.; Wu, C.; Liu, Z.; Fang, J.; Cai, H.; Zhang, A.; and Wang, X. 2026. Search and refine during think: Facilitating knowledge refinement for improved retrieval-augmented reasoning. Advances in Neural Information Processing Systems, 38: 155930–155958.arxiv.org/abs/2505.11277
- [32]Singh, A.; Fry, A.; Perelman, A.; Tart, A.; Ganesh, A.; El-Kishky, A.; McLaughlin, A.; Low, A.; Ostrow, A.; Ananthram, A.; et al. 2025. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267.arxiv.org/abs/2601.03267
- [33]Song, H.; Jiang, J.; Min, Y.; Chen, J.; Chen, Z.; Zhao, W. X.; Fang, L.; and Wen, J.-R. 2025. R1-searcher: Incentivizing the search capability in llms via reinforcement learning. arXiv preprint arXiv:2503.05592.arxiv.org/abs/2503.05592
- [34]Sun, J.; Chong, W.; Tu, S.; Zhang, Q.; Zhang, Y.; Chai, J.; Wang, X.; Lin, W.; Yin, G.; and Zhao, D. 2026. AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning. In Findings of the Association for Computational Linguistics: ACL 2026, 28059–28079.
- [35]Tan, Z.; Huang, J.; Wu, Q.; Zhang, H.; Zhuang, C.; and Gu, J. 2026. Rag-r1: Incentivizing the search and reasoning capabilities of llms through multi-query parallelism. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 33187–33195.
- [41]Wang, Z.; Zheng, X.; An, K.; Ouyang, C.; Cai, J.; Wang, Y.; and Wu, Y. 2025. Stepsearch: Igniting llms search ability via step-wise proximal policy optimization. arXiv preprint arXiv:2505.15107.arxiv.org/abs/2505.15107
- [42]Wei, J.; Sun, Z.; Papay, S.; McKinney, S.; Han, J.; Fulford, I.; Chung, H. W.; Passos, A. T.; Fedus, W.; and Glaese, A. 2025a. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516.arxiv.org/abs/2504.12516
- [43]Wei, Z.; Yao, W.; Liu, Y.; Zhang, W.; Lu, Q.; Qiu, L.; Yu, C.; Xu, P.; Zhang, C.; Yin, B.; et al. 2025b. Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 7920–7939.
- [45]Wu, J.; Yin, W.; Jiang, Y.; Wang, Z.; Xi, Z.; Fang, R.; Zhang, L.; He, Y.; Zhou, D.; Xie, P.; et al. 2025. Webwalker: Benchmarking llms in web traversal. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10290–10305.
- [46]Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388.arxiv.org/abs/2505.09388
- [47]Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.arxiv.org/abs/2210.03629
- [48]Yao, Y.; Zhu, H.; Wang, P.; Ren, J.; Yang, X.; Chen, Q.; Li, X.; Shi, D.; Li, J.; Wang, Q.; et al. 2026. O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL. arXiv preprint arXiv:2601.03743.arxiv.org/abs/2601.03743
- [49]Yue, Z.; Upasani, K.; Yang, X.; Ge, S.; Nie, S.; Mao, Y.; Liu, Z.; and Wang, D. 2026. Dr. Zero: Self-Evolving Search Agents without Training Data. arXiv preprint arXiv:2601.07055.arxiv.org/abs/2601.07055
- [50]Zeng, A.; Lv, X.; Hou, Z.; Du, Z.; Zheng, Q.; Chen, B.; Yin, D.; Ge, C.; Huang, C.; Xie, C.; et al. 2026. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763.arxiv.org/abs/2602.15763
- [52]Zhao, A.; Wu, Y.; Wu, T.; Xu, Q.; Yue, Y.; Lin, M.; Wang, S.; Wu, Q.; Zheng, Z.; and Huang, G. 2026a. Absolute zero: Reinforced self-play reasoning with zero data. Advances in Neural Information Processing Systems, 38: 105816–105879.arxiv.org/abs/2505.03335
- [53]Zhao, Q.; Wang, R.; Xu, D.; Zha, D.; Bowen, M.; Wang, Z.; Jia, S.; Liu, L.; and Wang, X. 2026b. R-search: Empowering llm reasoning with search via multi-reward reinforcement learning. In Findings of the Association for Computational Linguistics: ACL 2026, 38030–38046.arxiv.org/abs/2506.04185
- [54]Zheng, Y.; Fu, D.; Hu, X.; Cai, X.; Ye, L.; Lu, P.; and Liu, P. 2025. Deepresearch: Scaling deep research via reinforcement learning in real-world environments. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 414–431.arxiv.org/abs/2504.03160