Google Research RRSI Guide: Mastering Self-Improving AI Agents
In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the papers idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and...
In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the papers idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and its edit history, and then plug a simulated agent into RRSIs own Domain interface. Because we built the simulated environment ourselves, we know the true effect of every edit, which lets us audit RRSIs decisions against ground truth and compare them with an unregularized search that simply keeps whatever scores highest. Copy CodeCopiedUse a different Browser We install RRSI from the google-research repository, pinned to the commit this notebook was written against, since the package is not on PyPI. Its only dependency is the Anthropic client, which the search roles use to call Claude and which we never exercise. We then print the mapping the repository itself documents between the papers symbols and the functions that implement them: the empirical score and cost estimate in evaluate, the noise band in calibrate, Algorithm 2 in selection, the annealed edit budget in schedule, the leakage screen in critic, and the edit history with its yield, prune, stall and exploration summaries in history. RRSIConfig holds the papers hyperparameters, and every function below receives it exactly as the real loop does. Copy CodeCopiedUse a different Browser RRSI measures two numbers per harness: S, the reward averaged over every trial of every task, and C, the mean policy tokens per trial. TaskResult records one tasks trials and aggregates them. The detail worth copying into any agent evaluation is how missing trials are handled. When a candidate crashes on the hardest task, an estimator that drops the missing trials reports 0.750 and makes the crash look like an improvement. At the same time, RRSI counts each missing trial as zero reward with the full denominator and reports the same 0.500 as before, so a candidate cannot look better by destroying the trials it finds hard. Weighted rewards cover rubric-graded suites such as Harvey LAB, where S becomes the fraction of all criteria passed rather than the mean of per-task means. Copy CodeCopiedUse a different Browser Before any rule can separate a real gain from luck, it needs to know how far one harnesss score moves on its own. We build a small simulated agent, whose success on each task is a logistic function of harness skill minus task difficulty, and evaluate the unchanged starting harness six times on forty tasks with two trials each: the scores spread by 0.113 although nothing changed. calibrate turns repeated evaluations of the same harness into delta, twice the standard deviation of the difference between two runs. With 80 trials delta is about 0.108; with 3,200 trials it falls to about 0.013, in the range the paper reports for its instances (0.004 to 0.020). For selection purposes, any gain smaller than delta is indistinguishable from re-running the same harness. Copy CodeCopiedUse a different Browser Algorithm 2 is implemented as pure functions, so we can hand it candidates and read its reasons verbatim. We fix an incumbent at S 0.630 and 10,000 tokens, a best-ever score of 0.640 and delta 0.020, and send eight candidates through judge. A candidate below the floor, the best score ever seen minus delta, is rejected outright. A gain larger than delta must pay for its extra tokens under the rule that the relative cost change stays below 0.10 plus 40 times the gain; a +6 point gain at +20% tokens is admitted, and a +3 point gain at +150% tokens is not. Inside the band, scores are treated as a tie and a shaped score of 100 times the gain minus 15 times the cost change, plus a small bonus for a never-accepted structural component, decides; that is how RRSI admits candidate D, which scored lower than the incumbent but costs 20% fewer tokens, and how a new sub-agent breaks a tie that an equally scoring prompt tweak does not. A domain guard vetoes regardless of score. Copy CodeCopiedUse a different Browser select_round applies judge to every candidate in a round and keeps the highest-scoring admissible one. We give it the expensive candidate, the cheaper slightly worse one, and a candidate the critic already rejected. The highest scorer loses because its three-point gain does not pay for 150% more tokens, the critic-rejected candidate never reaches evaluation, and the winner moves the incumbent down by half a point while cutting its token cost by a fifth. The detail that keeps this safe is S*, which only ever rises: the floor is anchored to the best score ever measured rather than to the incumbent, so a chain of cheaper-but-slightly-worse swaps cannot walk the score away over many rounds. Copy CodeCopiedUse a different Browser The proposal side regularizes how edits are drafted, not which ones are kept. edit_budget implements the annealed L0 budget from the paper: a cosine schedule from b_max to b_min, and by default it allows up to four coordinated edits per candidate for the first eight rounds, three for the next five, and two for the last seven. The ceiling in the formula means the budget only reaches its minimum of one at t = T, one step after the run ends, which is easy to miss when reading the equation. Because every edit in a bundle inherits the bundles one measurement, the shrinking budget is what makes late-run history attributable to fewer components. The budget limits how many edits travel together and never restricts which mechanisms the harness may eventually contain. Copy CodeCopiedUse a different Browser The critic screens every candidate diff before any evaluation is spent, in two layers. The first is a deterministic precheck against a generic credential pattern plus the domains own denylist; with patterns for evolve-set task ids and grader paths it rejects a diff that memorises the answer to task_007, one that reads the expected output, one that leaks an API key, and an empty diff, all without calling a model. A clean diff falls through to the second layer, an intent review by Claude, and here the notebook surfaces a real gotcha: with RRSI_VERTEX_PROJECTS unset, rrsi.llm.generate computes an index modulo the number of projects outside its try block and raises ZeroDivisionError, so the helpful configuration error in the code is never reached. We also look at how edits are tagged: normalize keeps a declared component only when the diff carries evidence for it, so a prompt tweak cannot pose as a new skill to win the novelty bonus, and a mislabelled context-management change is tagged as what it is. Copy CodeCopiedUse a different Browser History writes one JSONL record per edit, and the loop derives four summaries from it. The tried set excludes edits the critic dropped, since they were never measured. The recent yield g_t is the best measured gain per component within the last n_prune rounds, and every component whose recent yield is not positive enters the prune set B_t, together with any machinery from that component that is still in the incumbent; in our history that flags a prompt edit accepted in round zero that has not paid off since. The stall flag fires when the score has moved less than delta over the last w rounds, and exploration then writes the directive the proposer receives, reserving a candidate slot for components the run has never exercised. Because the proposer is conditioned on all of this, a falsified