Next upAI x Bio Pitch Contest
News

Google Research open-sources EnvHarness to reshape agent training environments

EnvHarness wraps existing agent-training environments in programmable layers tailored to a policy’s weaknesses while leaving the original tasks and verifiers intact.

D
Sep 26, 2026 · 2 min read

Google Research has open-sourced EnvHarness, a framework that reshapes the environments used to train AI agents without rewriting the underlying benchmark, task or verifier. The public repository’s license file identifies Apache License 2.0 terms.

EnvHarness builds targeted practice conditions around an agent’s observed weaknesses while continuing to use the source environment’s grading logic. It changes the environment side of training; separate skill-induction or reinforcement-learning stages still change the agent or policy. That distinction also separates the project from Google Research’s related work on answer-first tool-use training data.

The framework sits at an environment’s standard reset and step interface. Its current repository describes three composable components. Setup changes the initial state. Rule changes permitted actions, action effects or observations. Link brings in tasks from another environment. These layers can alter what the agent encounters, but the task and success predicate remain intact, preserving the original verifier.

EnvRigger, an automated process, adapts those layers to a particular policy. It treats the policy as a black box, examines successful and failed trajectories, diagnoses a recurring weakness, writes candidate components and tests them with fresh rollouts before accepting them. The resulting experiences can then feed a distinct learning stage.

The researchers evaluated EnvHarness on ALFWorld, WebArena, SWE-bench Verified, OfficeQA and SpreadsheetBench. In their experiments, skills learned in modified environments improved held-out results by as much as 9.0 points over skills learned in the original environments. They also reported that average execution steps on SWE-bench Verified fell from 55.01 to 49.61, a 9.8% reduction. These are research-team results, not an independent reproduction; the repository says its main tables report means over three independent runs and hold the Gemini model fixed while comparing skill sources.

Adding a benchmark requires a bridge that implements the framework’s ActionableEnv interface, including methods for resetting, stepping, observing, evaluating and saving or restoring state. The repository says EnvHarness is not an officially supported Google product. The paper was initially submitted on August 20, 2026, while the repository dates the paper and project-site release to August 21.

More news