Google Research lays out unresolved privacy and security risks for AI agents
A Google Research workshop report sets out a research agenda for AI agents that handle personal data, interpret ambiguous inputs and take action through probabilistic plans.
Google Research published a workshop report on October 5 that maps unresolved privacy and security questions for AI agents across system, model, user and ecosystem levels. It is a research agenda, not a claim that Google has deployed or validated a complete solution.
The researchers point to three features that strain familiar software controls: Agents accept ambiguous, unstructured inputs; generative models produce probabilistic control flows; and delegated autonomy reduces continuous user oversight. Agents may also need personal data and permission to act across tools, leaving one-time permission prompts or repeated confirmations ill-suited to address every risk.
The report grew out of Google’s Contextual Agent Privacy and Security workshop, held in New York City in November 2025. Google says more than 30 academic and industry leaders took part, and the related Google Research publication record lists more than 50 authors.
Google frames the agenda around Contextual Integrity, a theory that defines privacy through the appropriateness of information flows under social norms. Those norms depend on who is involved, what information is moving and the conditions governing its transmission. The report applies the same lens to “contextual security”: whether an agent’s action is appropriate in the setting where it occurs.
One proposed architecture puts a contextual policy engine in a supervisor layer. It would generate policy in real time from the user’s request, the changing context and tools discovered during a task, then apply that policy before information leaves the user’s workspace. Google presents the design as a research subject, not a deployed product.
At the system level, the agenda calls for sandboxing to contain failures, reliable agent identity, dynamically restricted capabilities, and data access that can be granted or revoked as context changes. At the model level, it calls for agents that can clarify underspecified prompts and judge whether actions such as sharing data remain appropriate as circumstances change.
For users, the report argues that conventional notice-and-choice controls depend on foreseeable actions. Generative agents can produce too many possible actions and confirmation requests for that model to hold up, leading the authors to call for dynamic, contextual and personalized controls aligned with users’ mental models. At the ecosystem level, they identify open questions about agents colluding to violate norms, as well as mechanisms for sharing norms, resolving conflicts across domains, negotiating changes and verifying compliance.
The evaluation agenda includes standardized multi-agent benchmarks and dynamic “Agent Gym” environments, a separate line of work from Google Research’s EnvHarness agent-training environments. The proposed open-source sandboxes would test cascading interactions among agents over longer periods instead of treating isolated actions as sufficient evidence of safety.
Independent security guidance points to the same pressures around evaluation and authorization. NIST describes agent hijacking as indirect prompt injection caused by combining trusted instructions with untrusted task data. After running five injection tasks 25 times each, NIST said average attack success rose from 57% on a single-attempt measure to 80% after repeated attempts, showing how one run can understate risk in probabilistic systems.
OWASP defines excessive agency as harmful action enabled by unexpected, ambiguous or manipulated model output combined with excessive functionality, permissions or autonomy. Its guidance recommends least privilege, authorization in downstream systems and human approval for high-impact actions.
The opened Google article and publication record provide no deployment results for the proposed policy engine, implementation metrics or completed Agent Gym benchmark results. They also do not identify the workshop report as peer reviewed or accepted at an academic venue.
More news

Google Research open-sources EnvHarness to reshape agent training environments

Apple researchers turn failed agent attempts into training insights with RLTL;DR

Google Research introduces Diffusion Controller to steer image generation
