Next upPhysical AI VC <> Founders Pitch Night #SFTechWeek @Mission Robotics
News

Google reports uneven public-health gains from Earth AI embeddings

Google Research reported that Population Dynamics Foundation Model embeddings improved selected vaccination, dengue, cholera and postpartum-depression prediction tasks, but the five studies found inconsistent gains and did not test live public-health deployments.

D
Oct 6, 2026 · 4 min read

Google Research on Oct. 6 published results from five partner-led evaluations of its Population Dynamics Foundation Model, finding measurable gains in selected vaccination, dengue, cholera and postpartum-depression prediction tasks but little or no benefit in several other settings.

The 41-page supporting preprint evaluates versions of the model’s fixed location embeddings across the United States, Canada, Mexico and the Democratic Republic of the Congo. The studies measured predictive performance rather than changes in patient outcomes, outbreak control or resource allocation, and the opened sources do not document any of the five models running in a live public-health decision system.

PDFM compresses aggregated search trends, mobility and built-environment signals, weather and air-quality data into vectors representing places. Researchers added those vectors as covariates to existing statistical or machine-learning models without task-specific fine-tuning, according to Google Research’s release. For related Google Research coverage, DataPhoenix has also examined Google’s verifiable private federated-learning system for Gboard.

The clearest vaccination result came from 146 U.S. counties within 150 kilometers of Canada. Adding Canadian contextual embeddings to U.S. embeddings raised R-squared, a measure of explained variation, from 0.159 to 0.216 — about a 36% relative gain. The Canadian data was complementary rather than a replacement: across all U.S. counties, Canadian context alone produced an R-squared of 0.013, compared with 0.373 for U.S. embeddings alone, while combining both raised the national result only to 0.381.

For cardiovascular mortality, PDFM was competitive with census covariates but did not establish a statistically significant advantage over them. In a gradient-boosted model nowcasting 2023 county deaths, PDFM recorded mean absolute error of 18.66 and root mean squared error of 46.00, versus 19.13 and 57.69 for American Community Survey inputs. Direct PDFM-versus-ACS differences were not statistically significant. In a separate Bayesian interpolation test, ACS was nominally better, with mean absolute error of 42.07 versus 44.95 for PDFM. Adding all PDFM dimensions to ACS increased variance instead of improving cross-sectional accuracy.

That comparison tests a potential timing advantage. ACS five-year estimates pool 60 months of survey observations, while the embeddings were computed from one month of data and Google says they can refresh monthly. But most retrospective evaluations used a static October 2023 embedding snapshot, so the experiments did not demonstrate that monthly updates improve predictions as local conditions change.

The dengue study covered about 2,450 Mexican municipalities from January 2020 through August 2025. Adding PDFM to TimesFM improved the one-month Weighted Interval Score by 0.0051, with p below 0.001, but there was no robust overall improvement at three or six months. Only 47.6% of municipalities improved on average at the one-month horizon, and the median municipality was slightly worse; larger gains in a subset of active-transmission locations drove the aggregate result.

For cholera, the researchers prospectively tested 89 weeks across 403 health zones in the Democratic Republic of the Congo. At four weeks, PDFM raised precision-recall area under the curve from 0.2832 to 0.3107. At eight weeks, the share of correct zones among the model’s five weekly picks rose from 0.3556 to 0.4198, equivalent to 2.10 correct zones rather than 1.78. At one and two weeks, recent case history was already sufficient and PDFM did not significantly improve performance; one-week precision among the top five picks was 8% lower, though the decline was not statistically significant. Sparse digital signals in the DRC also required six months of search-data pooling and compression to 200 dimensions, limiting how quickly the representation could reflect acute shocks.

In individual postpartum-depression risk prediction, the measured gain was small. Among 332,970 respondents in the CDC Pregnancy Risk Assessment Monitoring System, PDFM increased area under the receiver operating characteristic curve by 0.0020 from a 0.62 baseline in states represented in training and by 0.0038 in unseen states. The gain did not persist when models trained on 2012–2019 births were tested on 2020–2021 births. PDFM did not replace individual income and insurance data, added little beyond those records and closed none of 21 pre-registered demographic performance gaps.

Screening simulations exposed an operational tradeoff rather than an unqualified improvement. At a fixed capacity to screen 20% of mothers, the PDFM model identified an estimated 5,640 more rural cases annually while generating 23,474 more false alarms. At a target of 80% sensitivity, it avoided 17,723 false alarms but flagged 1,722 fewer rural cases. The authors wrote that choosing between those outcomes is a policy decision, not a property of the model.

The manuscript was submitted as arXiv version 1 on Oct. 5 and has not been independently replicated in the supplied evidence. Google made the commercial dataset powered by PDFM, Population Dynamics Insights, available in Preview in April. Google says it covers 17 countries and updates monthly; its research release says academics and public-health researchers can request no-cost access for select, non-operational research uses.

More news