News

BMW runs daily cloud-cost anomaly detection across 14,000 accounts

BMW Group's CLEA system scans hundreds of thousands of account-service cost series every day, using forecasts and filters to flag anomalies while account owners decide whether spending changes were intentional.

D
Sep 22, 2026 · 3 min read

BMW Group runs daily cloud-cost anomaly detection across more than 14,000 accounts through CLEA, an in-house FinOps system built on AWS with Reply. According to a case study co-written by AWS, BMW Group and Data Reply, the system analyzes the previous day’s spending and emails account owners when costs depart far enough from an expected pattern.

The workload described in the case study starts with roughly 3 billion raw billing rows across 500 columns per month. CLEA aggregates billing exports from AWS and other unnamed cloud providers into daily cost for each account and service, producing hundreds of thousands of active account-service time series. Before analyzing the previous day’s data, the pipeline waits for the AWS Cost and Usage Report to finish delivery so it does not alert on a partial day. The case study does not identify the other cloud providers; managed private Azure-to-AWS links provide separate context on cross-cloud operations, not evidence about the CLEA provider mix.

For every account-service pair, CLEA trains a separate Prophet forecast on 365 days of daily history, with additive seasonality. A daily cost outside the model’s confidence interval becomes a potential anomaly. In practice, that means each service in each account is measured against its own recent spending pattern instead of a company-wide limit. Prophet’s documentation cautions that its uncertainty intervals assume future trend changes resemble historical ones, making the interval a screening signal rather than proof that spending is wasteful.

Several filters determine whether a candidate warrants an email. CLEA excludes services averaging less than $0.10 over the prior three days, services with fewer than 10 days of history and specified irrelevant charge types. A candidate normally must differ from expected spending by at least 40% and clear an absolute-impact floor based on the account’s trailing three-month average spend. Those floors range from more than $300 for accounts below $100,000 to more than $1,000 for accounts above $500,000.

BMW, AWS and Reply said legitimate daily variance led them to use a 60% deviation threshold for AWS Glue, Amazon Athena and Amazon EC2. Accounts on a reduced-sensitivity list must exceed three times the standard thresholds. A sustained increase may be flagged for its first few days, then absorbed into the baseline as the rolling training history catches up. That makes the approach strongest at catching spikes.

AWS Step Functions orchestrates the run. A preparation function writes the active-account list to Amazon S3, then a Distributed Map fans the work out to as many as 500 concurrent Lambda workers. AWS’s documentation for Distributed Map confirms that the mode runs parallel child workflows and supports configurable concurrency and failure thresholds. BMW, AWS and Reply said the roughly 14,000-account cycle finishes in about 20 minutes.

Each worker reads an account-level Parquet file and writes anomaly results as JSON to Amazon S3. AWS Glue consolidates the files into a daily Parquet dataset, Athena exposes the results, and dbt applies threshold, date-range and alert-labeling logic. A separate alert process deduplicates active anomalies before sending account owners the affected service, expected and actual spend, absolute impact, percentage deviation and accumulated impact from concurrent anomalies.

The co-authors put the system’s failure tolerance at five accounts per run, equivalent to a required success rate of about 99.96%, and estimated compute cost at about $50 per month. Those deployment figures, along with the 14,000-account scope and 20-minute runtime, were not independently audited in the public evidence reviewed for this story. The case study also did not disclose alert volume, false-positive rates, realized savings or remediation time.

CLEA has expanded since an earlier AWS and BMW account of its dashboard phase described coverage of 4,500 AWS accounts and a May 2023 dashboard launch. The current anomaly system automates detection, but not the final judgment. Its co-authors said CLEA can observe operations, usage types and costs but cannot determine intent. Account owners decide whether an increase was planned, and their feedback is used to tune the alert thresholds.

More news