Apple study says minimal coding agent matches complex ML harnesses
Apple researchers say a minimal coding agent matched or beat four elaborate harnesses on the autonomous ML benchmarks they tested, while limiting the finding to those settings.
Apple researchers report that a minimal, single-session coding agent matched or beat four more elaborate open-source harnesses when the systems used the same frontier model backbone, hardware and time budgets on the autonomous machine-learning benchmarks they evaluated.
The result puts the value of extra orchestration into question, but only within those experiments. The paper does not establish the same outcome for other agent tasks or for sophisticated coordination among multiple workers.
Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé compared a baseline called Malena with MLEvolve, AiScientist, Arbor and ScienceFlow. Their paper describes Malena as one long-running coding-agent session built on OpenCode. It can inspect and edit files, run and debug code, submit results and manage background jobs. A brief prompt tells it to continue when a turn ends.
In controlled ablations, direct access to a coding-agent environment produced the largest measured improvement. Once that environment was available, the search, autonomy and multi-agent mechanisms tested by the authors did not produce a statistically significant gain across the intervention families they examined.
The external comparison covered 30 MLE-bench tasks with a 24-hour budget and 40 NatureBench tasks with an eight-hour budget. The authors evaluated 17 harness-backbone pairs on MLE-bench and six on NatureBench. They reported that Malena matched or beat each tested harness at every evaluated pairing with a frontier backbone.
With GLM 5.2 on MLE-bench, Malena recorded a 62.5% self-selected any-medal rate, compared with 47.1% for AiScientist, the strongest external harness on that measure. The authors said paired confidence intervals favored Malena over all four external harnesses in that comparison.
The pattern did not hold uniformly with a weaker model. Using the Gemma 4 31B backbone, MLEvolve significantly exceeded Malena on mean percentile, while the medal-rate difference was not statistically distinguishable at the study’s power level. The authors also modeled Malena’s GLM 5.2 inference cost at $12.12, about 6.2 times AiScientist’s $1.95, because Malena repeatedly resent its growing context. They said hardware cost remained larger overall.
The study is an arXiv preprint. According to the authors, expensive 24-hour runs limited them to three or four seeds per task, leaving some confidence intervals too wide for strong effect-size claims. They frame the better-supported conclusion as Malena being no worse than the compared baselines, not as a precise estimate of superiority.
Most results rely on public Kaggle tasks in MLE-bench, which the authors say may be vulnerable to pretraining contamination and had task-preparation problems. They report repairing affected tasks, running two contamination checks and adding NatureBench as a mitigation. They also identify possible tuning asymmetry in favor of Malena.
Apple has separately published SCLATE work on continual-learning agent evaluation, which addresses shared event scheduling for longer-horizon benchmarks. In the harness study, the authors say their broad claim rests almost entirely on single-worker interventions. Their multi-worker experiments covered only a small set of coordination primitives.
More news

Databricks adds branch-based restores to Lakebase Postgres

AWS adds hierarchy filtering to Amazon Quick Sight dashboards

ServiceNow CoreAI introduces AutoSynthData for enterprise-agent training
