Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Apple researchers introduce SCLATE for continual-learning agent evaluation

Apple researchers introduced SCLATE, a shared event-scheduling substrate for long-horizon agent benchmarks and post-training.

D
Sep 30, 2026 · 2 min read

Apple researchers have introduced SCLATE, an execution substrate for evaluating and post-training continual-learning agents across sessions and simulated periods spanning days or weeks. Its shared scheduler lets researchers test combinations of models, agent harnesses and memory systems without building a separate scheduling loop for every pairing.

The September 26 preprint describes an open event scheduler: benchmark and agent adapters can both register events on a single timeline. These can include scenario tasks, environment changes, session boundaries, background jobs, memory consolidation, grading and state handoffs. The paper does not identify a public code repository or software license, so its phrase “open event scheduler” should not be read here as confirmation of an open-source release.

A hybrid simulated clock advances in real time while an agent or background job is working, then jumps to the next scheduled event when the system is idle. This compresses long scenarios without advancing the clock beneath an active tool or job.

According to the paper, SCLATE can run existing harnesses without source-code changes. Integration may still require an adapter, configuration files, environment variables, a capture-proxy endpoint and libfaketime. That is the limited sense in which the authors use “unmodified.”

The researchers ported seven public benchmarks to the substrate: MetaClaw, GAIA-2, AppWorld, LongMemEval, PersonaMem, SWE-Gym and SWE-bench Verified. They evaluated 10 harness-and-memory configurations across 10 models, using Claude Code, Hermes and Codex harnesses with native or disabled memory and, where applicable, Mem0, Hindsight or GBrain. The work adds to Apple research on measuring reset-free reinforcement-learning failures, while SCLATE focuses on coordinating long-horizon evaluation and training runs.

The performance findings are author-reported and have not been independently reproduced in the reviewed evidence. The paper says external memory did not reliably outperform a harness’s native memory across the comparisons. The authors also report that SCLATE delivered all 13,050 scheduled GAIA-2 environment events in their validation, with 98.7% of checkable events firing at the delay specified by the scenario.

For post-training, SCLATE routes model calls through a proxy inside each execution container. The proxy records prompt tokens, completions and log probabilities, tags calls by component, and assembles token-level trajectories. A rollout engine combines those trajectories with grading verdicts for external post-training frameworks while managing agent and memory state across container runs.

The paper lists Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang and Manjot Bilkhu as authors, all affiliated with Apple. The arXiv record shows that the preprint was first submitted on September 26, 2026, for ICLR 2027. The reviewed sources do not establish a peer-review decision or independent reproduction.

More news