News

Ai2 swaps GPU job priorities for time budgets

Ai2 has replaced priority-based scheduling across its research GPU clusters with time budgets, fair-share job ranking and preemption after a protected runtime.

D
Oct 9, 2026 · 3 min read

Ai2 has replaced priority-based GPU scheduling across its research compute environment with time budgets, hierarchical fair-share allocation and a time-slicing contract. Instead of tying access decisions to priority labels on individual jobs, the system gives research managers explicit control over how much GPU time each project receives.

Ai2 said that, over a 30-day test, teams received 98% of the GPU hours they were owed. Thirteen of 15 team allocations received at least 95%, while the lowest received 90%. These are internal results reported by Ai2; the institute did not publish raw telemetry, complete sample sizes or a reproducible dataset.

The institute manages thousands of Nvidia H100, B200 and B300 GPUs in clusters ranging from 88 to 1,024 GPUs for about 150 internal researchers. Ai2 said submitted demand at any moment is two to three times available capacity. Under its previous system, every scheduled workload eventually carried HIGH priority, users sometimes parked no-op jobs to retain access, and jobs could opt out of preemption.

The replacement system gives managers proportional budgets of GPU time that flow down through Ai2’s research hierarchy. A workload must draw from a budget to gain protection from preemption. By default, the scheduler looks back over seven days, placing jobs from groups that have used less than their budgeted share ahead of those from groups that have used more.

Fair-share scheduling itself is well established. Apache Hadoop’s Fair Scheduler distributes resources among hierarchical queues and can weight their shares, while Slurm’s Fair Tree algorithm ranks users through an account hierarchy using assigned shares and prior usage. Ai2’s implementation applies that approach to a management-set tree of research budgets.

Each Ai2 job also declares the minimum runtime it needs to make useful progress and whether it can resume later. The job is protected for that minimum runtime. Afterward, the scheduler can preempt and requeue resumable work to rebalance capacity. Ai2 capped the protected minimum at eight hours. Jobs that declare zero receive unallocated time, which is not charged to a budget but can be preempted immediately.

Ai2 said overall cluster occupancy held at 98% before and after the change, with unallocated work accounting for 18% of delivered GPU time. On its largest H100 cluster, median queue wait fell from five minutes to 24 seconds, while the 90th-percentile wait dropped from 2.8 hours to 1.8 hours. For small debug jobs, the reported 90th-percentile wait fell from two hours to 30 seconds in production. Ai2 cautioned that the baseline sample of debug workloads was smaller and therefore more variable.

According to Ai2, automatic draining after a job’s protected runtime cut repairs requiring a person in the loop by 74%. The change also left interactive sessions vulnerable to preemption after eight hours, forcing researchers to rebuild volatile working state. Ai2 said it plans a CPU-only development cluster and restorable sessions to address that problem.

Ai2 also recently released an open-weights model for scientific reports.

The institute is still investigating whether minimum-runtime protection fragments capacity and lengthens waits for the largest jobs. Its post did not provide a result or measurement for that investigation.

More news