Offered load ≈ 87%

Pull-based scheduler

Old · random node polling
CPU utilization
0%
Queued jobs00 vCPU
Loading the fleet…

Centralized scheduler

New · selective draining
CPU utilization
0%
Queued jobs00 vCPU
Loading the fleet…
Jobs Holds a reservation Draining host
Time 0.0 s
Queue waitPull-basedCentralized
Rolling · 2 min history
Shared scale · seconds

2 vCPU

4 vCPU

8 vCPU

16 vCPU

32 vCPU

Percentile tables · last 1 min

2 vCPU

Seconds Pull Centralized

4 vCPU

Seconds Pull Centralized

8 vCPU

Seconds Pull Centralized

16 vCPU

Seconds Pull Centralized

32 vCPU

Seconds Pull Centralized

Jobs started in the preceding 60 simulated seconds; queued jobs are counted separately. With fewer than 10,000 starts, empirical p99.99 equals the longest observed wait.

Incoming workload

Both fleets receive the same jobs. Changes apply to new arrivals; running jobs finish normally.

Job sizeJobs / minuteMean duration (min)
Offered load at perfect packing86.9%

1,779 / 2,048 average vCPU · 269 vCPU headroom

Arrival rates±10%
Mean durations±0.1 min

Forecasts assume perfect packing of average demand. Bursts and scheduling can add waiting below 100% load.

View both schedulers side by side, or focus on the pull-based or centralized scheduler alone. At least one fleet is always visible. Hidden fleets keep running on the same timeline; switching views preserves jobs, settings, and wait history. Charts and tables show only the visible schedulers. Fleet size is 64, 81, or 100 nodes, arranged as an 8×8, 9×9, or 10×10 grid. Changing it restarts all policies with the current workload and scheduler settings. Arrival rates do not scale with capacity. Reserved pool counts are capped to fit the selected fleet size. This continuous illustrative simulation sends identical random arrivals and job durations to three scheduling policies. The old pull-based scheduler is on the left; the new centralized scheduler is on the right. Selective draining uses centralized scheduling with a choice of tightest fit or aligned finishes for placement. Aligned finishes scores the 10 tightest eligible occupied hosts using the assigner’s vCPU-weighted completion-misalignment cost. It compares the incoming job’s predicted duration with the latest predicted finish among the host’s current jobs. Empty hosts are a fallback when no occupied host fits. Estimates use configured mean runtimes and elapsed time, with half the mean remaining for overdue jobs; they never use sampled completion times. The left fleet models the old Redis pull behavior: each node polls independently at random intervals averaging 0.25 simulated seconds. Each node takes the oldest fitting job until its oldest eligible job has waited longer than the configured drain threshold. Then it holds capacity for that job; smaller jobs wait behind it. The default threshold is 0 seconds (immediate draining). Never (∞) disables draining and keeps taking fitting jobs, which can starve larger jobs. Wait age is measured from arrival in simulated seconds, including time spent behind other jobs. No scheduler selects a host or optimizes packing. Yellow jobs hold active drain reservations of any size; pale yellow hatching marks hosts draining for those requests. A reserved job takes an eligible host using the selected placement algorithm and immediately releases its original reservation; older reservations retain priority. Centralized draining can select the least-full eligible host with random tie-breaking, estimate which host will have enough CPU free soonest from running jobs' mean durations, or be disabled. Hosts running 32-vCPU jobs are excluded from drain targets. Reserved openings go to the oldest waiting job of the same size. After a reservation is used or displaced, remaining reservations are reassigned to the best targets in arrival order. Jobs within a host are rearranged as needed to keep each job in one contiguous box. Whole-fleet draining pauses other placements while blocked jobs wait. The left-hand reserved-node toggle compares permanent pools for 16- and 32-vCPU jobs instead. Other job sizes cannot use those pools; large jobs can also use shared hosts. Each polling node considers the oldest job eligible for its pool and applies the same drain threshold. Pool counts, arrival rates, and mean job durations can change while running jobs finish naturally. The workload dialog forecasts average vCPU demand relative to perfect packing, including shared-capacity limits under reserved pools. Job and drain-target numbers are hidden; dedicated pool labels identify the host restrictions. The latency timelines plot the selected nearest-rank percentile (p75, p90, p99, or p99.99), over two minutes of history. The look-back window applies to both charts and tables: 30 seconds, 1 minute, 5 minutes, 15 minutes, or the entire simulation. Each chart point summarizes jobs started within that window at that point in time; entire-simulation points accumulate starts since time zero. Each table shows the selected window’s average, p75, p90, p99, and p99.99, plus started and still-queued counts. Jobs still queued are excluded from wait percentiles. Empty windows show a dash. With fewer than 10,000 starts, nearest-rank p99.99 is the maximum observed wait. All job sizes start with a two-minute mean duration; playback speeds range from 1 to 100 times simulated time, starting at 5×. Incoming jobs move slowly from the left, with three seconds of playback time to cross the queue row. A separate simulation predicts their landing positions. They become eligible, count as queued, and can hold reservations only at their arrival time. The animation looks ahead to account for departures during the approach; schedulers cannot see future arrivals. Placement uses straight paths and no motion trails. Reservation cues appear together on the waiting job and its host. Poisson arrivals and lognormal job durations are illustrative assumptions, not production measurements.