355 families · 3 websites · 6 conditions
SkillShift-Bench measures whether a web agent’s learned skills still execute after the website itself has evolved.
GitLab 162 · Magento 114 · WordPress 79 · src, src-perc, src-exec, tgt, tgt-perc, tgt-exec
Three named shifts
A skill has to be invoked, grounded, and executed. SkillShift-Bench names those three failure modes as run conditions. Execution is the steepest SR drop in every launch series.
- Contextual
- Same task after the site’s version or theme changes. Routes and chrome move; the goal does not.
- Perceptual
- Same task under altered layout, DOM, and locator evidence. The objects stay aligned.
- Execution oxide
- Same task with reproducible runtime blockers. This is where learned skills collapse on the launch board.
Execution drop, from the launch JSON
Two instrument banks, not one ranking. Lines are SR by condition. The table is the source of truth.
Mixed backbones — Table 1
| Method | src | src-perc | src-exec | tgt | tgt-perc | tgt-exec |
|---|---|---|---|---|---|---|
| ASI | 53.65 | 48.80 | 24.43 | 41.84 | 33.19 | 15.79 |
| AWM | 50.07 | 43.35 | 22.45 | 41.28 | 37.65 | 16.90 |
| WALT | 46.82 | 43.53 | 25.46 | 41.17 | 32.74 | 19.09 |
| SkillWeaver | 35.13 | 28.75 | 15.63 | 17.80 | 13.85 | 9.71 |
gpt-5-mini — Table 3
| Method | src | src-perc | src-exec | tgt | tgt-perc | tgt-exec |
|---|---|---|---|---|---|---|
| MPCR | 56.67 | 57.09 | 33.23 | 43.32 | 37.24 | 19.99 |
| ASI | 52.79 | 53.02 | 26.68 | 42.83 | 34.32 | 17.75 |
Getting started
pip install -e 'skillshift-bench[run]'
skillshift calibrate
skillshift run --agent your_module:Factory --all --out runs/latest Cite
@unpublished{chen2026skillshift,
title = {SkillShift-Bench: Benchmarking Skill Learning under Environment Evolution for Web Agents},
author = {Chen, Bo-Yu and Zhou, Zhi and Li, Yu-Feng},
year = {2026},
note = {Code, tasks, environments, and leaderboard: https://skillshift-bench.github.io/}
}