Ai Driven TestingBeginner

Self-Healing Test Automation: How AI Fixes Broken Selectors

How self-healing test automation actually works, which broken selectors it can recover, and the failures it cannot fix.

ObserveOne Team
11 min read

A test suite that breaks every time someone renames a CSS class is not protecting you. It is billing you. The pipeline goes red, somebody spends twenty minutes finding out that nothing is actually wrong, and the next red build gets a little less attention than it deserved.

Self-healing test automation is the attempt to stop paying that bill. This guide covers what it actually does, what ObserveOne's implementation does specifically, and the failure modes nobody advertises.

What self-healing test automation is#

A self-healing test can recover when the element it was looking for is still on the page but no longer matches the selector that used to find it.

That is a narrower claim than most of the marketing makes, and worth holding onto. Healing relocates an element. It does not understand your product. Everything else a test does, the assertions, the setup, the ordering, is untouched by it.

// Breaks the moment someone renames the class
await page.click(".submit-button-v2");
// Describes what the thing is, so there is something to recover from
await page.getByRole("button", { name: "Submit" });

The second version is not self-healing on its own. It is the version a self-healing system has a chance with, because it records intent rather than a coordinate.

What actually breaks a selector#

Before deciding whether healing is worth having, it helps to know what proportion of your red builds it could plausibly touch. In practice, broken locators come from a small set of causes:

A class or id gets renamed. This is the case everyone thinks of, and it is the easiest to recover from, because the element is unchanged apart from one attribute.

The DOM gets restructured. A wrapper div appears, a button moves inside a new toolbar component, and every positional selector below that point is now pointing at the wrong subtree. Recovery here depends on whether the test described the element or its address.

The visible text changes. "Submit" becomes "Place order". A locator built on accessible name breaks, but the element is trivially findable by role and position in the form.

The element became conditional. It renders only for accounts with a particular flag, or only after a first login, and your test account drifted. Nothing is broken. The element is genuinely not there.

The page is slow. The element arrives 200ms after the assertion looked for it. This is not a selector problem at all, and no amount of healing fixes it.

Only the first three are healable. The last two are the ones that produce most of the frustration, and a tool that claims to fix all five is describing something it does not do.

How ObserveOne heals a test#

Healing runs as a separate step after a run has finished, not as a fallback inside the locator call. You start it from the run page: a failed test gets its own fix button, and a run with several failures gets a bar offering to fix all of them. A run that is still going cannot be healed, because a test that has not finished may still pass on its own.

What happens after that is mostly a sequence of refusals, and the refusals are the interesting part. They are what stops a healer from quietly turning a real regression green.

It starts from the error text rather than from the page. Playwright names the failing locator in its own output: on the Locator: line when an expect fails, in the call log when an action times out, and in a strict mode violation that resolved to several elements. ObserveOne reads the locator out of that text, which needs no browser and no model. If the error names no locator at all, say a bot challenge or a bare navigation timeout, healing refuses before it spends anything. There is nothing to repair, and the only way to make that run pass would be to weaken the test around it.

It repairs one locator, never a file. The healer gets the broken expression, the error, the handful of lines around it in the spec, and a list of elements that were actually on the page. It returns one replacement expression. It never sees the spec as a file, so there is nothing there for it to restructure: it cannot reformat imports, invent a helper, or nudge a timeout on the way past. That is a property of the shape of the request, not a rule we audit afterwards.

It looks at the page as the failing run left it. The replacement is picked against the accessibility snapshot Playwright recorded at the moment of failure, read back out of the run's trace when a separate snapshot was not saved. That detail matters more than it sounds. The trace holds the page after the steps that led to the failure, with the dialog the test opened still open on it. Reloading the entry URL would show a page that never reached the failing step. If the captured page turns out to be an interstitial, a bot check rather than your site, healing refuses. Re-aiming your selector at a challenge page is exactly the sort of green build the feature exists to prevent.

Ambiguity is a refusal too. If the failing locator appears more than once in the spec there is no way to tell which call broke, so it stops. If it cannot account for every selector in the file, it stops there as well, rather than editing the ones it could read and leaving the rest unexamined.

Then the harder check, the one most tools skip: a replacement that aims somewhere else. Swapping a checkout button for a cart button is a selector-only change, so no diff check can catch it. ObserveOne reads the role and the accessible name each expression is aiming at and rejects the pairs that disagree. The healer also has a way to say the element is not on the page at all, which comes back as a refusal rather than a nearest-match guess. Without that option a model picks the closest thing it can see, which is how a product link becomes a "Refresh Page" link.

Last, it checks the patch is selector-only before writing it. The edited spec is compared against the original, and anything beyond the locator substitution fails the write. The new locator is linted for shapes that could never match anything, because a heal that pins a failing test red forever is worse than no heal at all.

Every one of those refusals comes back with the reason it fired, and the credit the heal held is refunded. A tool that never declines is guessing somewhere and not telling you.

What a heal gives you, and what it does not#

A finished heal is a saved edit to your test. It is not a verified fix. Nothing re-runs the test inside the heal loop, and the product says as much: when healing finishes, the message is that the tests were healed and you should re-run to verify.

I would rather that line were louder than it is, because it is the most important sentence in the feature. The check is yours. Re-run the suite. A heal that was right stays green with nothing further from you. A heal that relocated to the wrong element usually breaks on the next assertion downstream, which is one more reason to keep assertions specific rather than checking that something, anything, rendered.

What you get to look at in the meantime is the diff. A completed heal returns both the original script and the healed one, and the run page renders them as a line by line diff with the unchanged context collapsed. For a good heal that is two lines, one removed and one added, and you can tell in about three seconds whether the new selector is pointed at the same thing.

Two properties of that record are worth knowing before you lean on it.

The heal's progress messages and its diff are kept for roughly ten minutes after it reaches a final state, then dropped. It is a review window, not a history you can come back to next sprint.

And the edit replaces the stored script rather than stacking up a version history you can browse in the product. If you want a durable record of what a heal changed, read the diff while it is on screen, or keep your specs somewhere that versions them for you. The run's trace stays the ground truth for what actually executed.

Writing tests that can be healed#

Healing works on what the test recorded. A test that recorded nothing but a path cannot be recovered, because there is no intent to reason about.

// Nothing to work with. This is an address, not a description.
page.click("/html/body/div[1]/div[3]/button");
// Role and accessible name both survive a restyle
page.getByRole("button", { name: "Submit order" }).click();

Two habits do most of the work here.

Prefer role and accessible name over structure. getByRole, getByLabel and getByTestId describe the element in terms that survive a CSS refactor. Descendant chains and nth-child selectors describe a position in a tree that your framework is free to change.

Set a test id attribute and use it for the things that have no good semantic handle.

use: {
testIdAttribute: "data-testid",
actionTimeout: 10000,
navigationTimeout: 15000,
},

The timeouts are there for a reason worth being explicit about: they are not healing, they are the thing that stops you asking healing to fix a race condition. Get waiting right first, or you will spend your heals on tests that were never broken.

What self-healing does not fix#

This is the section most tool documentation leaves out, and it is the one that decides whether you can trust the feature.

The big one: it cannot tell an intended change from a regression. If someone removes the Submit button and replaces it with a Continue button that does something different, a healing system may well relocate to the new button and go green. The test now passes against behaviour nobody asked for. That is why the diff matters. A heal is a change to your test suite and deserves the same look you would give a pull request.

Assertions are not repaired. Healing finds an element. If the assertion about what that element should contain is now wrong, the test still fails, and it should.

Flaky timing is not a locator problem either. A test that fails one run in ten because of a slow network gives healing nothing to relocate. Fix the waits.

Then there are elements that legitimately disappeared: a feature flag turned off, a plan tier that no longer includes the button, an account in a different state. The element is absent. The test is right to fail.

Of those, the ones ObserveOne declines outright rather than attempts are the timing failure and the navigation failure, because their error text names no locator, and the missing element, because the healer can see it is not on the captured page and says so. Getting a refusal with a reason on it is a better outcome than a fix you then have to disprove, and it costs you nothing.

The last one is subtler. A green suite now has two possible meanings, that nothing changed or that something changed and was absorbed, and only the diff separates them. Self-healing buys you fewer false alarms at the cost of needing to read the results more carefully, not less.

Reviewing a heal#

Treat a completed heal as a proposed change rather than a resolved incident. It is a commit somebody else wrote against your test suite, and it deserves the ten seconds you would give a one line pull request.

A selector that moved from a class or an id to a role and an accessible name is usually fine, and is often better than what you had. The case to stop on is a selector that moved to a visibly different element, even one carrying a plausible name. Watch too for a swap that went the other way, from something descriptive to something structural, which is the page telling you it no longer has a good handle where you needed one.

If the same test heals in the same place every few weeks, the fix is upstream of the healer. A component whose tests keep needing to be re-aimed is telling you its markup has no stable handle, and adding a test id to it is a five minute change that ends the cycle for good.

Where to start#

Pick the test that has broken most often for reasons that turned out to be nothing. Rewrite its locators in terms of role and accessible name, add test ids where there is no sensible role, and let healing cover the residue. That ordering matters. Healing is worth having on a suite whose locators are already describing intent. On a suite of XPath chains it is mostly an expensive way to postpone a rewrite you are going to do anyway.

Ready for AI-Powered Testing?

ObserveOne monitors your selectors 24/7 and automatically heals them when websites change. Never deal with broken tests again.

Start Free Trial