Published Sep 19, 2026• Updated Sep 28, 2026

Building a Hang-Proof Playwright E2E Skill

How a loop of infinite timeouts turned into three nested deadlines, a watchdog, and a repeatable method for writing agent skills that survive production.

Building a Hang-Proof Playwright E2E Skill

How a loop of infinite timeouts turned into three nested deadlines, a watchdog, and a repeatable method for writing agent skills that actually hold up in production.


If you run AI agents against real infrastructure, you know the deal: you ask the agent to "verify the app in a browser", it starts driving Playwright, and ten minutes later it's still going. Stuck. Not crashed — stuck, in a loop of timeouts it genuinely believes are app bugs. You interrupt the task, resume it, and hope it takes a different path this time.

That was my experience every few runs with my playwright-e2e-verify skill — the instruction pack my coding agent reads before every end-to-end browser verification. This post is about the redesign that made the hangs impossible, the technical decisions behind it, and the step-by-step method I now use to build skills that survive contact with reality.

First, what is a "skill"?

A skill is just a Markdown file the agent loads when a task matches it: a name, a trigger description, and a body of instructions — commands, code snippets, rules, traps. No magic, no plugins in the runtime sense. The value is entirely in which words you put there and whether they're true.

Which is why the interesting question isn't "how do I write a skill?" but "how do I write one whose instructions survive a live run?"

The stage: one agent, two operating systems

Before the bugs, the machine itself is worth describing, because half the traps in this skill exist because of it — and an agent skill that doesn't encode its own host is a skill that fails on the first non-trivial task.

The setup is a common one and deceptively simple on paper:

  • Windows is the host — and the agent runs here, in the desktop app. Its tool calls execute on Windows: PowerShell is the first (and sometimes last) shell every command touches.
  • Ubuntu 24.04 runs in WSL2 — and this is where the projects live: repos, dev servers, SSH keys, dotfiles, the Linux-side toolchain. Real work happens on this side; the agent just has to cross a border to reach it.
  • Node exists twice, with different versions and different global modules. The browser-automation stack (playwright-cli and its bundled playwright-core) is installed on the Windows side; the projects on the WSL side have their own node_modules, compiled for their own Node.

One agent, one task, two operating systems, two toolchains. Every non-trivial operation crosses a boundary — and every boundary is an opportunity to fail in a way that looks like an app bug. The caveats we had to overcome, all of them now encoded as verified rules in the skill (and in its sibling infrastructure skills):

1. Nothing crosses the WSL boundary for free. Environment variables set on the Windows host do not appear inside WSL — verified with a five-second probe (echo ${#VAR} returns 0). Passing them requires an explicit pattern on the command line, and writing that pattern in PowerShell has its own trap: \" is not an escaped quote in PowerShell (its escape character is the backtick), so the naive wsl -e bash -lc "env VAR=\"$env:VAR\" ..." corrupts the value with literal backslashes and bash dies with unexpected EOF while looking for matching '"'. The verified working pattern: a single-quoted bash string with PowerShell string concatenation.

2. Nested quoting through two shells mangles anything complex. PowerShell → WSL → remote shell is three quoting contexts deep. Go templates like {{.Names}}, nested quotes, %{...} curl formats — all of them get chewed. The reliable pattern is boring on purpose: write the script to a temp file, then pipe it over stdin (wsl -e bash -lc "ssh ... 'bash -s' < /mnt/c/.../script.sh"). Inline cleverness loses; the file always wins.

3. Credentials and tools are per-OS. The SSH keys for the servers live in WSL's ~/.ssh, not in the Windows profile — running ssh from PowerShell directly fails with Permission denied (publickey,password), an error that points at the server while the problem is the client side. The lesson generalizes: when a tool "doesn't work", first ask which side of the boundary it lives on.

4. Same tool, two worlds, incompatible binaries. A native module (think better-sqlite3) compiled under the WSL Node refuses to load under the Windows Node with NODE_MODULE_VERSION mismatch — and vice versa. "It's installed" is not an answer; where and for which Node is the answer.

5. Paths and names shift at the border. Windows paths become /mnt/c/... inside WSL. And under WSL's mirrored networking, localhost can resolve to IPv6 and time out while 127.0.0.1 works — so every URL the skill constructs uses the numeric address.

6. Not every shell command exists on both sides. The timeout 150 node ... belt-and-suspenders guard is coreutils — it simply doesn't exist on the Windows host. The watchdog deadline has to be the one that's guaranteed everywhere.

The general lesson that came out of all six: the skill must encode the machine, not just the app. An E2E verification is never just "drive a browser" — it's "drive a browser from this specific environment, across these specific boundaries, without trusting any of them to behave." Every cross-boundary operation got a verified, copy-pasteable pattern in the skill, with the failure mode it replaces documented next to it. When the environment is complicated, the difference between a skill that works and one that doesn't is whether the environment is described as precisely as the task.

The symptom: an agent that wouldn't give up

The failure mode was never a crash. It was worse: an optimistic crash loop. The agent would:

  1. Launch a browser step and get a timeout or a null evaluation result
  2. Conclude the app was broken
  3. Retry — because that's what good agents do
  4. Get the same timeout
  5. Retry again…

Each individual step was "working". The loop was the bug. And since every step burned a 30-second Playwright default timeout, the run felt productive while going nowhere. The human — me — had to watch this happen, interrupt the task, and resume it. That cost is what turned a nice-to-have improvement into a redesign.

Root cause #1: a stateless agent driving a stateful browser

The skill originally drove playwright-cli, one command per tool call. That works great for a human in a terminal — the CLI keeps its browser alive inside the session that launched it.

But an agent environment is not a terminal. Every command runs in a fresh shell. The browser dies with the shell that spawned it. The next call silently opens a new browser on about:blank, the page state is gone, and every eval returns null against an empty document. The agent sees null and thinks: the selector must be wrong, let me try another one. Then another. Then a reload. Then another browser.

The fix is architectural, not tactical: stop driving the CLI across calls. Drive playwright-core from a single Node script instead.

js
const { chromium } = require('<global-root>/@playwright/cli/node_modules/playwright-core');

const browser = await chromium.launch({ channel: 'chrome', headless: true });
const context = await browser.newContext({ viewport: { width: 1600, height: 1100 } });
const page = await context.newPage();

One process, one browser, full state for the whole run. No shell quoting between steps — which on a Windows host is its own special circle of pain — and, crucially, a single place to enforce timeouts.

Note the path: the CLI bundles playwright-core as a dependency, so I'm not adding a second installation. The CLI still has its place — quick one-shot probes ("what does this page show right now") and interactive file-chooser modal handling — but anything multi-step goes through the script.

Two CLI quirks are worth encoding in the skill anyway, because they masquerade as app bugs when you do fall back to the CLI:

  • Chained statements fail: eval "a(); b()" dies with Passed function is not well-serializable. One eval per call.
  • Return values must be serializable: a DOM node or a CSSStyleDeclaration cannot cross the boundary. Extract primitives — strings, numbers, plain objects built inside a single expression.

Root cause #2: no timeouts anywhere

This is the embarrassing one. Playwright ships defaults (30s per action), and I had never set a single one. Thirty seconds sounds fine until an agent multiplies it by fifteen retries.

The redesigned skill enforces three nested deadlines:

LayerTimeoutWhat it catches
Per-action8sOne missing selector, one hung click
Watchdog120sThe whole script wedging on anything
Shell / tool call~150–180sEven a hard-crashed Node process

At the top of every script:

js
page.setDefaultTimeout(8000);
page.setDefaultNavigationTimeout(15000);

And the watchdog — the piece I now consider mandatory in any agent-driven long-running script:

js
const WATCHDOG_MS = 120000;
const timer = setTimeout(() => {
    console.log('WATCHDOG: run exceeded ' + WATCHDOG_MS + 'ms, aborting');
    console.log('RESULT ?/' + checks.length + ' (incomplete)');
    checks.filter((c) => !c.pass).forEach((c) => console.log('FAIL  ' + c.name));
    process.exit(2);
}, WATCHDOG_MS);

The watchdog doesn't just kill the process — it prints partial results. An aborted run that says "6 checks passed, hung on check 7, here's the transcript so far" is infinitely more useful than silence. Which is also why each check logs PASS/FAIL as it happens, not in a summary at the end: if the watchdog fires, the partial transcript shows exactly how far the run got.

One platform wrinkle I hit while testing: the belt-and-suspenders shell-level guard timeout 150 node e2e-test.js is coreutils — and on this machine the script runs on the Windows host, where it doesn't exist. So the effective ceiling there is the watchdog plus the tool-call timeout itself. The rule that falls out of this: keep the watchdog strictly shorter than the tool-call timeout, so it always fires first and you get the partial report instead of a generic tool timeout. (This is caveat #6 of the two-OS architecture — the deadlines have to be designed per platform, not assumed.)

The blocking constructs blacklist

Alongside the deadlines, the skill explicitly bans the constructs that hang forever:

  • page.pause() — opens the inspector and waits for a human. An agent is nominally the driver; nobody is coming to click "resume".
  • waitForLoadState('networkidle') — fine on a static page, forever-pending on anything with SSE, websockets, or polling. Wait for a specific element instead: page.waitForSelector('#app-ready', { timeout: 8000 }), or a bounded waitForTimeout(500).
  • page.waitForNavigation() — racy and easily outlived. page.waitForURL('**/expected', { timeout: 8000 }) states both the intent and the deadline.
  • Unbounded polling — any while (!cond) needs an attempt counter, or it's a hang with extra steps.

Plus hygiene that turns out to matter more than expected:

  • browser.close() in a finally — a throwing check must never leak a browser, and an explicit process.exit() guarantees the process terminates.
  • Headless by default for agent runs — a headed browser steals the user's focus and can block on first-run dialogs (profile picker, crash-restore bubble). Screenshots capture everything either way.
  • setInputFiles over the modal dance — in a script there's no need to click a dropzone and wait for a file-chooser modal state; assign the input directly.

The rule that mattered most: bounded retries

All the timeouts above are code. The deepest fix was a behavioral rule in the skill, aimed at the agent itself:

If the same check fails or times out twice in a row, STOP. Report what was verified, what hung, and the last observed state. Never loop a failing step.

Two attempts, then report. That's it. A hung run costs more than an incomplete report — and a partial report is honest, while a loop is theater.

It pairs with an early-death detector: if the very first page.evaluate on a URL I know is live returns null or times out, the browser died — relaunch once inside the same script, and if it fails again, print and exit. No loop. And before any browser work: a one-second curl of the target. A dead server makes every Playwright step time out; you want to learn that in one second, not after burning a full run.

Assert by measurement, not by eye

A skill for an agent can't rely on screenshots as evidence — models may not view images reliably, and eyeballing is unreliable even for humans. So every assertion is a measurement:

js
const m = await page.evaluate(() => ({
    scrollWidth: document.documentElement.scrollWidth,
    innerWidth: window.innerWidth
}));
check('no horizontal overflow @390', m.scrollWidth <= m.innerWidth);

Results accumulate in a check list — { name, pass, detail } — and end in a RESULT n/m line that's both machine-parseable and human-readable. Screenshots still get taken, but for the human reading the report, not as proof. Even the report format is a rule: separate verified (what a measurement proved), not verified (what couldn't be exercised, and why), and bugs found (concrete input, observed vs expected). A green run that only covered the happy path is not a verification.

The same philosophy covers the unglamorous stuff: triage console noise before calling it a bug (ERR_BLOCKED_BY_CLIENT is the user's adblocker; analytics beacons are injected by the CDN), scan visible text for NaN/undefined/TODO placeholders (they're real defects), and — a direct consequence of the two-OS architecture — prefer 127.0.0.1 over localhost, for the reasons covered earlier.


Why a skill beats starting from scratch every time

The obvious alternative to a skill is: describe the goal to the agent each time and let it figure the browser work out on the fly. It works — the first time. The problems show up on every run after that.

Every from-scratch derivation re-rolls the dice on solved bugs. The CLI death loop from the first section is the perfect example: the fix is known, it's written down, it's one architectural decision — but an agent without the skill doesn't know that. Each fresh conversation re-derives the approach from first principles, and each re-derivation can reintroduce a bug you already paid to fix. A skill is the difference between paying once and paying every run.

The lessons compound — or they evaporate. This is the core argument. Every trap captured in the skill (the null eval, the networkidle that never fires, the hidden submit button) was paid for with a real failed run. In a skill, that cost is amortized across every future run, forever. Starting from scratch throws the investment away and silently re-buys it, run after run, at full price.

Consistency becomes possible. A from-scratch agent produces a different report every time: prose here, screenshots there, a verdict that can't be compared with last week's. The skill's contract — the check list, the incremental PASS/FAIL log, the RESULT n/m line, the verified/not-verified/bugs split — makes runs comparable and auditable. You can diff two runs of the same app and see exactly what regressed.

The human review moves from every-run to once. Without a skill, you audit the agent's behavior on every single task. With a skill, you audit the document once — and every future run inherits the review. That's the same leverage you get from a code review: expensive once, cheap forever after.

Failure is cheaper. The pre-flight curl, the bounded retries, the watchdog — these exist because the skill's author (me) already watched full runs burn on dead servers and hung scripts. A from-scratch agent rediscovers those failure modes the expensive way, live, on your time.

The honest framing: a skill is not a shortcut around competence, it's competence that persists between conversations. LLMs don't remember; the files you hand them do.

One skill, any app: making the E2E flow generic

A fair objection at this point: doesn't a skill full of specific checks only work on the app it was written for? The answer is no — if you keep a strict separation between the method and the target. The skill owns the how; every what arrives at run time. Concretely, three design rules do the work.

Rule 1: nothing about the app is hardcoded — everything is run input. The skill's entry point takes a URL (or a local port), a one-line description of what the app is, and optionally a critical flow the user cares about. That's the entire surface of app-specificity. The module path, the deadlines, the check-list scaffolding, the reporting contract — those belong to the skill and never mention a particular app.

Rule 2: the skill discovers the app instead of assuming it. This is what makes the same skill work on a React SPA, a server-rendered site, or a canvas-heavy dashboard. The script doesn't hardcode selectors — it runs a DOM census and walks what it finds:

js
// discover, don't assume
const nav     = await page.evaluate(() => Array.from(document.querySelectorAll('nav a, [role=tab]')).map(a => a.textContent.trim()));
const buttons = await page.evaluate(() => Array.from(document.querySelectorAll('button, [role=button]')).filter(b => b.offsetParent !== null).map(b => (b.innerText || b.id).trim().slice(0, 30)));
const inputs  = await page.evaluate(() => Array.from(document.querySelectorAll('input, select, textarea')).map(i => ({ tag: i.tagName, type: i.type, name: i.name || i.id })));

Then it drives the generic surface: click through the discovered navigation, exercise the discovered controls with synthetic data, toggle what toggles, submit what submits (via the app's own handlers when the submit button is hidden — that trap again). Every app gets a different walk, from the same walking instructions.

Rule 3: the defect detectors are universal; the app-specific assertions are per-run. Some checks are app-agnostic by nature and live in the skill permanently:

  • console errors and pageerrors (after noise triage)
  • horizontal overflow at desktop and phone widths
  • placeholder rot: NaN, undefined, TODO, stale product names in visible text
  • broken images, dead links, blank screens
  • the stale-page-after-redeploy trap (reload before asserting on fresh assets)

App-specific behavior — "the elevation chart should end where the data ends", "importing this GPX shows three segments" — is generated for that run and wrapped in the skill's scaffold: same check list, same deadlines, same watchdog, same report. If a critical flow is worth keeping, it becomes a project-local companion file the skill knows how to consume — the skill stays clean, the app keeps its specifics.

The test of genericity is simple: point the skill at an app it has never seen and watch the cold run. That's exactly what the verification run below was — the skill had never met the target app; it discovered the navigation, the controls, and the layout live, and produced a full verdict in seconds.


The method: building the skill, step by step

The redesign above didn't come from a burst of inspiration. It came from a repeatable loop — the same one I now use for every skill. Here it is, step by step.

Step 0 — Scope the skill to one job, and write the trigger honestly

One skill, one job. This skill exists to answer: "does this app actually work in a real browser?" — nothing else. The trigger description lists the phrases that should summon it ("verify E2E", "check the app in a browser", "test all the pages") and, just as important, implies what should not summon it. A skill with a fuzzy scope accumulates unrelated rules until nothing in it is findable, by the agent or by you.

Step 1 — Capture failures as they happen, verbatim

This is the whole game. Every time the agent misbehaves, don't fix the task — fix the skill, and record the failure first: the exact command, the exact error, the wrong conclusion the agent drew from it. "The eval returned null and the agent blamed the selector" is a lesson. "Quoting is tricky" is not. I keep these raw, with the real error strings, because error messages are what the agent will actually see next time — the skill must match on those, not on a paraphrase.

Step 2 — Draft v1 from the traps you already know

Turn the captured failures into a first draft. The draft structure that has held up:

  1. The one architectural decision (here: one Node script, not CLI calls), stated first with the reason — because if the reason is there, the agent can generalize when it hits a variant.
  2. The verified snippet for the common case — copy-pasteable, with real paths.
  3. The traps, each as symptom → cause → fix. Symptom first, because that's what the agent matches against.
  4. The reporting contract — what the run must output, so the human can audit it.

Rule of thumb for the draft: every trap needs its real error message in the text. If you can't quote the error, you haven't captured the trap yet — go back to Step 1.

Step 3 — Verify every rule live before it enters the skill

This is the discipline that separates a skill from a guess. Never write a rule you haven't run. Every command in the final skill was executed on the real machine first: the module path was ls'd, the broken quoting pattern was deliberately reproduced to confirm the exact error text, the env-var inheritance assumption was tested with a one-liner (echo ${#VAR} returning 0 — assumption killed in five seconds). Where a rule was platform-dependent, both platforms were tested and both are documented.

The cost is minutes. The payoff is that the agent never has to rediscover any of it — and, critically, never has to trust a rule that might be folklore.

Step 4 — Design the failure modes out, don't document around them

With the traps known, ask for each one: can I make this impossible instead of unlikely? That's where the three nested deadlines, the watchdog, the finally-close, and the headless default came from. "Don't get stuck" is a wish. "8s → 120s → 180s deadlines, watchdog prints partial results, two-strike rule" is a design. Code > prose, everywhere a construct can replace an instruction.

Step 5 — Test the skill with a fresh run on a real target

Then the moment of truth: hand the skill to the agent cold — no memory of this conversation, just the document — and point it at a real, deployed app. Watch the behavior, not the output: Does it do the pre-check? Does it set the timeouts? Does it stop after two failures or loop?

The run I did against a live app returned a clean sweep:

PASS page loads (HTTP 2xx/3xx)  [status=200]
PASS page has a title  [[app title]]
PASS page has visible text  [970 chars]
PASS no raw placeholders (NaN/undefined/TODO)
PASS no horizontal overflow @1440+  [scrollWidth=1590 innerWidth=1600]
PASS no horizontal overflow @390  [scrollWidth=380 innerWidth=390]
PASS page exposes interactive controls  [9 found]
PASS no app-origin console errors
RESULT 8/8

Seconds, not minutes. No hang, no interruption, no resume. The watchdog never fired because nothing got near it — which is the point.

Step 6 — Harvest the test back into the skill

The test run is also a failure harvest, because the meta-failures show up there: a documented pattern that was subtly wrong (a shell-quoting example that works in bash but not in PowerShell), an assumption that held on one platform only, a guard that doesn't exist on the host OS. Each one became a new verified rule in the next version. Then version the skill as an immutable update — same id, new revision — so there's always one current truth and a history of what changed and why.

Step 7 — Keep knowledge in exactly one place

The final structural rule: when two skills need the same knowledge, one owns it and the other points to it. Fix it once, everything inherits. The fastest way to rot a skill library is the same snippet pasted in three files, drifting.


The meta-lesson

Agents don't need to be smarter to stop hanging. They need instructions that assume the worst case and make the worst case terminate with a report.

And skills are code, in the way that matters: they should be versioned, their rules should be verified before they ship, their failure modes should be designed out where possible, and every bug found in the wild should become a rule — with the real error string attached — instead of a war story.

The loop is: observe a real failure → reproduce it deliberately → encode the fix as a verified rule → test the skill cold → harvest the test back in. Run it a few times and the skill stops being a pile of tips and becomes what it should have been all along: production code that happens to be written in prose.


The watchdog, the deadline table, and the two-strike rule from this post ship at the top of the skill — the first thing the agent reads, before it writes a single line of test code.

WRITTEN BY

Luca

Exploring the future of quality assurance and testing automation through deep technical insights.