Playwright suite audit: playwright-e2e-demo
This is what the Playwright Suite Audit report looks like. The subject is my own public demo repo, audited the same way I would audit yours. I kept every finding, including the ones that do not flatter me.
- Repo:
- RAJUSHANIGARAPU/playwright-e2e-demo @ 7824f71
- Date:
- 2026-10-01
- Suite:
- 8 tests × 3 browsers = 24, Playwright 1.61.1
- Method:
- config and code review, CI history, typecheck, lint
Summary
The code is well structured: page objects, fixtures, stable locators and no sleeps. The weak part is everything around it. The suite has run in CI once, its retries would hide flakiness, and three quarters of the pipeline time is spent downloading browsers. The biggest wins cost under an hour in total.
3 high, 4 medium, 3 low. Not run: the tests themselves (they hit a live third-party site); the 24/24 pass figure comes from the last CI run.
Since this report: H1 is fixed; a weekly scheduled CI run now keeps the badge current (d0221dc). H2 is still open.
Findings, ranked by impact
The green badge is 75 days old and nothing re-checks it
CI runs on push and pull request only. The repo has had one CI run (2026-07-18, 24/24 passed). Every test targets a third-party site that can change at any time, so the badge says nothing about today.
Fix: Add a daily `schedule:` trigger so drift in the target site shows up as a red run, not as a surprise on the next PR.
Retries turn flaky tests green without telling anyone
CI retries every failure twice. A test that fails twice and passes on the third attempt is reported as a passing build. The config comment says this absorbs network flakiness, which is how a suite drifts into the state where nobody trusts red.
Fix: Keep the retries for diagnosis, but set `failOnFlakyTests: !!process.env.CI` (supported by the pinned Playwright 1.61.1) or at least publish the flaky count from the report on every run.
75% of CI time is spent installing browsers
In the only CI run, `playwright install --with-deps` took 65 s and the tests took 22 s. The browsers and OS packages are downloaded from scratch on every run, for all three engines.
Fix: Run the job in the `mcr.microsoft.com/playwright:v1.61.1-noble` image (browsers preinstalled), or cache `~/.cache/ms-playwright` keyed on the Playwright version.
The checkout test never checks the order it places
The test fills the form, clicks through the overview step and asserts only the 'Thank you' header. Item list, subtotal, tax and total on the overview page are never read, so a pricing bug would pass.
Fix: Add an overview assertion: the item name and the total for the backpack.
Prepared but untested: sorting and the broken-user path
`sortBy()` and `USERS.problem` are defined and never used. `problem_user` is the account the demo app ships specifically to break images and sorting, so the most interesting cases are the ones left out.
Fix: Write the two tests (sort by price, problem user sees broken images) or delete the dead code.
No lint gate: a missing `await` would compile and pass
CI runs `tsc` only. I ran ESLint 9 with eslint-plugin-playwright (recommended) and `no-floating-promises`: 0 errors, 1 warning (`expect-expect` on login.spec.ts:5, where the assertion is hidden inside a page object). The code is clean today; nothing keeps it clean.
Fix: Add the same ESLint config and a `lint` step to CI.
CI runs on an end-of-life Node version
Node 20 reached end of life on 2026-04-30 and no longer receives security fixes.
Fix: Move to Node 22 or 24.
Every cart and checkout test logs in through the UI
Fine at 24 tests. At a few hundred it becomes the biggest cost in the suite, and a login regression fails every test at once instead of one.
Fix: Log in once in a setup project and reuse `storageState`.
A negative assertion that would pass on a typo
`toBeHidden()` on a test id also passes when the id does not exist. Today cart.spec.ts:7 proves the locator; if that test changes, this one passes no matter what.
Fix: Assert something positive as well, for example that the cart link is visible and the badge has count 0.
The README claims more than the suite shows
It calls 8 tests against a demo site 'production-style' and says the patterns 'scale to a much larger suite unchanged'. H2, H3 and L1 are the parts that would change at scale.
Fix: Describe it as a pattern demo and list what a larger suite would add.
What is already right
- Stable locators: everything goes through `getByTestId()` with `testIdAttribute: 'data-test'` (playwright.config.ts:31).
- No `waitForTimeout` anywhere; only web-first, auto-waiting assertions.
- `forbidOnly` on CI (playwright.config.ts:14) and strict TypeScript (tsconfig.json:9-11). `tsc --noEmit` passes.
- Login negatives are one data-driven loop instead of four copies (tests/login.spec.ts:12-25).
- Trace, screenshot and video only on failure or retry (playwright.config.ts:37-39).
Want this for your suite?
A paid audit also covers your last 30 days of CI runs and ends with a debrief call.
Book a 30-min fit call