You already have Playwright tests. How to make them a suite you trust
Most teams that ask us about test automation are not starting from zero. Developers have been writing Playwright or Cypress tests for a year or two, one feature at a time. There are a few hundred of them now. Nobody owns them, nobody can say what they cover, a handful fail at random, and when the build goes red the first reaction is to run it again. The tests exist, but nobody believes them.
Throwing that work away and starting a new framework is the expensive answer, and usually the wrong one. Those tests encode knowledge about your product that took real time to collect. This guide explains how to organize Playwright tests you already have into a suite people trust, in the order we would do it, and what to stop automating along the way.
What do our existing tests actually cover?
Before fixing anything, find out what you have. It is the step teams skip most often, because it is slow, but every later decision depends on it. You cannot decide which tests to fix, delete or tag until you know what each one is for.
Build a coverage map. A spreadsheet is fine to begin with. One row per test file, and for each one write down:
- The feature or requirement it checks, in product language: "a customer can apply a discount code at checkout", not "checkout.spec.ts line 40".
- Its current state: passes reliably, fails sometimes, always fails, or skipped.
- How long it takes to run. Playwright reports this per test, so it costs nothing to collect.
- Who last touched it, from the git history. That person is often the only one who knows what it was meant to prove.
Then turn the map around and read it from the product side. List the journeys that would hurt most if they broke, such as sign-up, payment or permissions, and check which of them have any test at all. Teams often find that their suite is thick around the features that were easy to automate and thin around the ones that matter. That gap is the real finding of the exercise, more useful than any count of tests.
A spreadsheet goes stale the week after you build it. That is why we keep existing automated tests as test cases in BugBoard, the test management tool we built, next to the manual and new test cases, each linked to the requirement it covers. Coverage then becomes something you can look up instead of something you have to reconstruct every quarter.
Which tests should we delete, and what do we do with the flaky ones?
The coverage map will show three kinds of problem. Each needs a different response.
Duplicates
When several developers write tests over time, the same journey gets tested three or four times, each a little differently. Logging in, adding to the cart and checking the total might appear in a dozen files. Keep the clearest version of each check and remove the rest.
Dead tests
Some tests check features that no longer exist, or have been marked as skipped for so long that nobody remembers why. A skipped test is worse than no test, because it appears in the count and suggests coverage that is not there. If nobody can say what a test is for after reading it and asking its author, delete it. Git keeps the history if you need it back.
Flaky tests
A flaky test passes and fails on the same code. It is the main reason teams stop trusting a suite, because once people learn that red sometimes means nothing, they start treating all red as nothing. The answer is not to retry until green. Retries hide the problem and slow every run.
Instead, move a flaky test out of the blocking run into a quarantine group straight away, so it stops training people to ignore failures. Then track it like a bug, with an owner and a deadline. Most flakiness comes from a small set of causes: fixed waits instead of waiting for a condition, tests that depend on data another test created, shared accounts, and selectors that match more than one element. Fix the cause and bring the test back. If it sits in quarantine for weeks untouched, it was not worth keeping: delete it.
Why do our tests break every time the interface changes?
Usually because they find elements by their position in the page. A selector such as div.main > div:nth-child(3) > button describes where a button happens to sit today. Move one wrapper and the test fails, even though nothing a user cares about has changed.
Find elements the way a user would instead. Playwright's own guidance on locators recommends roles and visible names first ("the button called Pay now"), and dedicated test ids where the interface gives you nothing better. Agree on one convention, add test ids to the components that need them, and replace the brittle selectors file by file, starting with the tests that break most often.
The second big source of breakage is repeated setup. If thirty tests each log in by filling the form, a change to the login page breaks thirty tests. Move login into one shared step: log in once, save the session, and reuse it across tests. Do the same for creating test data. Each test should set up what it needs through a shared helper or an API call, and should never rely on what an earlier test left behind.
How should the suite run so people read the results?
A suite that takes forty minutes and runs once a night is a suite nobody waits for. Split it by purpose. Tag each test with its speed and priority, so different runs can pick different sets:
- Smoke: a small set covering the journeys that must never break. Fast enough to run on every pull request and block the merge if it fails.
- Regression: the wider suite, run before a release or on a schedule.
- Quarantine: the flaky tests being fixed, run separately and never blocking anyone.
Then make the results readable. Run the suite in your existing CI, not on someone's laptop. Publish the HTML report, and keep traces and screenshots for failed tests so whoever looks at a failure can see what happened without running it again. If the only way to diagnose a red build is to ask the author, the suite still belongs to one person, not the team.
New tests should join the same suite rather than start a parallel one. When we record new browser tests with Flows, our test recorder, they export to Playwright, so they land in the same framework, the same repository and the same CI run as everything else. Flows also repairs its own tests when the interface changes, which takes some of the selector maintenance off the team.
What should we stop trying to automate?
Part of organising a suite is deciding what does not belong in it. Automation pays off on checks that are repeated often, have a clear right answer and sit on a stable part of the product. It costs more than it saves elsewhere.
Leave out features that are still changing every sprint; test them by hand until they settle. Leave out checks that need human judgement, such as whether a layout reads well or whether an error message makes sense. Leave out one-off migrations and rare admin tasks that will run twice in their life. And be careful with long end-to-end journeys that cross several systems: they are slow and break for many reasons, and a few targeted integration tests often prove the same thing with far less upkeep.
A smaller suite that passes when the product works and fails when it does not is worth more than a large one people have learned to ignore. That trust is the actual deliverable.
Should we do this ourselves or bring someone in?
Your developers can do all of the above. The problem is rarely skill. It is that a test clean-up never wins against the next feature, so it gets started and stopped for months. What helps most is someone whose job is the suite.
That is the work our test automation engineers do. We work in your repository and your CI, we build the coverage map with your team, clean up and stabilise the existing tests, and then keep adding to the same suite. What we build belongs to you. Flows and BugBoard are included with the engagement at no extra licence cost, and every generated test case is reviewed by a person before it counts. Rates are $25 to $45 an hour, and a team is usually working on your product about two weeks after the first call. If you would rather we own the whole testing function, see managed testing services.
Frequently asked questions
Turn the tests you have into a suite you trust
Tell us what your suite looks like today and where it lets you down. We will tell you what we would map, fix and delete first.
Book a call See our testing servicesNeed help with software testing?
BetterQA provides independent QA services across manual testing, automation, security audits, and performance testing. ISO 27001, 9001, 14001 and 13485 certified.