Pre-launch store QA brief
Everything else you outsource before launch is graded by taste. This one is graded by whether a stranger on a phone can complete an order and be charged the right amount. That is a pass or a fail, which makes it the only item on a launch list you can check properly — and the first one dropped when the launch date gets close.
When to order it, and when to do it yourself
Outsource when the store is otherwise finished and frozen: products loaded, theme done, apps installed, policies written, and nothing queued to change while someone is testing. Do it yourself when products are still going in. QA against a moving target produces a report that is obsolete before it is read — half the defects were closed by an unrelated edit, and the tester cannot tell you which half.
Frozen is an hour and a promise, not a mood. Name the time the freeze starts and hold it: no theme edits, no new products, no app installs, no shipping-rule changes mid-pass. A defect that cannot be reproduced the next day gets dismissed as a mistake, and the one real bug in the batch leaves with it.
There is a case for running the early passes yourself that has nothing to do with saving an order. Walking your own checkout on your own phone is how you find out which parts of the store you assembled without understanding — the shipping profile you copied, the app you installed at midnight, the market you switched on and forgot. Hand that stage to a tester and you buy a list of things you would have found in twenty minutes, while the defects that genuinely need a stranger stay hidden. QA sits last in the 7-task launch outsourcing checklist for the same reason.
Why this is the item that gets cut
A QA pass produces no asset. Every other order on a launch list ends in a file you can look at: a logo, a photograph, a page, an edit. This one ends in a document listing things that are wrong, which is the least satisfying thing anyone has ever bought, and the easiest to postpone until after the first sale — a moment that, for a store with a broken checkout, never arrives. Insurance is always the simplest line to cut, because the thing it prevents has not happened yet.
The argument for buying it is not that testers are clever. It is that you are the worst available tester of your own store. Your address autofills, your session is cached, the theme you edited an hour ago is still sitting in your browser, and you have walked one route through the store so many times that you no longer see it. You test the path you built. A buyer arrives on a handset you do not own, from a market you have barely thought about, and takes the path you did not.
And a buyer does not file defects. A shipping step that returns no rates for their country produces no support email; it produces a closed tab. The failure is silent, and it is silent precisely where the money is.
What to put on the coverage list
"Test my store and tell me what is broken" returns whatever the tester happens to notice, which is nearly always the visual layer, because that is the layer that is visible. Coverage has to be enumerated. Write the list, ask for a stated result on every line, and insist on a result even where nothing was wrong.
- Checkout completed end to end on a real mobile device. Not a desktop window narrowed to phone width — a physical handset, model named in the report, with the order number as evidence. A resized browser does not reproduce the keyboard covering the card field, the autofill landing in the wrong line, or the payment sheet that never opens.
- Tax and shipping rules for every market you sell to. Each destination you have switched on, at more than one basket size if your rates are banded. The failure to hunt for is not a wrong figure; it is no rates returned at all, which reads to the buyer as a store that will not ship to them.
- Analytics and conversion tracking firing on the right events. Installed is not a test result. Ask which events fired, in what order, with what value and currency attached, and whether a completed purchase records once rather than twice. A tag that fires on every page view will hand you a conversion rate you then make decisions on for months.
- Broken links, bad redirects, and pages reachable from nothing. Navigation, footer, policy pages, links inside product descriptions, and anything you moved while building. Orphan pages matter because they tend to be the ones still carrying a discontinued variant or a superseded claim.
- Layout and tap targets at a small viewport. Name the failures you want checked: a sticky bar or consent notice sitting over the buy button, a field the keyboard hides, a swipe gallery that traps the page scroll, text that needs a pinch to read.
- Loading behaviour on the templates that matter. Home, the collection you intend to advertise, the product page you intend to advertise, cart, and checkout. Name the tool and the network condition in the brief, and take a reading before any fixes so a later comparison means something. Thresholds inside these tools change, so confirm the current ones in the tool you both agreed on rather than trusting a figure quoted from memory.
- The transactional emails a real order triggers. Order confirmation and shipping notification at minimum, plus whatever a refund or cancellation sends. Check the sender name, the reply-to address, whether the branding survived, and where each message actually landed on more than one mail provider. These are not your welcome and abandoned-cart sequences; those are a separate order with a separate brief, covered in the email flow brief.
Where to look: website development listings.
Ask for a defect list, not a screen recording
A narrated screen recording feels generous. It is homework. You watch forty minutes to extract eleven items, write them down yourself, then send that written list to whoever is fixing them — so the work of turning observation into tasks happened twice, and you did it the second time.
A list can be sorted, assigned, quoted against, and closed. A recording cannot be handed to a developer who was not on the call, and a defect nobody else can reproduce gets closed as "works for me". Require these fields on every row:
- Severity, taken from the agreed scheme rather than from feeling.
- Steps to reproduce, numbered, starting from a URL, written so a stranger can follow them without asking a question.
- Expected versus actual, one line each. A good share of what arrives as a defect is really a disagreement about what the store is supposed to do, and this field surfaces that before anyone spends an hour on it.
- Device, operating system, and browser, named. A defect that exists on one combination is a different job from one that exists everywhere, and without this you cannot tell which you have.
- One screenshot, annotated when the problem is a position rather than a value.
Video has one legitimate place: attached to a single row where the problem is movement — an animation that jumps, a gallery that swallows a swipe, a checkout that flashes and reloads. A ten-second clip there is evidence. A tour of the store is not.
Severity that means something
High, medium and low describe the tester's mood. You cannot sort by them and you cannot argue with them, because someone who cares about typography and someone who cares about failed payments will hand you two incompatible lists under the same three labels. Define the tiers by consequence instead, put the definitions in the brief, and require the tier on every row.
| Tier | Definition, by consequence | The kind of defect that belongs here | Blocks launch |
|---|---|---|---|
| Blocks an order | A buyer on a path you are sending traffic to cannot complete a purchase. | The shipping step returns no rates for a country you sell to, so checkout cannot proceed. | Yes |
| Takes money incorrectly | The order completes, but the amount, tax, currency, or discount applied is not what either side intended. | A discount code stacks with an automatic offer and the total lands below what fulfilling the order costs you. | Yes |
| Misleads the buyer | The store states something the order will not honour. | A delivery estimate on the product page that the shipping rules contradict, or a returns window the policy page denies. | Yes — this is a refund and dispute source, not a wording preference |
| Damages trust | The transaction is correct, but the store reads as unfinished or unsafe. | Placeholder text left in a policy page, a broken image in the footer, an order confirmation arriving from an unbranded address. | Anything on the path to checkout, yes; the rest can ship |
| Cosmetic | A visual imperfection with no effect on comprehension or completion. | A collection card sitting a couple of pixels out of line at one viewport width. | No |
The point of consequence tiers is that you can disagree with one. If a row is filed as blocking an order, you can walk the steps and watch whether an order completes. "Medium" gives you nothing to check.
Test orders, refunds, and clean books
Someone has to attempt a purchase, which means money moves or is deliberately made not to. Decide which route you are taking before you order, and write it into the brief.
If your platform or payment provider offers a test mode, confirm how it behaves today and whether it exercises the same path a real card would, then decide whether that is enough coverage — the parts most likely to be broken are often the ones a simulator skips. The alternatives are a real order placed by the tester on a payment method you control, or one you place yourself while the tester watches and records the result. Any of the three is defensible. Leaving it to be worked out mid-pass is not.
Most of this pass needs no admin access at all. A tester needs the storefront, the storefront password if the store is still gated, and permission to attempt a purchase. If a defect turns out to need admin visibility to diagnose, treat that as a fresh decision rather than a default — the safer way to hand over a store, and the permissions to use when you must, are in the landing page build brief.
Who fixes what, and the one re-check
QA and repair are two orders. A tester finding a defect does not mean the tester is fixing it, and a listing that mentions testing may not include a single edit. Ask in writing before you order, because assuming it ends the job in a dispute over work nobody agreed to.
Keeping them separate also keeps the list honest: a tester who is also the repairer has a quiet incentive to file what is easy to fix and to be vague about what is not. Combining them is reasonable when the same person built the store, since they already know the theme — but then the defect list is being written by the person the defect list is about, and it should be read that way.
Agree a single re-check pass at the outset, and define it: reproduce every defect in the top two tiers from the original steps, mark each one closed or still open, and confirm nothing new appeared on the paths that were touched. Left until the fixes land, a re-check becomes a negotiation you conduct from the weak position of a launch date, against someone who knows you have one.
If the repair goes to someone else, the defect list is the brief. That is the second reason to insist on steps to reproduce: they are what a repairer quotes against.
Where to look: website development listings for the pass and the repair, ordered separately. For working out which of the two a listing is actually selling, see reading a listing and vetting a seller.
Ordering the pass
Brief
STORE Store URL: Storefront password, if the store is still gated: Frozen from (date and time) — no theme, product, app, or rule changes: SCOPE Markets to test (countries), and the currency each should display: Payment methods to attempt, in priority order: Templates that matter: home, top collection, top product, cart, checkout Devices required (at least one real phone, model named, plus desktop): Browsers required on each: TEST ORDERS Test orders permitted: yes / no If yes, the payment method to use and any code created for this pass: Who refunds test orders and cancels any test fulfilment: How test orders must be labelled so I can find them afterwards: COVERAGE — report a result on every line, including the clean ones Checkout completed end to end on the named phone, with order number: Tax and shipping rates returned for every market listed above: Tracking: which events fired, in what order, with what value: Broken links, redirects, and pages reachable from no menu: Layout and tap targets at a small viewport: Loading behaviour on the templates above, with the tool named: Transactional emails a real order triggers, and where each landed: DELIVERABLE Format: one defect per row, sorted by severity, worst first Per defect: severity, steps to reproduce, expected versus actual, device and browser, one screenshot Severity scheme: the five consequence tiers agreed, not high/medium/low Re-check pass after fixes: included / quoted separately Deadline:
Check on delivery
- Every defect reproduces from the steps as written, on the device named, by someone who was not there. A defect you cannot reproduce is not a report; it is a rumour, and it will be closed as one.
- Severities match the agreed definitions. Spot-check two rows by reading the consequence rather than the label.
- At least one completed checkout is evidenced on a real mobile device: order number recorded, handset model named, not a resized desktop window.
- Tracking is verified by events observed with their values, not by a snippet being present in the theme.
- Every market you listed appears in the report, including the ones where nothing was wrong. Silence is not a pass.
- Test orders are accounted for one by one: refunded or cancelled by whoever agreed to, and excluded or tagged anywhere they would otherwise distort your numbers.
- The report is a list you can work through in order, top row first, rather than a narrative you have to convert into tasks yourself.