Skip to content

LTRAC-2085: test(core) - Give each E2E run its own customer - #3258

Merged
jorgemoya merged 7 commits into
canaryfrom
LTRAC-2085/fix/e2e-run-scoped-customer
Oct 1, 2026
Merged

jorgemoya merged 7 commits into
canaryfrom
LTRAC-2085/fix/e2e-run-scoped-customer

Conversation

@jorgemoya

@jorgemoya jorgemoya commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Linear: LTRAC-2085

What/Why?

E2E Functional Tests (default) keeps going red or flaky on account tests (wishlists, addresses, orders, register). The usual error: right after customer.login(), the next account page lands on /login.

Root cause: BigCommerce keeps only one live customer access token per customer. A new login revokes the customer's previous token. I checked this against the API with a throwaway customer: after a second login, the first token returns "A customer access token was provided in the headers, but it was not valid." Every E2E run, on every PR and integrations/* branch, logged in as the same repo-level TEST_CUSTOMER on the same store. When two runs overlapped, each one's logins revoked the other's cached session. The app's GraphQL onError then redirected to /api/auth/signout, and the test landed on /login. Overlapping runs also wiped each other's data (deleteAllCustomerOrders, deleteAllWishlists, …).

The timing matches. Every red or flaky default job overlapped another run's default job. #3255 ran 15:06–15:16, inside the b2b-makeswift run (14:39–15:32). Both clean #3254 runs had nothing running beside them.

Fix: each run creates its own customer the first time a test needs one, and global teardown deletes it. Tests in the run share that one customer, so session reuse still works. Global setup also deletes a customer left over from an interrupted run. TEST_CUSTOMER_* is now only used with TESTS_READ_ONLY, where tests can't create customers. No test depended on the shared customer's existing data; they all seed their own through the API.

Worth reviewing:

  • If a CI job is cancelled mid-run, teardown never runs and that run's @example.com customer stays in the store. The next run on the same machine cleans it up, but a fresh CI runner can't. A periodic cleanup could follow if it matters.
  • Within a run, tests that log in without the session store still rotate the token (UI login in product.spec.ts:234, JWT login in login.spec.ts). The existing reuse check catches that and logs in again. This PR doesn't change it.

CI: run E2E in the Playwright container

On the first push, the default job sat on "Install Playwright browsers" for 40+ minutes. The job has no timeout, so it would have kept going. The log showed the cause: --with-deps webkit pulls 125 MB of apt packages (GStreamer, ffmpeg codecs, …), and azure.archive.ubuntu.com was serving them at ~50 kB/s. That step took ~1 min on Sep 29, then 29 min on #3255, and timed out here. All of that is for a single WebKit test (wishlists.mobile.spec.ts).

The E2E job now runs in mcr.microsoft.com/playwright:v<version>-noble, which ships every browser and its system libraries, so the job installs nothing. A small playwright-version job reads the locked @playwright/test version from pnpm-lock.yaml, so a dependency bump can't leave the image behind. The job runs with --user 1001, the runner's UID, as Playwright's CI docs show, so checkout and the pnpm cache keep working.

(Two commits along the way first added a browser cache and a 15-minute install timeout. The timeout worked, but the cache didn't address the apt download. The container change replaces both.)

Blog specs: match only visible text

In the container, the first full run passed everything except blog.spec.ts (Blog post page displays … failed, Blog can be filtered by tags flaked). The blog post page streams in, and React keeps a revealed boundary in a hidden copy for a moment before swapping it in. getByText matches hidden elements too, so breadcrumbs.getByText('Home') resolved to two elements. A strict-mode violation fails immediately instead of retrying. The error context's ARIA tree shows only one breadcrumb, so the second match is the hidden copy. The test failed 3/3 locally too, on the author name, for the same reason. The blog index test already works around this with .first(). These two tests now use .filter({ visible: true }), and pass 8/8 locally.

Shipping spec: handle both form states after an address change

Update shipping estimates flaked on the green container run and failed 2/10 locally, for two reasons:

  • Test bug (fixed): after changing to a new random address, the test checked once for the shipping options with isVisible(), which doesn't wait, and always clicked "Update shipping". The form only shows that label while an option is still selected. When BigCommerce drops the option for the new address, the button reads "Add shipping" and the test timed out. It now waits for the options or the summary, then clicks whichever button is shown.
  • Upstream (not fixed): some runs hit "There was a server error!" because the store's API failed mid-request: TypeError: terminated (SocketError: other side closed in CI, NGHTTP2_INTERNAL_ERROR locally) or BigCommerce API returned 502. Retries cover it. Whether the cart should retry dropped requests is a separate question.

After the fix: 14/15 passed locally. The one failure was an upstream 502 before the changed code.

Facets spec: compare counts instead of hard-coding them

The facets tests asserted the heading "Shop All 13", and "Shop All 5" after filtering. Two overlapping runs (this PR and LTRAC-2145, which has none of these changes) both saw "Shop All 12" for about 2 minutes, and all four tests failed every retry. The default product, [Sample] Smith Journal 13, was modified at 19:26:19, 2 seconds after a checkout.spec guest checkout. My guess is a re-index briefly dropped it from the listing, but I couldn't reproduce that with an API order or a local guest checkout. The tests now read the unfiltered count from the heading, check that the filter lowers it, and check that resetting restores it.

Testing

Locally against the E2E test store (next build + next start, chromium, --retries=0):

  • account-settings.spec.ts + wishlists.spec.ts:29 (the hard failure on Version Packages (canary) #3255): 5/5 pass. Teardown deleted the run's customer, and setup deleted a customer left by an interrupted run.
  • Full account/ + auth/ + product.spec.ts: 56 passed, 14 failed. 10 of the failures (auth/login, register) fail the same way on canary without this change; they're waitForURL timeouts from my local env. product.spec.ts:234 fails the same way, but I didn't run it on canary. wishlist-details.spec.ts:99 failed once in two tries on a "removed" toast; it's unrelated to sessions and also flaked in CI before this change.

The real check is this PR's CI, ideally while another E2E run overlaps it.

Migration

None. Tests and CI only. TEST_CUSTOMER_* can stay in e2e.yml; it's ignored unless TESTS_READ_ONLY is set.

🤖 Generated with Claude Code

@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
catalyst Ready Ready Preview Oct 1, 2026 8:03pm UTC

Request Review

@jorgemoya
jorgemoya marked this pull request as ready for review October 1, 2026 17:14
@jorgemoya
jorgemoya requested a review from a team as a code owner October 1, 2026 17:14
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Unlighthouse Performance Comparison — Vercel

Comparing PR preview deployment Unlighthouse scores vs production Unlighthouse scores.

Summary Score

Aggregate score across all categories as reported by Unlighthouse.

Prod Desktop Prod Mobile Preview Desktop Preview Mobile
Score 98 98 99 99

Category Scores

Category Prod Desktop Prod Mobile Preview Desktop Preview Mobile
Performance 100 100 100 100
Accessibility 95 92 95 92
Best Practices 100 100 100 100
SEO 100 100 100 100

Core Web Vitals

Metric Prod Desktop Prod Mobile Preview Desktop Preview Mobile
LCP 0.4 s 0.7 s 0.3 s 0.4 s
CLS 0.05 0 0 0
FCP 0.3 s 0.3 s 0.3 s 0.4 s
TBT 0 ms 20 ms 0 ms 0 ms
Max Potential FID 50 ms 70 ms 40 ms 40 ms
Time to Interactive 0.3 s 0.6 s 0.3 s 0.4 s

Full Unlighthouse report →

@changeset-bot

changeset-bot Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 9212186

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

jorgemoya and others added 7 commits October 1, 2026 15:03
BigCommerce keeps one live customer access token per customer, so a new
login revokes every other session for that customer. Every E2E run logged
in as the shared TEST_CUSTOMER, so runs that overlapped on the store
revoked each other's sessions, and the next account page bounced to
/login. They also wiped each other's orders, addresses and wishlists.

Create a customer the first time a run needs one and delete it in the
global teardown. TEST_CUSTOMER_* is now only used in read-only mode,
where tests can't create customers.

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d-only mode

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Since Sep 30 the "Install Playwright browsers" step has taken anywhere
from 6 minutes to 40+ (it normally takes about a minute). The job has no
timeout, so a hung install can hold the E2E check for hours. Cache
~/.cache/ms-playwright per Playwright version and browser set, so a
cache hit only installs the system deps. Give either install step 15
minutes before it fails.

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The slow install step was WebKit's ~125 MB of apt dependencies, which
download at ~50 kB/s from the runner's Ubuntu mirror. The browser cache
only skipped the browser download, so the default job still hit the
15-minute timeout. The official Playwright image ships browsers and
their system libraries, so the job no longer installs anything.

A small job reads the locked @playwright/test version, so the image tag
stays in step with dependency bumps.

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The blog post page streams its content. React keeps a revealed boundary
in a hidden copy for a moment before swapping it in, and text locators
match hidden elements too. If an assertion lands in that window,
breadcrumbs.getByText('Home') resolves to two elements, and the
strict-mode violation fails the test right away instead of retrying. It
failed 3/3 locally and in the Playwright container. Filter those
locators to visible elements, as the blog index test already does with
.first().

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dress change

After the address changes, "Update shipping estimates" checked once
for the shipping options with isVisible(), which doesn't wait, so it
could skip re-selecting an option. It also always clicked "Update
shipping". The form only shows that label while an option is still
selected. BigCommerce drops the option when it doesn't apply to the new
random address, and the button then reads "Add shipping", so the test
timed out. Wait for the options or the summary, then click whichever
submit button the form shows.

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ng them

The facets specs asserted the heading "Shop All 13" (and "Shop All 5"
after filtering). Two overlapping CI runs saw "Shop All 12" for about
two minutes, right after a checkout modified the default product, and
all four tests failed every retry. Read the unfiltered count from the
heading, then check that the filter narrows it and that resetting
restores it.

Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jorgemoya
jorgemoya added this pull request to the merge queue Oct 1, 2026
Merged via the queue into canary with commit f5991b7 Oct 1, 2026
20 checks passed
@jorgemoya
jorgemoya deleted the LTRAC-2085/fix/e2e-run-scoped-customer branch October 1, 2026 22:18

This branch was successfully deployed

1 active deployment
Preview — 92121869 Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants