Repository navigation
LTRAC-2085: test(core) - Give each E2E run its own customer - #3258
Merged
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
jorgemoya
marked this pull request as ready for review
October 1, 2026 17:14
Contributor
Unlighthouse Performance Comparison — VercelComparing PR preview deployment Unlighthouse scores vs production Unlighthouse scores. Summary ScoreAggregate score across all categories as reported by Unlighthouse.
Category Scores
Core Web Vitals
|
|
BigCommerce keeps one live customer access token per customer, so a new login revokes every other session for that customer. Every E2E run logged in as the shared TEST_CUSTOMER, so runs that overlapped on the store revoked each other's sessions, and the next account page bounced to /login. They also wiped each other's orders, addresses and wishlists. Create a customer the first time a run needs one and delete it in the global teardown. TEST_CUSTOMER_* is now only used in read-only mode, where tests can't create customers. Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…d-only mode Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Since Sep 30 the "Install Playwright browsers" step has taken anywhere from 6 minutes to 40+ (it normally takes about a minute). The job has no timeout, so a hung install can hold the E2E check for hours. Cache ~/.cache/ms-playwright per Playwright version and browser set, so a cache hit only installs the system deps. Give either install step 15 minutes before it fails. Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The slow install step was WebKit's ~125 MB of apt dependencies, which download at ~50 kB/s from the runner's Ubuntu mirror. The browser cache only skipped the browser download, so the default job still hit the 15-minute timeout. The official Playwright image ships browsers and their system libraries, so the job no longer installs anything. A small job reads the locked @playwright/test version, so the image tag stays in step with dependency bumps. Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The blog post page streams its content. React keeps a revealed boundary
in a hidden copy for a moment before swapping it in, and text locators
match hidden elements too. If an assertion lands in that window,
breadcrumbs.getByText('Home') resolves to two elements, and the
strict-mode violation fails the test right away instead of retrying. It
failed 3/3 locally and in the Playwright container. Filter those
locators to visible elements, as the blog index test already does with
.first().
Refs LTRAC-2085
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dress change After the address changes, "Update shipping estimates" checked once for the shipping options with isVisible(), which doesn't wait, so it could skip re-selecting an option. It also always clicked "Update shipping". The form only shows that label while an option is still selected. BigCommerce drops the option when it doesn't apply to the new random address, and the button then reads "Add shipping", so the test timed out. Wait for the options or the summary, then click whichever submit button the form shows. Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ng them The facets specs asserted the heading "Shop All 13" (and "Shop All 5" after filtering). Two overlapping CI runs saw "Shop All 12" for about two minutes, right after a checkout modified the default product, and all four tests failed every retry. Read the unfiltered count from the heading, then check that the filter narrows it and that resetting restores it. Refs LTRAC-2085 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jorgemoya
force-pushed
the
LTRAC-2085/fix/e2e-run-scoped-customer
branch
from
October 1, 2026 20:03
4db69ea to
9212186
Compare
parthshahp
approved these changes
Oct 1, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Linear: LTRAC-2085
What/Why?
E2E Functional Tests (default)keeps going red or flaky on account tests (wishlists, addresses, orders, register). The usual error: right aftercustomer.login(), the next account page lands on/login.Root cause: BigCommerce keeps only one live customer access token per customer. A new
loginrevokes the customer's previous token. I checked this against the API with a throwaway customer: after a second login, the first token returns "A customer access token was provided in the headers, but it was not valid." Every E2E run, on every PR andintegrations/*branch, logged in as the same repo-levelTEST_CUSTOMERon the same store. When two runs overlapped, each one's logins revoked the other's cached session. The app's GraphQLonErrorthen redirected to/api/auth/signout, and the test landed on/login. Overlapping runs also wiped each other's data (deleteAllCustomerOrders,deleteAllWishlists, …).The timing matches. Every red or flaky
defaultjob overlapped another run'sdefaultjob. #3255 ran 15:06–15:16, inside the b2b-makeswift run (14:39–15:32). Both clean #3254 runs had nothing running beside them.Fix: each run creates its own customer the first time a test needs one, and global teardown deletes it. Tests in the run share that one customer, so session reuse still works. Global setup also deletes a customer left over from an interrupted run.
TEST_CUSTOMER_*is now only used withTESTS_READ_ONLY, where tests can't create customers. No test depended on the shared customer's existing data; they all seed their own through the API.Worth reviewing:
@example.comcustomer stays in the store. The next run on the same machine cleans it up, but a fresh CI runner can't. A periodic cleanup could follow if it matters.product.spec.ts:234, JWT login inlogin.spec.ts). The existing reuse check catches that and logs in again. This PR doesn't change it.CI: run E2E in the Playwright container
On the first push, the
defaultjob sat on "Install Playwright browsers" for 40+ minutes. The job has no timeout, so it would have kept going. The log showed the cause:--with-deps webkitpulls 125 MB of apt packages (GStreamer, ffmpeg codecs, …), andazure.archive.ubuntu.comwas serving them at ~50 kB/s. That step took ~1 min on Sep 29, then 29 min on #3255, and timed out here. All of that is for a single WebKit test (wishlists.mobile.spec.ts).The E2E job now runs in
mcr.microsoft.com/playwright:v<version>-noble, which ships every browser and its system libraries, so the job installs nothing. A smallplaywright-versionjob reads the locked@playwright/testversion frompnpm-lock.yaml, so a dependency bump can't leave the image behind. The job runs with--user 1001, the runner's UID, as Playwright's CI docs show, so checkout and the pnpm cache keep working.(Two commits along the way first added a browser cache and a 15-minute install timeout. The timeout worked, but the cache didn't address the apt download. The container change replaces both.)
Blog specs: match only visible text
In the container, the first full run passed everything except
blog.spec.ts(Blog post page displays …failed,Blog can be filtered by tagsflaked). The blog post page streams in, and React keeps a revealed boundary in a hidden copy for a moment before swapping it in.getByTextmatches hidden elements too, sobreadcrumbs.getByText('Home')resolved to two elements. A strict-mode violation fails immediately instead of retrying. The error context's ARIA tree shows only one breadcrumb, so the second match is the hidden copy. The test failed 3/3 locally too, on the author name, for the same reason. The blog index test already works around this with.first(). These two tests now use.filter({ visible: true }), and pass 8/8 locally.Shipping spec: handle both form states after an address change
Update shipping estimatesflaked on the green container run and failed 2/10 locally, for two reasons:isVisible(), which doesn't wait, and always clicked "Update shipping". The form only shows that label while an option is still selected. When BigCommerce drops the option for the new address, the button reads "Add shipping" and the test timed out. It now waits for the options or the summary, then clicks whichever button is shown.TypeError: terminated(SocketError: other side closedin CI,NGHTTP2_INTERNAL_ERRORlocally) orBigCommerce API returned 502. Retries cover it. Whether the cart should retry dropped requests is a separate question.After the fix: 14/15 passed locally. The one failure was an upstream 502 before the changed code.
Facets spec: compare counts instead of hard-coding them
The facets tests asserted the heading "Shop All 13", and "Shop All 5" after filtering. Two overlapping runs (this PR and LTRAC-2145, which has none of these changes) both saw "Shop All 12" for about 2 minutes, and all four tests failed every retry. The default product,
[Sample] Smith Journal 13, was modified at 19:26:19, 2 seconds after acheckout.specguest checkout. My guess is a re-index briefly dropped it from the listing, but I couldn't reproduce that with an API order or a local guest checkout. The tests now read the unfiltered count from the heading, check that the filter lowers it, and check that resetting restores it.Testing
Locally against the E2E test store (
next build+next start, chromium,--retries=0):account-settings.spec.ts+wishlists.spec.ts:29(the hard failure on Version Packages (canary) #3255): 5/5 pass. Teardown deleted the run's customer, and setup deleted a customer left by an interrupted run.account/+auth/+product.spec.ts: 56 passed, 14 failed. 10 of the failures (auth/login, register) fail the same way oncanarywithout this change; they'rewaitForURLtimeouts from my local env.product.spec.ts:234fails the same way, but I didn't run it oncanary.wishlist-details.spec.ts:99failed once in two tries on a "removed" toast; it's unrelated to sessions and also flaked in CI before this change.The real check is this PR's CI, ideally while another E2E run overlaps it.
Migration
None. Tests and CI only.
TEST_CUSTOMER_*can stay ine2e.yml; it's ignored unlessTESTS_READ_ONLYis set.🤖 Generated with Claude Code