Found while tracing a flaky test, PgLogEventsLivePostgresTests.TheSelfHostedCollector_ReadsTheTargetsOwnLog_EndToEnd (CI run 36500483761, 2026-09-29 00:02Z).
What happens
PgServerLogTail.TailCteSql reads the last 4 MB of the newest stderr log file only (ORDER BY modification DESC LIMIT 1). After PostgreSQL rotates its log, the next read looks only at the new file. Anything written to the old file after the previous read is never read. That can be up to one collection interval of lock waits, errors, deadlocks and plans.
When
- At every rotation. By default (
log_rotation_age = 1d) that is daily at midnight in log_timezone.
- Also at every
log_rotation_size rotation, which can be often on a busy server.
The type header already documents the gap between two windows when the log grows faster than the window. It doesn't cover this one. The managed route (RdsLogSource) keeps a resume marker and isn't affected.
How the test showed it
- The workflow's workload step wrote a lock wait at 23:59:33Z.
- PostgreSQL rotated its log at 00:00:00Z.
- The test read at 00:02:05Z and saw only the new file, so its LockWait assertion failed.
On a local rig, running pg_rotate_logfile() between the workload and the read reproduces the failure every time.
The test is being fixed on its own: it will drive its own events, so it no longer depends on the workflow step. This issue is the product gap, which that fix does not change.
Possible directions (need a design call)
- When the newest file has changed since the last read, also read the previous file's tail from where the last read ended.
- Or keep a per-target resume marker (file name and offset), as the RDS route does.
Found while tracing a flaky test,
PgLogEventsLivePostgresTests.TheSelfHostedCollector_ReadsTheTargetsOwnLog_EndToEnd(CI run 36500483761, 2026-09-29 00:02Z).What happens
PgServerLogTail.TailCteSqlreads the last 4 MB of the newest stderr log file only (ORDER BY modification DESC LIMIT 1). After PostgreSQL rotates its log, the next read looks only at the new file. Anything written to the old file after the previous read is never read. That can be up to one collection interval of lock waits, errors, deadlocks and plans.When
log_rotation_age = 1d) that is daily at midnight inlog_timezone.log_rotation_sizerotation, which can be often on a busy server.The type header already documents the gap between two windows when the log grows faster than the window. It doesn't cover this one. The managed route (
RdsLogSource) keeps a resume marker and isn't affected.How the test showed it
On a local rig, running
pg_rotate_logfile()between the workload and the read reproduces the failure every time.The test is being fixed on its own: it will drive its own events, so it no longer depends on the workflow step. This issue is the product gap, which that fix does not change.
Possible directions (need a design call)