Repository navigation
Leave a destroyed awaiter alone when its cancellation runs late - #56
Conversation
|
Thanks for tracking this down — the GoalStream fix works: on main, However, the early return in
Without it, a later These tests pass on main and crash with this PR (segfault 5/5 runs in a plain Debug build, heap-use-after-free under ASan): static Task<void> wait_event(Event & event) { co_await event.wait(); }
static Task<void> lock_mutex(Mutex & mutex) { co_await mutex.lock(); }
TEST_F(TaskDestroyTest, EventSetAfterWaiterDestroyed)
{
Event event(*ctx_);
{
auto task = ctx_->create_task(wait_event(event));
spin_for(50ms);
}
spin_for(50ms); // let the deferred cancellation run
event.set();
spin_for(50ms);
}
TEST_F(TaskDestroyTest, MutexUnlockAfterWaiterDestroyed)
{
Mutex mutex(*ctx_);
auto holder = ctx_->create_task(lock_mutex(mutex));
spin_for(50ms);
ASSERT_TRUE(mutex.is_locked());
{
auto task = ctx_->create_task(lock_mutex(mutex));
spin_for(50ms);
}
spin_for(50ms);
mutex.unlock();
spin_for(50ms);
EXPECT_FALSE(mutex.is_locked());
}( Since the existing tests only destroy TopicStream / GoalStream waiters, CI doesn't catch this. Suggestion: do what you already did for TEST_F(TaskDestroyTest, EventSetBeforeDeferredCancellationRuns)
{
Event event(*ctx_);
{
auto task = ctx_->create_task(wait_event(event));
spin_for(50ms);
}
event.set(); // the deferred cancellation has not run yet
spin_for(50ms);
}Could you add these tests along with the fix? |
|
Thanks for the careful review — you're right on all counts. Skipping I've followed your suggestion: each of those awaiters now withdraws its waiter in its destructor ( I added your three tests, plus Two related things I noticed but left out of this PR:
Happy to send follow-ups for either. |
|
Thanks a lot for the quick and thorough fix, and for the extra tests! |
~Task() destroyed the coroutine frame even while it was suspended in an awaiter. Everything that awaiter had handed the handle to -- a stream's waiter, an Event queue, a pending service response, a posted resumption, a TF pending request -- then pointed at freed memory, and each awaiter grew its own guard against that (is_done predicates, weak StopCb checks, the destructors from #56). A running frame is now never destroyed by its owner. ~Task() cancels it and lets go: the awaiter resumes it on the executor, CancelledException unwinds the body, and FinalAwaiter frees the frame. With a live frame behind every handed out handle, the per-awaiter destructors are no longer needed and are removed. create_task now posts the start instead of resuming on the spot, so a task starts after whatever was scheduled before it -- the unwinding of a task cancelled just before, say -- and a task released before its start never runs. This relies on the callable passed to create_task being kept alive (previous commit). Add tests for the ownership rule, the start order, and the remaining cases where a resumption was posted before the task was destroyed (Channel, TfBuffer, send_goal), which crashed on main.
register_cancelqueues the cancellation on the executor and, when it runs, callsis_done()andaction()before it checks that the awaiter is still alive.TopicStreamandGoalStreampass closures that capture the awaiter (this), so a task destroyed aftercancel()but before that queued callback runs makes it read the destroyed frame.~Task()itself requests a stop before destroying the frame, so dropping a suspended task is enough to trigger it.The same streams also keep the destroyed frame as their
waiter_, since only the deferred cancellation would clear it. A stream that outlives the coroutine (held outside it) then resumes the destroyed frame on its next message.Check that the awaiter is alive before calling
is_done()/action(), and clear the stream'swaiter_from theNextAwaiterdestructor when it is still the one registered.TimerStreamandChannelalready clear their waiter synchronously in the stop callback and check the awaiter before resuming, so they are unchanged.Add two
TopicStreamtests that destroy a waiting task (after and withoutcancel()) and then keep using the stream. Built with-fsanitize=address, the first reports a heap-use-after-free in theregister_cancelclosure without this change and passes with it. (LeakSanitizer was off for that run:create_timer's stream and timer keep each other alive, which reports leaks independently of this change.)GoalStream's deferred cancellation also cancels the goal (set_auto_cancel_on_stop). Destroying a task that waits for feedback relied on that callback running against the destroyed awaiter. With the awaiter check, it no longer runs, soNextAwaiter's destructor now cancels the goal itself when it is destroyed while still waiting.test_task_destroygainsDestroyingATaskAwaitingFeedbackCancelsTheGoal, which fails without that and passes with it. (Without it, the test server's feedback thread outlived the test and crashed the lyrical Release job at exit.)