[https://nvbugs/6487038][fix] Stop single-rank disagg errors from crashing all gen ranks - #16834
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughThe change tolerates transient active-request overshoot, handles completed response threads and event-loop errors, adds regression tests, and updates disaggregated NIXL transfer settings. ChangesExecutor safeguards and transfer configuration
Estimated code review effort: 3 (Moderate) | ~20 minutes Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
adad095 to
d290731
Compare
|
/bot run --post-merge --stage-list "GB200-16_GPUs-4_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE2-GPU8-GEN1-NODE2-GPU8-Post-Merge, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge" |
|
PR_Github #61811 [ run ] triggered by Bot. Commit: |
|
PR_Github #61811 Bot args parsing error: CI requested by |
05f16c5 to
9727efc
Compare
|
/bot run --post-merge --stage-list "GB200-16_GPUs-4_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE2-GPU8-GEN1-NODE2-GPU8-Post-Merge, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge" |
|
PR_Github #61826 [ run ] triggered by Bot. Commit: |
|
PR_Github #61826 Bot args parsing error: CI requested by |
|
Removed the "ci: post-merge approved" label because @niukuo could not be verified as an active member of NVIDIA/trt-llm-ci-approvers. Ask a member of that team to apply it. |
9727efc to
e14fc1b
Compare
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-1, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-3" |
|
PR_Github #65189 [ run ] triggered by Bot. Commit: |
|
PR_Github #65189 [ run ] completed with state
|
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-1, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-3" |
|
PR_Github #65514 [ run ] triggered by Bot. Commit: |
|
PR_Github #65514 [ run ] completed with state
|
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-1, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-3" |
|
PR_Github #65538 [ run ] triggered by Bot. Commit: |
|
PR_Github #65538 [ run ] completed with state
|
|
/bot run --add-multi-gpu-test --disable-fail-fast --extra-stage "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-1, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-3" |
|
PR_Github #65573 [ skip ] triggered by Bot. Commit: |
|
PR_Github #65574 [ run ] triggered by Bot. Commit: |
|
PR_Github #65573 [ skip ] completed with state |
|
PR_Github #65574 [ run ] completed with state
|
|
/bot skip --comment "Infra-only failures on GB300 capacity, no regression from this PR." |
|
/bot run --disable-fail-fast --extra-stage "GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-1, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-2, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-3, GB300-20_GPUs-5_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE4-GPU16-Post-Merge-4, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-1, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-2, GB300-4_GPUs-PyTorch-PerfSanity-Post-Merge-3" |
|
PR_Github #65812 [ skip ] triggered by Bot. Commit: |
|
PR_Github #65814 [ run ] triggered by Bot. Commit: |
|
PR_Github #65812 [ skip ] completed with state |
|
PR_Github #65814 [ run ] completed with state
|
|
/bot skip --comment "Infra-only failures on GB300 capacity, no regression from this PR." |
|
PR_Github #65853 [ skip ] triggered by Bot. Commit: |
|
PR_Github #65853 [ skip ] completed with state |
Description
Three separate changes, one per commit so each can be bisected and reverted on its own. They
are in one PR so the disagg CI cost is paid once rather than three times.
1.
[fix]Do not crash every ADP rank on a disagg transfer error —py_executor.py_pad_attention_dp_dummy_request()assertedexpected_num_active_requests >= len(active_requests). Requests whose KV transfer errored can linger inactive_requestsfor atick before cleanup drains them, so a single-rank disagg error tripped the assert and took down
the generation loop on every attention-DP rank. The overshoot is transient and
self-correcting, and the padding keys on the schedulable count rather than the raw length, so
it now warns once and continues.
2.
[fix]Surface the engine event-loop error fromstart_thread—worker.pyA
ManagedThreadthat already ran cannot be restarted, sostart_threadfell through tothread.start()and raisedRuntimeError: threads can only be started onceout ofsubmit(),hiding the engine failure that actually killed the thread. It now reports the stashed
_event_loop_error, wrapped asRequestError(str(err)) from err—submit()callsstart()every time, and re-raising the same object would append a frame to its
__traceback__on eachcall. This also matches how
base_worker.pyreports its other submit-path failures. Thepost-shutdown path (
stop()setstop_event, no error stashed) returns quietly.3.
[chore]Tune gb300 disagg perf-sanity transfer configs — three YAMLsA behavior change on three CI perf cases, so calling it out explicitly:
CPPtoPYTHON;kv_cache_bounce_size_mb: 2048andTRTLLM_KV_TRANSFER_NUM_THREADS=4;kv_transfer_timeout_ms: 600000. The 8k1k cases previously relied on the60 s default while switching transceiver, which is the exact bound NVBug 6487038 reports being
exceeded; transfer time varies with the environment, so all three are bounded the same way.
The 128k8k case deliberately gets neither the bounce buffer nor the extra transfer threads, and
says so inline, because measurement says they hurt this shape.
Test Coverage
Unit tests, all CPU-only:
tests/unittest/executor/test_event_loop_error_broadcast.py::TestStartThreadAfterExit(new):engine error →
RequestErrorwith__cause__intact and no restart; repeated calls → distinctexception objects, guarding the
__traceback__accumulation;_event_loop_error is None→quiet return; fresh thread → still starts.
tests/unittest/_torch/executor/test_py_executor.py: the overshoot test now pins the outcome(
add_dummy_calls == [],active_requestsunchanged) instead of only asserting no raise, andtest_pad_dummy_added_when_overshoot_has_no_schedulable_requestscovers the branch thatmatters — overshoot with zero schedulable requests must still add the pad dummy.
A/B for the 128k8k transfer config, GB300, 3 nodes / 12 GPUs,
disagg-e2e, one run per arm:One sample per arm on a shared cluster, so the exact percentages are soft — but every metric
moves the same way, and the slower arm finishes in 9513 s against the case's 180 min budget,
which is why the tuning is not applied there.
Dev Engineer Review
RequestErrorand avoid invalid restarts.600000ms timeouts to the 8k1k configurations.QA Engineer Review
test_pad_dummy_tolerates_active_request_overshoot().test_pad_dummy_added_when_overshoot_has_no_schedulable_requests().test_surfaces_engine_error_as_request_error().test_repeated_calls_do_not_accumulate_traceback().test_post_shutdown_exit_returns_quietly().test_fresh_thread_is_started().test_does_not_start_when_enqueueing_is_disabled().test-db/orqa/entries were provided for these tests.