-
Notifications
You must be signed in to change notification settings - Fork 819
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(jax): interpolate the tensor-sequence-parallelism warning
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3486
opened Sep 5, 2026 by
Anai-Guo
Loading…
[PyTorch] Fix NaN expert weight gradients at num_groups == 1 with SReLU
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3485
opened Sep 5, 2026 by
GarlGuo
Loading…
6 of 13 tasks
[Common] Fix Grouped MXFP8 tensor-map validation and synchronization
#3483
opened Sep 4, 2026 by
Oleg-Goncharov
Collaborator
Loading…
3 of 13 tasks
[PyTorch] Split FusedAttnFunc into single-argument forward/backward helpers
2.20
#3480
opened Sep 4, 2026 by
pggPL
Collaborator
Loading…
6 of 13 tasks
Drop the unsupported window_size argument from attention backend queries
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3479
opened Sep 4, 2026 by
Anai-Guo
Loading…
[PyT] Linear Attention API
2.20
#3477
opened Sep 4, 2026 by
KshitijLakhani
Collaborator
•
Draft
13 tasks
[PyTorch] Enable fused activation recompute for ScaledTanhSReLU
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[All] Guard THD learnable dSink on older cuDNN
attention
bug
Something isn't working
#3470
opened Sep 3, 2026 by
KshitijLakhani
Collaborator
Loading…
2 of 13 tasks
[PyTorch] Fix FP8 illegal memory access in single-process multi-GPU execution
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3469
opened Sep 3, 2026 by
SuperGoodGame
Loading…
Add opt-in rowwise-only quantized primary weights
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3468
opened Sep 3, 2026 by
xiuhu17
Contributor
Loading…
8 of 13 tasks
[Common] row-scaled nvfp4 path: add single-launch group fused amax
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3467
opened Sep 3, 2026 by
cael-ling
Contributor
Loading…
2 of 13 tasks
[JAX] Make MoEBlock aware of Dense TP axes
#3462
opened Sep 2, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[Common] Group NVFP4 Quantize Kernels
#3458
opened Sep 1, 2026 by
Oleg-Goncharov
Collaborator
Loading…
9 of 13 tasks
[JAX] Add SiTU-GLU support to JAX
#3457
opened Sep 1, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[PyTorch] Schedule delayed-scaling updates after backward
#3456
opened Sep 1, 2026 by
pggPL
Collaborator
Loading…
9 of 13 tasks
[Common] row-scaled nvfp4 path: fuse row/col amax into a single TMA-tiled kernel
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3454
opened Sep 1, 2026 by
cael-ling
Contributor
Loading…
2 of 13 tasks
[JAX] Attention support for explicit sharding
#3446
opened Aug 31, 2026 by
jberchtold-nvidia
Collaborator
•
Draft
8 of 13 tasks
[PyTorch] Fix mutable QB bounds in CUDA graphs
org-contribution
#3426
opened Aug 26, 2026 by
harryzhou2000
Contributor
Loading…
Support paged stashing for GroupedLinear activations
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3423
opened Aug 25, 2026 by
lhb8125
Contributor
Loading…
[PyTorch] Allow CP P2P transport group overrides
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3420
opened Aug 24, 2026 by
xiaoyao0115
•
Draft
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.