Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
NVIDIA
/
cudnn-frontend
Public
Notifications
You must be signed in to change notification settings
Fork
297
Star
946
Code
Issues
107
Pull requests
94
Actions
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Security and quality
Insights
All pull requests
New pull request
Search pull requests
is
:
pr
state
:
open
is:pr state:open
Clear filter
Search
Pull requests
Open
94
(94)
Closed
816
(816)
Author
Label
Projects
Milestones
Reviews
Assignee
Sort by
Newest
descending
More items
Comfortable display density
Compact display density
[FROST] Preserve Int64 THD output strides through descriptor setup
area:frost
area:global_attention
cat-bugfix
orig-nv-eng
#1224
·
YangXu1990uiuc
opened
Sep 25, 2026
Collaborator
·
·
·
Frontend 1.31.0
4
frost(sdpa_fwd_sm80): serve dense GQA/MQA natively, drop the per-execute K/V head expansion
#1223
·
harshithkantamneni
opened
Sep 25, 2026
·
·
2
Reduce Python graph execution overhead with ordered tensor bindings
area:frost
area:global_attention
cat-enhancements
orig-nv-eng
#1222
·
YangXu1990uiuc
opened
Sep 25, 2026
Collaborator
·
·
·
Frontend 1.31.0
24
fix(bsa): stabilize SM100 forward softmax for large logits
#1221
·
Butterfingrz
opened
Sep 24, 2026
Contributor
·
·
1
Update Benchmarking Artifacts - backend 9.27.0.21, frontend ab43be0
#1219
·
brandonfzhang
opened
Sep 24, 2026
Collaborator
·
·
1
Reduce FROST THD host overhead with native argument binding
area:frost
area:global_attention
cat-enhancements
orig-nv-eng
#1217
·
YangXu1990uiuc
opened
Sep 24, 2026
Collaborator
·
·
·
Frontend 1.31.0
9
[FROST] Prepare SM100 D128 FP8 forward host launches
area:frost
area:global_attention
cat-perf-bug
orig-nv-eng
#1216
·
YangXu1990uiuc
opened
Sep 24, 2026
Collaborator
·
·
·
Frontend 1.31.0
26
frost(sdpa): paged KV on the SM100 MXFP8 flavors (F8_128x4 descale pools, page_size % 128)
area:frost
area:global_attention
cat-feature
orig-nv-eng
#1214
·
Adnios
opened
Sep 24, 2026
Collaborator
·
·
·
Frontend 1.30.0
13
BSA: optimize SM90 blk64 forward performance
#1213
·
Butterfingrz
opened
Sep 24, 2026
Contributor
·
·
11
frost(sdpa/bwd): sm107 (Rubin) d=256 backward -- bf16/fp16 and per-tensor FP8 kernels ported from the reference CTM pipelines, engine rows, tests, perf
area:frost
cat-feature
mod-cutedsl
mod-frost
orig-nv-eng
#1212
·
RomanAnders90
opened
Sep 24, 2026
Contributor
·
·
2
test(sdpa/frost): run the Stats assertion of test_fp8_thd_sliding_window before _check (follow-up to #1209)
area:frost
cat-cleanup
mod-frost
orig-nv-eng
#1211
·
RomanAnders90
opened
Sep 24, 2026
Contributor
·
·
·
Frontend 1.30.0
1
Add C++ and Python APIs for custom numerical weight dequantization
#1210
·
scottyokim
opened
Sep 24, 2026
·
·
1
Add bounded causal GQA backward for SM107
#1207
·
layalir
opened
Sep 23, 2026
·
·
7
[DSv4.1] Reduce Engram backward register pressure on SM100
area:frost
cat-enhancements
orig-nv-eng
#1204
·
YangXu1990uiuc
opened
Sep 23, 2026
Collaborator
·
·
·
Frontend 1.31.0
4
FROST GEMM heuristic: decline split-K on block-scale chains; keep block-scale MMA-K at 32 bytes
#1201
·
brandonfzhang
opened
Sep 23, 2026
Collaborator
·
·
·
Frontend 1.31.0
3
[DSv4.1] Support native N-major BF16 expert input gradients
area:frost
cat-feature
orig-nv-eng
#1199
·
YangXu1990uiuc
opened
Sep 23, 2026
Collaborator
·
·
·
Frontend 1.31.0
7
Add aligned HCA backward for GB300 and Rubin
area:sparse_attention
#1198
·
layalir
opened
Sep 22, 2026
·
·
·
Frontend 1.31.0
16
Rule 8 (gated_attention_block): quant words in a workspace slot, per-device graph handle with locked re-stream
area:global_attention
cat-cleanup
orig-nv-eng
#1186
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
64
Rule 8 (HSTU): caller-owned workspace for block-sparse metadata and the bwd fp32 accumulator, R5 declines, compile-time fakes
area:global_attention
cat-cleanup
orig-nv-eng
#1185
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
83
Rule 8 (flex_attention): compile() from fake tensors, no device memory at compile
area:global_attention
cat-cleanup
orig-nv-eng
#1184
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
72
Rule 8 (NSA/CSA): plan-time host ints, SWA caller-owned workspace and THD ragged offsets; wrappers unchanged
area:sparse_attention
cat-cleanup
orig-nv-eng
#1183
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
87
Rule 8 (DSA): plan-style classes own no device memory -- required outputs, caller-owned workspaces, plan-time host ints; wrappers unchanged
area:sparse_attention
cat-cleanup
op: DSA
orig-nv-eng
#1182
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
68
Rule 8 core: caller-owned workspaces and no host blocking in the graph API, FROST MoE GEMM and SDPA, grouped GEMM and Hopper KDA; R9 detector and recipes R10/R11
area:frost
area:global_attention
area:linear_attention
cat-cleanup
orig-nv-eng
#1181
·
YangXu1990uiuc
opened
Sep 22, 2026
Collaborator
·
·
·
Frontend 1.31.0
88
Specialize top-k row stride for SM100 DSA H16/H32/H96
area:sparse_attention
#1166
·
icavan
opened
Sep 21, 2026
Contributor
·
·
·
Frontend 1.31.0
6
BSA: add native SM120 blk128 Sage FP8 forward
area:sparse_attention
cat-feature
orig-nv-eng
#1163
·
tiffany940107
opened
Sep 21, 2026
Contributor
·
·
·
Frontend 1.31.0
3
Previous
1
2
3
4
Next
You can’t perform that action at this time.