CUDA kernel configuration reference#
Both CUDA module launches and exported cuda_host wrappers use this registry.
LaunchConfig describes each launch; KernelAttributes describes compile-time
CUDA kernel attributes.
Optional launch attributes default to unspecified, preserving CUDA’s inherited defaults.
LaunchConfig#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
dim3 |
|
Grid dimensions in CTAs; an integer or one to three dimensions. |
|
dim3 |
|
Threads per CTA; an integer or one to three dimensions. |
|
dim3 |
|
CTAs per cluster. None differs from an explicit unit cluster. Encoder requires CUDA 12.0+. |
|
dim3 |
|
Preferred substitute cluster dimensions. Encoder requires CUDA 12.8+. |
|
int |
|
Dynamic shared memory bytes; None infers allocation requirements. |
|
handle |
|
CUDA stream; None uses the current tvm-ffi stream. |
|
bool |
|
Request a cooperative kernel launch. Encoder requires CUDA 12.0+. |
|
bool |
|
Enable programmatic dependent launch. Encoder requires CUDA 12.0+. |
|
enum |
|
Cluster scheduling preference. Values: |
|
int |
|
Launch priority; CUDA may clamp it to the supported range. Encoder requires CUDA 12.0+. |
|
enum |
|
Logical memory synchronization domain. Values: |
|
MemSyncDomainMap |
|
Mapping of logical to physical memory domains. Encoder requires CUDA 12.0+. |
|
AccessPolicyWindow |
|
Per-launch L2 access policy window. Encoder requires CUDA 12.0+. |
|
int |
|
Preferred shared-memory carveout percentage, 0 through 100. Encoder requires CUDA 12.8+. |
|
bool |
|
Best-effort NVLink utilization scheduling hint. Encoder requires CUDA 13.2+. |
|
enum |
|
Override cluster portability policy for this launch. Values: |
|
enum |
|
Override shared-memory resource mode; oversized modes require CUDA 13.4. Values: |
|
ProgrammaticEvent |
|
Record a programmatic dependency event. Encoder requires CUDA 12.0+. |
|
LaunchCompletionEvent |
|
Record an event associated with blocks beginning execution. Encoder requires CUDA 12.4+. |
KernelAttributes#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
int |
|
Second launch-bounds operand: minimum resident CTAs per SM. |
|
int |
|
Third launch-bounds operand; requires min_blocks_per_sm. |
|
int |
|
Emit __maxnreg__; incompatible with explicit launch bounds and required block size. |
|
bool |
|
Fix block and cluster dimensions at compilation using __block_size__. |
MemSyncDomainMap#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
int |
|
Physical domain for the default logical domain. |
|
int |
|
Physical domain for the remote logical domain. |
AccessPolicyWindow#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
handle |
|
Device pointer to the access-policy window. |
|
int |
|
Window size in bytes. |
|
float |
|
Fraction of accesses receiving the hit policy. |
|
enum |
|
Cache policy for hits. Values: |
|
enum |
|
Cache policy for misses. Values: |
ProgrammaticEvent#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
handle |
|
Caller-owned, timing-disabled CUDA event. |
|
int |
|
Event record flags; external-event recording is unsupported. |
|
bool |
|
Trigger when each block starts. |
LaunchCompletionEvent#
Field |
Type |
Default |
Description |
|---|---|---|---|
|
handle |
|
Caller-owned, timing-disabled CUDA event. |
|
int |
|
Event record flags; external-event recording is unsupported. |
Stream-only synchronization policy and graph-only device-updatable nodes are outside this kernel-launch API. CUDA checks device-specific support and resource limits.