.. Licensed to the Apache Software Foundation (ASF) under one
   or more contributor license agreements. See the NOTICE file
   distributed with this work for additional information
   regarding copyright ownership. The ASF licenses this file
   to you under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

   http://www.apache.org/licenses/LICENSE-2.0

   Unless required by applicable law or agreed to in writing,
   software distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
   limitations under the License.

.. Generated by launch/generate.py; edit table.py instead.

CUDA kernel configuration reference
===================================

Both CUDA module launches and exported ``cuda_host`` wrappers use this registry.
``LaunchConfig`` describes each launch; ``KernelAttributes`` describes compile-time
CUDA kernel attributes.
Optional launch attributes default to unspecified, preserving CUDA's inherited defaults.

LaunchConfig
------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``grid``
     - dim3
     - ``required``
     - Grid dimensions in CTAs; an integer or one to three dimensions.
   * - ``block``
     - dim3
     - ``required``
     - Threads per CTA; an integer or one to three dimensions.
   * - ``cluster``
     - dim3
     - ``None``
     - CTAs per cluster. None differs from an explicit unit cluster. Encoder requires CUDA 12.0+.
   * - ``preferred_cluster``
     - dim3
     - ``None``
     - Preferred substitute cluster dimensions. Encoder requires CUDA 12.8+.
   * - ``dynamic_smem_bytes``
     - int
     - ``None``
     - Dynamic shared memory bytes; None infers allocation requirements.
   * - ``stream``
     - handle
     - ``None``
     - CUDA stream; None uses the current tvm-ffi stream.
   * - ``cooperative``
     - bool
     - ``None``
     - Request a cooperative kernel launch. Encoder requires CUDA 12.0+.
   * - ``programmatic_stream_serialization``
     - bool
     - ``None``
     - Enable programmatic dependent launch. Encoder requires CUDA 12.0+.
   * - ``cluster_scheduling_policy``
     - enum
     - ``None``
     - Cluster scheduling preference. Values: ``default``, ``spread``, ``load_balancing``. Encoder requires CUDA 12.0+.
   * - ``priority``
     - int
     - ``None``
     - Launch priority; CUDA may clamp it to the supported range. Encoder requires CUDA 12.0+.
   * - ``mem_sync_domain``
     - enum
     - ``None``
     - Logical memory synchronization domain. Values: ``default``, ``remote``. Encoder requires CUDA 12.0+.
   * - ``mem_sync_domain_map``
     - MemSyncDomainMap
     - ``None``
     - Mapping of logical to physical memory domains. Encoder requires CUDA 12.0+.
   * - ``access_policy_window``
     - AccessPolicyWindow
     - ``None``
     - Per-launch L2 access policy window. Encoder requires CUDA 12.0+.
   * - ``preferred_shared_memory_carveout``
     - int
     - ``None``
     - Preferred shared-memory carveout percentage, 0 through 100. Encoder requires CUDA 12.8+.
   * - ``nvlink_util_centric_scheduling``
     - bool
     - ``None``
     - Best-effort NVLink utilization scheduling hint. Encoder requires CUDA 13.2+.
   * - ``portable_cluster_size_mode``
     - enum
     - ``None``
     - Override cluster portability policy for this launch. Values: ``default``, ``require_portable``, ``allow_non_portable``. Encoder requires CUDA 13.2+.
   * - ``shared_memory_mode``
     - enum
     - ``None``
     - Override shared-memory resource mode; oversized modes require CUDA 13.4. Values: ``default``, ``require_portable``, ``allow_non_portable``, ``allow_oversized``, ``prefer_oversized``. Encoder requires CUDA 13.2+.
   * - ``programmatic_event``
     - ProgrammaticEvent
     - ``None``
     - Record a programmatic dependency event. Encoder requires CUDA 12.0+.
   * - ``launch_completion_event``
     - LaunchCompletionEvent
     - ``None``
     - Record an event associated with blocks beginning execution. Encoder requires CUDA 12.4+.

KernelAttributes
----------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``min_blocks_per_sm``
     - int
     - ``None``
     - Second launch-bounds operand: minimum resident CTAs per SM.
   * - ``max_blocks_per_cluster``
     - int
     - ``None``
     - Third launch-bounds operand; requires min_blocks_per_sm.
   * - ``max_registers_per_thread``
     - int
     - ``None``
     - Emit __maxnreg__; incompatible with explicit launch bounds and required block size.
   * - ``required_block_size``
     - bool
     - ``False``
     - Fix block and cluster dimensions at compilation using __block_size__.

MemSyncDomainMap
----------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``default``
     - int
     - ``0``
     - Physical domain for the default logical domain.
   * - ``remote``
     - int
     - ``1``
     - Physical domain for the remote logical domain.

AccessPolicyWindow
------------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``base_ptr``
     - handle
     - ``required``
     - Device pointer to the access-policy window.
   * - ``num_bytes``
     - int
     - ``required``
     - Window size in bytes.
   * - ``hit_ratio``
     - float
     - ``1.0``
     - Fraction of accesses receiving the hit policy.
   * - ``hit_prop``
     - enum
     - ``'persisting'``
     - Cache policy for hits. Values: ``normal``, ``streaming``, ``persisting``.
   * - ``miss_prop``
     - enum
     - ``'normal'``
     - Cache policy for misses. Values: ``normal``, ``streaming``, ``persisting``.

ProgrammaticEvent
-----------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``event``
     - handle
     - ``required``
     - Caller-owned, timing-disabled CUDA event.
   * - ``flags``
     - int
     - ``0``
     - Event record flags; external-event recording is unsupported.
   * - ``trigger_at_block_start``
     - bool
     - ``False``
     - Trigger when each block starts.

LaunchCompletionEvent
---------------------

.. list-table::
   :header-rows: 1

   * - Field
     - Type
     - Default
     - Description
   * - ``event``
     - handle
     - ``required``
     - Caller-owned, timing-disabled CUDA event.
   * - ``flags``
     - int
     - ``0``
     - Event record flags; external-event recording is unsupported.

Stream-only synchronization policy and graph-only device-updatable nodes are
outside this kernel-launch API. CUDA checks device-specific support and resource limits.
