Native CUDA compiler arguments#
CUDA uses six fixed configuration keys; see CUDA backend configuration. Pass other
compiler settings as native arguments in the nvcc, nvrtc, or ptxas
lists:
backend_config = {"cuda": {
"arch": "sm_100a",
"compiler": "nvcc",
"nvcc": ["--use_fast_math", "--ftz=false", "--std=c++20", "-I/include"],
"nvrtc": ["--use_fast_math", "--ftz=false", "--std=c++20", "-I/include"],
"ptxas": ["-O3", "--register-usage-level=10"],
}}
Each list contains argv items, without shell parsing. The selected frontend
receives its own list and the forwarded ptxas list. Each list replaces the
inherited list completely: nvrtc=[] clears the default fast-math flag and
ptxas=[] clears the default assembler arguments.
TVM validates fixed keys and types. Architecture, output format, and toolchain routing are controlled by the corresponding configuration keys and the compiler adapter. Native flags that override those controls are reserved. The selected compiler validates other arguments, including their values and version support; new compiler flags can be used directly through these lists.
TVM adds the integration arguments needed for CUDA headers, architecture, output retrieval, and NVSHMEM device linking. Supported outputs are PTX, cubin, and NVCC fatbin. Additional artifact or linking workflows require backend support beyond passing their compiler flags.
Consult the NVCC compiler options and NVRTC compilation options for the
installed toolkit version. Use ptxas --help to inspect its assembler options.