CUDA backend configuration#

Pass a nested mapping directly to tvm.compile(..., backend_config=...) or T.device_entry(..., backend_config=...). Each backend owns its configuration; CUDA currently provides these fixed keys:

CUDA keys#

Key

Type

Meaning

arch

str

One real architecture, such as sm_100a; used during lowering and compilation.

compiler

"nvcc" or "nvrtc"

CUDA frontend; defaults to NVRTC.

target_format

"ptx", "cubin", or "fatbin"

NVRTC supports PTX and cubin; NVCC also supports fatbin.

nvcc

list[str]

Native NVCC arguments; defaults to ["--use_fast_math"].

nvrtc

list[str]

Native NVRTC arguments; defaults to ["--use_fast_math"].

ptxas

list[str]

Native assembler arguments forwarded through the selected frontend. Defaults to ["-v", "--warn-on-local-memory-usage", "--register-usage-level=10"].

BackendConfig is a TypedDict for completion and static typing. Its constructor returns an ordinary dictionary; all keys are optional. Compilation and entry construction validate and snapshot the supplied mapping. Native compiler arguments are not modeled as separate Python fields.

from tvm.backend.cuda import BackendConfig

cuda_config = BackendConfig(
    arch="sm_100a", compiler="nvrtc",
    nvrtc=["--use_fast_math", "--ftz=false"],
)
executable = tvm.compile(func, backend_config={"cuda": cuda_config})
class tvm.backend.cuda.BackendConfig#

Optional CUDA overrides. Toolchain arguments are passed as individual argv items.

arch: str#
compiler: Literal['nvcc', 'nvrtc']#
target_format: Literal['ptx', 'cubin', 'fatbin']#
nvcc: list[str]#
nvrtc: list[str]#
ptxas: list[str]#
class tvm.backend.cuda.compiler.CompilationResult(binary: bytes, target_format: str, config: dict)#

Device artifact and the effective settings used to produce it.

binary: bytes#

Alias for field number 0

target_format: str#

Alias for field number 1

config: dict#

Alias for field number 2

tvm.backend.cuda.compiler.compile_source(source: str, backend_config=None, *, target=None) → CompilationResult#

Compile source with per-backend overrides and optional Target/tag defaults.

When target is omitted, use the current CUDA Target or detect the device. Compiler logs are reported on failure and are not part of the artifact.

See Compiling and inspecting for target-tag defaults, inheritance, per-entry compilation, and artifact behavior. See Native CUDA compiler arguments for native NVCC/NVRTC/ptxas argument usage and official compiler references.