Core TVMScript#

TIRx kernels use tvm.script.tirx for the parser and core IR builders:

from tvm.script import tirx as Tx

Tx.alloc_tensor(...)

Tile primitives and backend-specific namespaces are documented separately in Tensor Instruction Authoring API, CUDA Authoring and Support APIs, and Direct PTX Instructions. For the relationship between these authoring layers and TIRx IR, see The Programming Model.

Parser entry points#

Public canonical TVMScript dialect namespace.

tvm.script.tirx.function(function: LambdaType | None = None, **options: Any) → Any

Parse a Python function into a function of the selected IR language variant.

Parameters:
  • function (Callable, optional) – The function to be parsed. May be omitted to use the decorator with keyword options, such as @T.function(private=True).

  • private (bool, optional) – Whether the function should be treated as private. A private function has no global symbol attribute; a public function has a global symbol matching its name. Defaults to False.

  • check_well_formed (bool, optional) – Whether to check that the constructed function is well formed. Defaults to True.

  • persistent (bool, optional) – For T.function, mark the resulting function as a persistent kernel. Defaults to False. See tvm.tirx.script.ir_builder.function().

  • pure (bool, optional) – For R.function, declare whether the function is pure, meaning that it has no observable side effects. Defaults to True. See tvm.relax.script.ir_builder.function_().

  • **options – Keyword options are forwarded to the selected language variant’s function_() hook, except check_well_formed, which controls parser validation. Options supported by only one language variant are not shared between T.function and R.function.

Returns:

result – The parsed function, or a decorator when function is omitted. Class members retain their Python functions until the enclosing module is constructed.

Return type:

Function or relax.Function or Callable

tvm.script.tirx.inline(function: LambdaType | None = None, **options: Any) → Callable[[...], Any]

Decorate a helper that constructs IR in its caller’s active frames.

Parameters:
  • function (Callable, optional) – The helper function. May be omitted to supply keyword options.

  • hygienic (bool, optional) – Whether the helper resolves symbols in its definition environment instead of its calling environment. Defaults to True. T.macro and R.macro capture values at definition time; T.inline refreshes captured closure cells when called.

  • **options – Keyword configuration for this helper decorator. hygienic controls name lookup as described above; other options are accepted but are not consumed or forwarded to builder hooks.

Returns:

result – The construction helper, or a decorator when function is omitted.

Return type:

Callable

Notes

T.inline follows Python lexical scoping with late binding of captured closure cells. Its return statements produce Python values, as do those of R.macro. T.macro emits returns in the active primitive function.

Examples

An inline helper can read values from its enclosing scope:

import tvm
from tvm.script import tirx as T
x_value = 128

@T.inline
def capture(A, B):
    B[()] = A[x_value]  # x_value resolved from enclosing scope

@T.function
def use(A: T.Tensor((1024,), "int32"), B: T.Tensor((), "int32")) -> None:
    capture(A, B)       # Produces B[()] = A[128]
tvm.script.tirx.jit(func: LambdaType | None = None, *, private: bool = False, check_well_formed: bool = True, persistent: bool = False) → TIRJit | Callable[[LambdaType], TIRJit]

Decorator: capture the kernel and defer parsing until .specialize().

Use @T.jit (instead of @T.function) when the kernel takes compile-time parameters annotated with T.constexpr or runtime parameters that may be removed with T.Optional. The resulting object exposes .specialize(**const_args), which returns a tvm.tirx.Function.

Parameters:
  • func (types.FunctionType or None, optional) – Function supplied by a definition-site @T.jit application. None, the default, returns a decorator for @T.jit(**options).

  • private (bool, optional) – Omit the function’s public global symbol when True. Default is False.

  • check_well_formed (bool, optional) – Validate each constructed specialization. Default is True; this option controls the parser separately from the function builder kwargs.

  • persistent (bool, optional) – Mark constructed specializations as persistent kernels. Default is False.

Returns:

Deferred kernel, or its definition-site decorator when func is None.

Return type:

TIRJit or Callable[[types.FunctionType], TIRJit]

Raises:
  • TypeError – If the decorated value is not a Python function.

  • SyntaxError – If application occurs after definition or uses an unsupported bare or preconfigured callable alias instead of a qualified namespace decorator.

  • OSError – If the function’s source cannot be recovered.

Examples

Specialize compile-time dimensions before compiling the kernel:

from __future__ import annotations

from tvm.script import tirx as T

@T.jit
def add(
    A: T.Tensor((N,), "float32"),
    B: T.Tensor((N,), "float32"),
    *,
    N: T.constexpr,
):
    for i in T.serial(N):
        B[i] = A[i] + 1.0

kernel = add.specialize(N=1024)  # returns a Function

@T.jit
def guarded(
    optional: T.Optional(T.Tensor((1,), "int32")),
    output: T.Tensor((1,), "int32"),
):
    if T.constexpr(optional is not None):
        output[0] = optional[0]
    else:
        output[0] = 0

present = guarded.specialize()
absent = guarded.specialize(optional=None)
tvm.script.tirx.macro(function: LambdaType | None = None, **options: Any) → Callable[[...], Any]

Decorate a helper that constructs IR in its caller’s active frames.

Parameters:
  • function (Callable, optional) – The helper function. May be omitted to supply keyword options.

  • hygienic (bool, optional) – Whether the helper resolves symbols in its definition environment instead of its calling environment. Defaults to True. T.macro and R.macro capture values at definition time; T.inline refreshes captured closure cells when called.

  • **options – Keyword configuration for this helper decorator. hygienic controls name lookup as described above; other options are accepted but are not consumed or forwarded to builder hooks.

Returns:

result – The construction helper, or a decorator when function is omitted.

Return type:

Callable

Notes

T.inline follows Python lexical scoping with late binding of captured closure cells. Its return statements produce Python values, as do those of R.macro. T.macro emits returns in the active primitive function.

Examples

An inline helper can read values from its enclosing scope:

import tvm
from tvm.script import tirx as T
x_value = 128

@T.inline
def capture(A, B):
    B[()] = A[x_value]  # x_value resolved from enclosing scope

@T.function
def use(A: T.Tensor((1024,), "int32"), B: T.Tensor((), "int32")) -> None:
    capture(A, B)       # Produces B[()] = A[128]

Core IR builder#

Common statement frames and construction helpers live in tvm.script.ir_builder.frame and tvm.script.ir_builder.stmt. The TIRx namespace re-exports them alongside its function, tensor/layout, and execution scope builders. Native shared statement builders live under include/tvm/script/ir_builder and src/script/ir_builder; dialect-specific thread-placement and tensor alias policies remain in the TIRx/S-TIR builders.

Concrete TIRx types, tensors, allocations and construction metadata.

class tvm.tirx.script.ir_builder.ir.Axis(*args: Any, **kwargs: Any)

Layout axis wrapper.

property name: str

Return the canonical axis name.

classmethod get(name: str) → Axis

Get or create the axis singleton named name.

Unknown names are registered without thread or memory attributes.

is_thread() → bool

Check if the axis is a thread axis.

is_memory() → bool

Check if the axis is a memory axis.

property backend: str | None

Backend owning this axis, or None for a generic axis.

get_scope() → ExecScope | None

Get the scope of the axis.

get_subscope() → ExecScope | None

Get the subscope of the axis.

class tvm.tirx.script.ir_builder.ir.ComposeLayout(per_element: int, swizzle_len: int, atom_len: int, tile_layout: TileLayout, swizzle_inner: bool = True)

A memory layout that swizzles a tile layout.

per_element / swizzle_len / atom_len / swizzle_inner carry the swizzle (formerly the standalone SwizzleLayout); tile_layout is the tiled memory map the swizzle is applied to. A bare swizzle is a ComposeLayout over a trivial identity tile.

class tvm.tirx.script.ir_builder.ir.DtypeConstructor(ffi_name: str, dtype_str: str)

Callable + subscriptable dtype object.

Replaces the plain functions previously returned by func_gen.

  • T.float32() — same FFI call as before (returns tirx.Var).

  • T.float32[N] — returns LocalVectorAnnotation("float32", (N,)).

  • T.float32[M, N] — returns LocalVectorAnnotation("float32", (M, N)).

  • x: T.float32 — parser calls this object, gets a tirx.Var.

class tvm.tirx.script.ir_builder.ir.ExecScope(backend: str, name: str)

A node in one backend’s registered execution hierarchy.

class tvm.tirx.script.ir_builder.ir.FloatImm(dtype: str | PrimType, value: float, loc: Location = UnknownLoc())

Float constant.

Parameters:
  • dtype (str) – The data type

  • value (float) – The constant value.

  • loc (Location, optional) – The location of this expression in the source code.

class tvm.tirx.script.ir_builder.ir.IntImm(dtype: str | PrimType, value: int, loc: Location = UnknownLoc())

Int constant.

Parameters:
  • dtype (str) – The data type

  • value (int) – The constant value.

  • loc (Location, optional) – The location of this expression in the source code.

class tvm.tirx.script.ir_builder.ir.Iter(extent: Expr, stride: Expr, axis: Axis | str)

A memory layout that tiles data across devices.

tvm.tirx.script.ir_builder.ir.Lambda(parameter_types, function, *, ret_type=None)

Build a shared staging lambda from explicit types and a Python callable.

Scalar constructors such as T.float32 may be used as parameter types. An optional return annotation checks the body type without inserting casts.

class tvm.tirx.script.ir_builder.ir.Layout
verify_well_formed() → bool

Verify if the layout is well-formed.

Returns:

True if the layout is well-formed, False otherwise

Return type:

bool

size(axis_name: str | None = None)

Get the size of the layout.

Parameters:

axis_name (Optional[str]) – The name of the axis to get the size of. If not provided, the default input size will be returned.

span(axis_name: str | None = None)

Get the span of the layout.

Parameters:

axis_name (Optional[str]) – The name of the axis to get the span of. If not provided, the default span will be returned.

apply(*coord: list[Expr], shape: list[Expr] | None = None) → dict[str, Expr]

Apply the layout on the input coordinate and get the mapped output.

Input cases: - coord is a single element -> will be treated as a 1D coordinate - coord is a list of elements -> will be treated as a multi-dimensional coordinate - shape is provided -> turn the coord with shape into a 1D coordinate - shape is not provided -> use the default shape

Returns:

The mapped output (axis name -> value on the axis)

Return type:

Dict[str, Expr]

apply_to_shape(coord: list[Expr], input_shape: list[Expr]) → list[Expr]

Compute the per-shard value that each shard would take if coord were interpreted against input_shape.

Tries self.group(input_shape) first. On success, each group owns exactly one input_shape entry, so coord[d] can be split within that group’s shard extents (bounds stay local to one input dim — simpler analyzer simplification, no cross-dim complications).

Falls back to FlattenCoord(coord, input_shape) + SplitCoord on self’s raw shard shape when the group call fails (e.g. when input_shape does not align with the layout’s factor boundaries).

Returns a list of length len(self.shard); each entry is the value that shard would iterate.

canonicalize() → Layout

Canonicalize the layout by simplifying and fusing iterators where possible.

Returns:

The canonicalized layout

Return type:

Layout

tile(outer: TileLayout, outer_shape: list[Expr], inner_shape: list[Expr]) → TileLayout | ComposeLayout

Tile the current layout with an outer layout.

Parameters:
  • outer (TileLayout) – The outer layout to tile with

  • outer_shape (List[Expr]) – The shape of the outer layout

  • inner_shape (List[Expr]) – The shape of the inner layout

Returns:

The resulting tiled layout

Return type:

Union[TileLayout, ComposeLayout]

direct_sum(left: TileLayout, left_shape: list[Expr], right_shape: list[Expr]) → TileLayout | ComposeLayout

Direct-sum on the tiling domain (unscaled composition): A + B.

This layout is treated as the right addend B grouped by right_shape. The left layout is treated as A grouped by left_shape. The resulting layout is evaluated over the interleaved domain S_A ⊗ S_B, without span scaling (unlike tiling).

is_tile_inner(tile_layout: TileLayout | ComposeLayout, tiled_shape: list[Expr], inner_shape: list[Expr]) → TileLayout | None

Check if a layout is the inner layout of a tiled layout.

Parameters:
  • tile_layout (Union[TileLayout, ComposeLayout]) – The tiled layout to check

  • tiled_shape (List[Expr]) – The shape of the tiled layout

  • inner_shape (List[Expr]) – The shape of the inner layout

Returns:

The outer layout if it is the inner layout of the tiled layout, None otherwise

Return type:

Optional[TileLayout]

is_tile_outer(tile_layout: TileLayout | ComposeLayout, tiled_shape: list[Expr], outer_shape: list[Expr]) → Layout | None

Check if a layout is the outer layout of a tiled layout.

Parameters:
  • tile_layout (Union[TileLayout, ComposeLayout]) – The tiled layout to check

  • tiled_shape (List[Expr]) – The shape of the tiled layout

  • outer_shape (List[Expr]) – The shape of the outer layout

Returns:

The inner layout if it is the outer layout of the tiled layout, None otherwise

Return type:

Optional[Layout]

is_direct_sum_right(sum_layout: TileLayout | ComposeLayout, interleaved_shape: list[Expr], right_shape: list[Expr]) → TileLayout | None

Check if this layout is the right addend B in a direct-sum A + B.

Returns the left addend A if recognized, otherwise None.

is_direct_sum_left(sum_layout: TileLayout | ComposeLayout, interleaved_shape: list[Expr], left_shape: list[Expr]) → Layout | None

Check if this layout is the left addend A in a direct-sum A + B.

Returns the right addend B if recognized, otherwise None.

slice(shape: list[Expr], region: list[tuple[Expr, Expr]]) → Layout | None

Slice the layout with a given shape and region.

Parameters:
  • shape (List[Expr]) – The shape of the layout

  • region (List[Tuple[Expr, Expr], tvm.ir.Range]) – The region to slice, each element is (begin, end)

Returns:

The sliced layout, or None if slicing is not possible

Return type:

Optional[Layout]

tile_to(to_shape: list[Expr], current_shape: list[Expr]) → Layout

Tile the current layout to the given shape.

Parameters:
  • to_shape (List[Expr]) – The shape to tile to

  • current_shape (List[Expr]) – The current shape of the layout

is_swizzle() → bool

Check if the layout is a bare swizzle (ComposeLayout over a trivial tile).

is_trivial() → bool

Check if the layout is trivial.

unpack(num: int) → Layout

Unpack the layout, where a single element in the layout is unpacked into num contiguous elements.

Parameters:

num (int) – The number of elements to unpack into

Returns:

The unpacked layout

Return type:

Layout

broadcast(num: int, position: int = -1, axis: 'Axis' | str = 'm') → Layout

Insert a stride-0 broadcast dim of extent num at position.

position follows Python list-insert semantics (negative indices count from the end; -1 appends after the last shard dim). The new dim has stride 0 — accessing along it doesn’t move the byte offset, so the same physical element is “seen” num times.

Useful for layouts where a consumer reads the same SMEM datum multiple times (e.g. sf_reuse over MMA-K steps).

pack(num: int) → Layout

Pack the layout, where num contiguous elements in the layout are packed into a single element.

Parameters:

num (int) – The number of elements to pack into

Returns:

The packed layout

Return type:

Layout

class tvm.tirx.script.ir_builder.ir.LocalVectorAnnotation(dtype: str, shape: tuple)

Marker for local vector/tensor allocation via type annotation subscript.

Created when a DtypeConstructor is subscripted, e.g. T.float32[N] or T.float32[M, N]. The declaration protocol recognizes this annotation and allocates local storage with T.alloc_local(shape=..., dtype=...).

tvm.tirx.script.ir_builder.ir.Ptr(element_type, storage_scope='global', *, loc: LocationEntry | Location = UnknownLoc())

Construct a pointer type from an element type or its annotation constructor.

For example, Ptr(int32) and Ptr(TensorMap) construct pointers in global scope; Ptr(int32, "shared") selects shared scope. Use I.Var(name, ty) to construct a variable of this type. String element dtypes remain accepted, as in Ptr("float32", "shared").

class tvm.tirx.script.ir_builder.ir.Range(begin: Expr, end: Expr | None = None, loc: Location = UnknownLoc())

Represent a range in TVM.

You do not need to create a Range explicitly. Python lists and tuples will be converted automatically to a Range in API functions.

Parameters:
  • begin (Expr) – The begin value of the range when end is None. Otherwise it is the length of the range.

  • end (Optional[Expr]) – The end value of the range.

  • loc (Location, optional) – The location of this node in the source code.

Note

The constructor creates the range [begin, end) if the end argument is not None. Otherwise, it creates [0, begin).

static from_min_extent(min_value: Expr, extent: Expr, loc: Location = UnknownLoc()) → Range

Construct a Range by min and extent.

This constructs a range in [min_value, min_value + extent)

Parameters:
  • min_value (Expr) – The minimum value of the range.

  • extent (Expr) – The extent of the range.

  • loc (Location, optional) – The location of this node in the source code.

Returns:

rng – The constructed range.

Return type:

Range

tvm.tirx.script.ir_builder.ir.Tensor

alias of _tensor_type

tvm.tirx.script.ir_builder.ir.TensorLoad(tensor: Var, indices: list[Expr], loc: Location = UnknownLoc()) → TensorLoad

Construct a validated tensor load.

Parameters:
  • tensor (tirx.Var) – The tensor to be loaded.

  • indices (List[Expr]) – The tensor indices to load values from.

  • loc (Location, optional) – The location of this expression in the source code.

tvm.tirx.script.ir_builder.ir.TensorMap

alias of TensorMapType

class tvm.tirx.script.ir_builder.ir.TileLayout(spec: _LayoutSpec)

A memory layout that tiles data across devices.

static from_iters(shard: Sequence[Iter] = (), replica: Sequence[Iter] = (), offset: dict[Axis | str, Expr] | None = None) → TileLayout

Construct a TileLayout from pre-built Iter objects.

is_trivial() → bool

Check if the layout is trivial.

group(shape: list[Expr]) → tuple[Layout, list[int]]

Group the current layout by the given shape.

Parameters:

shape (List[Expr]) – The shape to group by

Returns:

The grouped layout and the separators

Return type:

Tuple[Layout, List[int]]

group_many(shapes: Sequence[Sequence[Expr]]) → tuple[TileLayout, list[list[int]]]

Group the layout by the minimal common refinement of several shapes.

Repeated cumulative product boundaries are retained, so an extent-one dimension in any input shape becomes a real unit iterator in the refined layout. This operation only splits existing shard iterators; it does not canonicalize or reorder the layout.

Parameters:

shapes (Sequence[Sequence[Expr]]) – Logical shapes with provably equal total products.

Returns:

The commonly refined layout and one separator list per input shape.

Return type:

Tuple[TileLayout, List[List[int]]]

get_scope() → tuple[ExecScope, ExecScope] | None

Get the scope pair of the layout.

permute_dims(perm: list[int]) → TileLayout

Permute the dimensions of the layout.

permute_by_groups(seps: list[int], perm: list[int]) → TileLayout

Permute groups of shard iters defined by seps.

seps follows the convention of group()’s second return value: seps[0] == 0 and group i covers shard indices [seps[i], seps[i + 1]). The number of groups is len(seps) - 1.

Parameters:
  • seps (list[int]) – Group boundary positions in the shard list.

  • perm (list[int]) – Permutation of range(len(seps) - 1) selecting the new group order.

tvm.tirx.script.ir_builder.ir.Tuple(*fields: Type) → Type

Construct a tuple type for a TIRx function or binding annotation.

class tvm.tirx.script.ir_builder.ir.Var(name: str | None = None, ty: Type | str | None = None, loc: Location = UnknownLoc(), *, name_hint: str | None = None)

A canonical local variable in the IR.

Parameters:
  • name (str) – The name of the variable.

  • ty (Optional[Type or str]) – The exact type of the variable. A string denotes a primitive dtype.

  • loc (Location, optional) – Location that points to the original source code.

tvm.tirx.script.ir_builder.ir.alloc_cast_frag(src, dtype)

Allocate a register frag holding src value-cast to dtype.

Inherits src’s logical shape and its (lane, register) layout — only the element dtype changes — so Tx.cast(dst, src) is a per-thread element-wise cast with no cross-lane movement. .permute(...) the result to the axis order a downstream consumer (e.g. stmatrix via Tx.copy(dispatch="ldstmatrix")) expects.

Parameters:
  • src (tirx.Var) – Source register frag (e.g. from alloc_tcgen05_ldst_frag).

  • dtype (str) – Destination element dtype.

Returns:

Fresh local frag, src.shape shaped, src.layout, dtype-cast.

Return type:

tirx.Var

tvm.tirx.script.ir_builder.ir.alloc_local(shape, dtype='float32', data=None, strides=None, elem_offset=None, byte_offset=None, *, scope='local', align=-1, offset_factor=0, layout=<MISSING>, allocated_addr=None, annotations=None)

Implement separately named allocation conveniences using the canonical Op.

tvm.tirx.script.ir_builder.ir.alloc_scalar(dtype: str = 'float32', scope: str = 'global', *, annotations: dict[str, Any] | None = None) → TensorLoad

Allocate a zero-dimensional tensor (scalar), with optional allocation annotations.

tvm.tirx.script.ir_builder.ir.alloc_shared(shape, dtype='float32', data=None, strides=None, elem_offset=None, byte_offset=None, *, scope='shared', align=-1, offset_factor=0, layout=<MISSING>, allocated_addr=None, annotations=None)

Implement separately named allocation conveniences using the canonical Op.

tvm.tirx.script.ir_builder.ir.alloc_tensor(placement=None, *, ty_args, attrs=None, annotations: dict[str, Any] | None = None, loc: Location = UnknownLoc(), ty=None) → Call

Allocate the tensor described by one explicit Tensor type argument.

Optional placement is an ordinary tuple operand. Allocation annotations belong to attrs; annotations accepts a Python dictionary whose scalar values are normalized to typed IR constants.

tvm.tirx.script.ir_builder.ir.boolean(expr: Expr | None = None) → Expr

Construct a new tirx.Var with type boolean or cast expression to type boolean.

Parameters:

expr (Expr) – The expression to be cast.

Returns:

res – The new tirx.Var with type boolean or casted expression with type boolean.

Return type:

Expr

tvm.tirx.script.ir_builder.ir.decl_scalar(dtype, data, scope, elem_offset=None, byte_offset=None) → TensorLoad

Declare a zero-dimensional tensor (scalar) from a pointer.

tvm.tirx.script.ir_builder.ir.decl_tensor(data, *, ty_args, attrs=None, loc: Location = UnknownLoc(), ty=None) → Call

Declare a pointer-backed tensor with one explicit Tensor type argument.

tvm.tirx.script.ir_builder.ir.handle(dtype: str | None = None, storage_scope: str = 'global') → Var

Create a TIR var that represents a pointer.

Parameters:
  • dtype (str | None) – The data type of the pointer. If omitted, construct an opaque handle.

  • storage_scope (str) – The storage scope of the pointer.

Returns:

res – The new tirx.Var with type handle or casted expression with type handle.

Return type:

Expr

tvm.tirx.script.ir_builder.ir.index_map(mapping: Callable, *, inverse_index_map: Callable | None = None, index_dtype: str = 'int64') → IndexMap

Create a TIR Index mapping

tvm.tirx.script.ir_builder.ir.local_scalar(dtype: str = 'float32', *, annotations: dict[str, Any] | None = None) → TensorLoad

Allocate a zero-dimensional tensor in local memory.

tvm.tirx.script.ir_builder.ir.meta_class(cls)

Decorator for utility classes used inside @T.function.

Instances of decorated classes are treated as parser meta values.

tvm.tirx.script.ir_builder.ir.meta_var(value: T) → T

Return a Python metadata value without binding, naming or relocating it.

Parameters:

value (T) – Any host or IR object, including an unpackable sequence.

Returns:

  • T – The exact input object. No frame is required, no IR is emitted, and existing names and source locations are retained.

  • .. code:: python – # Source and generated Python (the value retains its identity) value = I.meta_var(existing_value) a, b = I.meta_var((left, right))

tvm.tirx.script.ir_builder.ir.ptr(dtype: str, storage_scope: str = 'global') → Var

The pointer declaration function.

Parameters:
  • dtype (str) – The data type of the pointer.

  • storage_scope (str) – The storage scope of the pointer.

Returns:

res – The pointer.

Return type:

tirx.Var

tvm.tirx.script.ir_builder.ir.shared_scalar(dtype: str = 'float32', *, annotations: dict[str, Any] | None = None) → TensorLoad

Allocate a zero-dimensional tensor in shared memory.

tvm.tirx.script.ir_builder.ir.smem(shape, dtype='float32', data=None, strides=None, elem_offset=None, byte_offset=None, *, scope='shared', align=-1, offset_factor=0, layout=<MISSING>, allocated_addr=None, annotations=None)

Implement separately named allocation conveniences using the canonical Op.

tvm.tirx.script.ir_builder.ir.target(target_config: dict | str, host: dict | str | Target | None = None) → Target

Create a target

Parameters:
  • target_config (Union[Dict, str]) – The target configuration.

  • host (Optional[Union[Dict, str, Target]]) – The target configuration.

Returns:

res – The target.

Return type:

Target

tvm.tirx.script.ir_builder.ir.tmem(shape, dtype='float32', data=None, strides=None, elem_offset=None, byte_offset=None, *, scope='tmem', align=-1, offset_factor=0, layout=<MISSING>, allocated_addr=None, annotations=None)

Implement separately named allocation conveniences using the canonical Op.

tvm.tirx.script.ir_builder.ir.void(expr: Expr | None = None) → Expr

Construct a new tirx.Var with type void or cast expression to type void.

Parameters:

expr (Expr) – The expression to be cast.

Returns:

res – The new tirx.Var with type void or casted expression with type void.

Return type:

Expr

tvm.tirx.script.ir_builder.ir.wg_reg_tile(elem_per_thread: int, dtype: str = 'float32') → Var

Warpgroup-wide (128, elem_per_thread) register tile in local scope.

Sugar for the recurring pattern:

T.alloc_tensor(ty_args=[T.Tensor(
    (128, elem_per_thread), dtype,
    layout=wg_local_layout(elem_per_thread), scope="local",
)])

Used to stage a tcgen05 load: each of the 128 threads in a warpgroup owns one row of elem_per_thread contiguous elements.

class tvm.tirx.script.ir_builder.ir.LetAnnotation(type_spec=None)#

Marker used by Tx.let and Tx.let[dtype] annotations to construct an explicit LetStmt.

tvm.tirx.script.ir_builder.ir.alloc_tcgen05_ldst_frag(instr_shape, tensor_shape, dtype)#

Allocate a local register fragment whose layout matches a tcgen05.{ld,st} atom. instr_shape accepts "32x32b", "16x64b", "16x128b", or "16x256b". For example, a two-CTA Layout-B accumulator and its readback fragment can be allocated as:

C = tmem_pool.alloc_tcgen05_mma_D(
    (64, 128), "float32", M=128, cta_group=2)
frag = Tx.alloc_tcgen05_ldst_frag("32x32b", (64, 128), "float32")
Tx.cuda.tile.tcgen05.ld(frag[:, :], C[:, :], scope="warpgroup")