Tensor Instruction Registration

Tensor Instruction Registration#

Backend authors define an Instruction with its canonical operator name, ordered Operand list, reflected static attribute schema, and one lowerer. The shared implementation is tvm.tirx.tile.instruction; CUDA and Trainium’s tile/instructions.py files contain the concrete contracts.

Registration installs the fixed argument count, void return type, opaque side effects, category, printer name, and native Call validator on the Op. The validator also runs for directly constructed Calls and printer entry. Runtime expressions must be operands. New static qualifiers require a typed attribute field rather than an arbitrary dictionary.

Lowerers receive a semantic view and DispatchContext and return a tvm.tirx.Function body. They can prepare pointers/descriptors and expand layouts for the selected core instruction. They must report unsupported contracts rather than select a different instruction family.

Named workspace operands have fixed optional slots. Declare a workspace policy only for instructions that own such scratch storage; do not use it to hide caller-owned algorithms. See Tensor Instruction Lowering for callbacks and the statement replacement boundary.

The common Python API is under tvm.tirx.tile (also T.tile): Instruction, Operand, TensorCall, and DispatchContext. Backend implementations live under tvm.backend.cuda.tile and tvm.backend.trn.tile. Their user-facing instruction paths remain T.cuda.tile and T.trn.tile. Operator names and serialized type keys remain unchanged.

Execution scopes are generic backend/name identities. Backend load hooks register their supported scopes and ordering; CUDA exposes T.cuda.ExecScope("warp") and Trainium exposes T.trn.ExecScope("thread"). A string instruction scope is resolved by the instruction’s backend. CUDA predicate analysis, index bindings, and active thread sets live in its private tile implementation; Trainium does not use CUDA’s execution-context model.