Adds support for NVIDIA Ampere Architecture features. CUDA 11 Toolkit recommended.
CUTLASS 2.1 contributes: - BLAS-style host-side API added to CUTLASS Library - Planar Complex GEMM kernels targeting Volta and Turing Tensor Cores - Minor enhancements and bug fixes