tilelang is a Pythonic DSL plus compiler infrastructure built on Apache TVM for writing high-performance GPU/CPU/NPU kernels such as GEMM and FlashAttention productively. It compiles to CUDA, HIP, Metal, and Ascend targets and works as a Python library with a one-line pip install.
Project overview
Kernel developers usually trade Python-level productivity for hand-tuned low-level performance; tilelang addresses that trade-off by pairing Pythonic syntax with a TVM-based compiler stack while still targeting multiple backends including CUDA, HIP, Metal, and Ascend.
Project type
Model Development · Infrastructure
Deployment
Refer to project documentation
License
License pending
Best for
Developers who write custom GPU/CPU/NPU kernels and want Pythonic productivity without losing state-of-the-art low-level optimizations.
Key capabilities
Compilation targets include CUDA, HIP, Metal, and Ascend, with an auto target that detects these devices from the environment.
Limitations and risks
The Ascend 950 backend requires building from source with USE_ASCEND=ON, along with CANN and torch_npu, rather than the pip-installed package.
Nightly builds may be less stable than official releases; check this before adopting a nightly version.
Getting started
Install with a single pip install: pip install tilelang. Verify with python -c "import tilelang; print(tilelang.__version__)". Then run the Quick Start GEMM+ReLU example, which uses PyTorch CUDA tensors as kernel inputs.
Evidence and sources
GitHub project description: Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
README: Tile Language (**tile-lang**) is a concise domain-specific language designed to streamline the development of high-performance GPU/CPU/NPU kernels (e.g., GEMM, Dequant GEMM, Flash…
README: The default `auto` target detects CUDA, HIP, Metal, and Ascend devices; select an explicit target when compiling for another backend or architecture.
README: The following example defines, compiles, runs, and verifies an FP16 GEMM kernel with FP32 accumulation and a fused ReLU epilogue. It uses PyTorch CUDA tensors
README: **2025-02-12 — [TileLang v0.1.0](https://github.com/tile-ai/tilelang/releases/tag/v0.1.0):** published the first v0.1 public release.