Productive, portable, and performant GPU programming in Python.
-
Updated
Jul 6, 2026 - C++
Productive, portable, and performant GPU programming in Python.
High-performance large-scale embedding acceleration for JAX on Google TPU SparseCores.
Compressed Sparse Row (CSR) Sparse Matrix-Dense Matrix Multiplication (SpMM) acceleration kernel with row-pointer parallelism and load-balanced binning.
Compressed Sparse Row (CSR) Sparse Matrix-Dense Matrix Multiplication (SpMM) acceleration kernel with row-pointer parallelism and load-balanced binning.
Leveraging Taichi Lang to customize brain dynamics operators.
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation [SC'25]
Sparse Lie algebra engine for G₂, F₄, E₆, E₇, E₈ — 913× compression, lattice gauge theory, equivariant GNN layers. pip install dhl-mm
Relation-based language modeling, RelationLex tokenization, stateful decode, and fused Triton kernels
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
Research code for testing whether causal token surprisal can guide adaptive computation, sparse refinement, and learned compute allocation in byte-level language models.
Design record for ECSIE, an entropy-controlled execution runtime for sparse Mixture-of-Experts models — architecture, public API, benchmark harness and measured results.
A biologically inspired R&D blueprint for sparse, grounded, continual, energy-efficient AI.
To associate your repository with the sparse-computation topic, visit your repo's landing page and select "manage topics."