Productive, portable, and performant GPU programming in Python.
-
Updated
Jul 6, 2026 - C++
Productive, portable, and performant GPU programming in Python.
High-performance large-scale embedding acceleration for JAX on Google TPU SparseCores.
Leveraging Taichi Lang to customize brain dynamics operators.
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation [SC'25]
Sparse Lie algebra engine for G₂, F₄, E₆, E₇, E₈ — 913× compression, lattice gauge theory, equivariant GNN layers. pip install dhl-mm
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
Relation-based language modeling, RelationLex tokenization, stateful decode, and fused Triton kernels
Design record for ECSIE, an entropy-controlled execution runtime for sparse Mixture-of-Experts models — architecture, public API, benchmark harness and measured results.
Research code for testing whether causal token surprisal can guide adaptive computation, sparse refinement, and learned compute allocation in byte-level language models.
Add a description, image, and links to the sparse-computation topic page so that developers can more easily learn about it.
To associate your repository with the sparse-computation topic, visit your repo's landing page and select "manage topics."