A GEMM interface and implementation on NVIDIA GPUs for multiple small matrices

@article{Jhurani2015AGI,
  title={A GEMM interface and implementation on NVIDIA GPUs for multiple small matrices},
  author={Chetan Jhurani and Paul Mullowney},
  journal={J. Parallel Distrib. Comput.},
  year={2015},
  volume={75},
  pages={133-140}
}
We present an interface and an implementation of the General Matrix Multiply (GEMM) routine for multiple small matrices processed simultaneously on NVIDIA graphics processing units (GPUs). We focus on matrix sizes under 16. The implementation can be easily extended to larger sizes. For single precision matrices, our implementation is 30% to 600% faster than the batched cuBLAS implementation distributed in the CUDA Toolkit 5.0 on NVIDIA Tesla K20c. For example, we obtain 104 GFlop/s and 216… CONTINUE READING

From This Paper

Figures, tables, results, and topics from this paper.

Key Quantitative Results

  • For single precision matrices, our implementation is 30% to 600% faster than the batched cuBLAS implementation distributed in the CUDA Toolkit 5.0 on NVIDIA Tesla K20c.

Citations

Publications citing this paper.
Showing 1-10 of 11 extracted citations

References

Publications referenced by this paper.
Showing 1-10 of 13 references

R

  • P. Bientinesi, V. Eijkhout, K. Kim, J. Kurtz
  • van de Geijn, Sparse direct factorizations…
  • 2010

Sparse direct factorizations through unassembled hypermatrices

  • V. Eijkhout, K. Kim, J. Kurtz
  • Computer Methods in Applied Mechanics and…
  • 2010

Computing with hp-ADAPTIVE FINITE ELEMENTS: Vol- 19 ume II Frontiers: Three Dimensional Elliptic and Maxwell Problems with Applications

  • L. Demkowicz, W. Rachowicz, D. Pardo, M. Paszynski, J. Kurtz, A. Zdunek
  • Chapman & Hall/CRC Press
  • 2007
2 Excerpts

Computing with hp-ADAPTIVE FINITE ELEMENTS: Volume I: One and Two Dimensional Elliptic and Maxwell Problems

  • L. Demkowicz
  • Chapman & Hall/CRC Press
  • 2006
2 Excerpts

LAPACK Users’ Guide

  • E. Anderson, Z. Bai, +8 authors D. Sorensen
  • 3rd Edition, SIAM, Philadelphia, PA
  • 1999
1 Excerpt

Spectral/hp element methods for CFD

  • G. Karniadakis, S. Sherwin
  • Oxford University Press, USA
  • 1999
2 Excerpts

Similar Papers

Loading similar papers…