A High Performance, Portable Distributed BLAS Implementation


In this paper, we give a report on recent developments for the Distributed BLAS (DBLAS) project. These include a powerful distributed matrix representation which yields a simple interface to the DBLAS, and the redesign the DBLAS algorithms terms of powerfuìspread' and`reduce' matrix communication operations for reasons of programmability. The DBLAS codes… (More)