Tensor decomposition and applicationsParallel Computing and Optimization TechniquesPolynomial and algebraic computation

M. Kubale, Damian Niemczyk

2026.3.30Archives of Control Sciences

DOI: 10.24425/acs.2026.158422

Abstract

The aim of this paper is to analyze the development of algorithms for Fast Matrix Multiplication (FMM) in both historical and technical contexts, as well as to compare available solutions on consumer-grade computer hardware. We review advancements in estimating the theoretical computational complexity of FMM and optimization techniques that are used in widely adopted algorithms, with a particular focus on optimal cache memory usage and leveraging Graphics Processing Units (GPU). The methodology of tests and their analysis highlight the performance differences of the considered algorithms depending on the matrix size and the nature of the data stored in them. Results indicate the significant role of tailoring the chosen algorithm to the available hardware and the specific application in which the algorithm is being performed. Also, we emphasize that the FMM algorithms can be applied not only to linear algebra problems but also to current problems in science and engineering, such as artificial intelligence, databases, parallel computations, computational biology, pattern recognition, and compiler construction, to mention just a few examples.

Citation format

KUBALE, M.; NIEMCZYK, Damian. Practical aspects of fast matrix multiplication. Archives of Control Sciences, 2026: 85–98.