Academic Journal

Design and Optimization of GEMM for Complex Numbers on Ascend NPU.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Design and Optimization of GEMM for Complex Numbers on Ascend NPU.
Συγγραφείς: Zhang, Erkun, Zhang, Yu, Xu, Pengxiang, Lu, Lu
Πηγή: Computers (2073-431X); Jul2026, Vol. 15 Issue 7, p407, 17p
Θεματικοί όροι: Matrix multiplications, Complex variables, Numerical solutions for linear algebra, Signal processing, Neural computers, Scheduling, Mathematical optimization, Optimization algorithms
Περίληψη: It is widely acknowledged that General Matrix Multiplication (GEMM) serves as a foundational kernel across numerous application domains. Complex numbers exhibit distinctive mathematical properties that enable their widespread adoption across engineering computing scenarios, including signal processing and signal transformation. This study investigates high-efficiency CGEMM, namely, complex-valued GEMM, for NPU hardware, broadening the application scope of NPUs beyond mainstream low-precision AI computation workloads. The major contributions of this study are as follows: (i) numerical precision and hardware utilization of the 3M and 4M decomposition schemes on Ascend NPUs are analyzed, and the 4M method is selected as the preferred CGEMM implementation under our tested hardware constraints to fit the bandwidth limitations of modern accelerators for both precision-sensitive and performance-critical matrix computation scenarios; (ii) a complete high-performance CGEMM design based on the 4M scheme tailored for Ascend NPUs is proposed, with an AIC/AIV dual-stream pipeline scheduling strategy equipped to coordinate padding operations, matrix–matrix multiplications, and element-wise instructions across multi-level memory hierarchies and compute units; (iii) a fine-grained task scheduling and assignment mechanism is implemented to maximize Cube core occupancy across diverse matrix dimensions, improving hardware utilization for various computation workloads. Our experimental measurements show that the proposed CGEMM achieves a competitive hardware utilization rate of 83.6% across all tested matrix configurations, enabling efficient exploitation of available computing resources. Meanwhile, we observe a measured average speedup of 1.14× relative to the AscendSipBoost implementation tested on an identical Ascend NPU, alongside a measured 3.17× speedup compared with cuBLAS running on the Nvidia GPU platform adopted in our experiments across all evaluated matrix sizes. These results reflect the promising capability of Ascend NPUs for high-precision complex-valued computing workloads within the tested experimental setup. [ABSTRACT FROM AUTHOR]
Copyright of Computers (2073-431X) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edb&genre=article&issn=2073431X&ISBN=&volume=15&issue=7&date=20260701&spage=407&pages=407-423&title=Computers (2073-431X)&atitle=Design%20and%20Optimization%20of%20GEMM%20for%20Complex%20Numbers%20on%20Ascend%20NPU.&aulast=Zhang%2C%20Erkun&id=DOI:10.3390/computers15070407
    Name: Full Text Finder (for New FTF UI) (ns324271)
    Category: fullText
    Text: Full Text Finder
    MouseOverText: Full Text Finder
Header DbId: edb
DbLabel: Complementary Index
An: 195794706
RelevancyScore: 1082
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 1082.42175292969
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Design and Optimization of GEMM for Complex Numbers on Ascend NPU.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Erkun%22">Zhang, Erkun</searchLink><br /><searchLink fieldCode="AR" term="%22Zhang%2C+Yu%22">Zhang, Yu</searchLink><br /><searchLink fieldCode="AR" term="%22Xu%2C+Pengxiang%22">Xu, Pengxiang</searchLink><br /><searchLink fieldCode="AR" term="%22Lu%2C+Lu%22">Lu, Lu</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: Computers (2073-431X); Jul2026, Vol. 15 Issue 7, p407, 17p
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Matrix+multiplications%22">Matrix multiplications</searchLink><br /><searchLink fieldCode="DE" term="%22Complex+variables%22">Complex variables</searchLink><br /><searchLink fieldCode="DE" term="%22Numerical+solutions+for+linear+algebra%22">Numerical solutions for linear algebra</searchLink><br /><searchLink fieldCode="DE" term="%22Signal+processing%22">Signal processing</searchLink><br /><searchLink fieldCode="DE" term="%22Neural+computers%22">Neural computers</searchLink><br /><searchLink fieldCode="DE" term="%22Scheduling%22">Scheduling</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+optimization%22">Mathematical optimization</searchLink><br /><searchLink fieldCode="DE" term="%22Optimization+algorithms%22">Optimization algorithms</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: It is widely acknowledged that General Matrix Multiplication (GEMM) serves as a foundational kernel across numerous application domains. Complex numbers exhibit distinctive mathematical properties that enable their widespread adoption across engineering computing scenarios, including signal processing and signal transformation. This study investigates high-efficiency CGEMM, namely, complex-valued GEMM, for NPU hardware, broadening the application scope of NPUs beyond mainstream low-precision AI computation workloads. The major contributions of this study are as follows: (i) numerical precision and hardware utilization of the 3M and 4M decomposition schemes on Ascend NPUs are analyzed, and the 4M method is selected as the preferred CGEMM implementation under our tested hardware constraints to fit the bandwidth limitations of modern accelerators for both precision-sensitive and performance-critical matrix computation scenarios; (ii) a complete high-performance CGEMM design based on the 4M scheme tailored for Ascend NPUs is proposed, with an AIC/AIV dual-stream pipeline scheduling strategy equipped to coordinate padding operations, matrix–matrix multiplications, and element-wise instructions across multi-level memory hierarchies and compute units; (iii) a fine-grained task scheduling and assignment mechanism is implemented to maximize Cube core occupancy across diverse matrix dimensions, improving hardware utilization for various computation workloads. Our experimental measurements show that the proposed CGEMM achieves a competitive hardware utilization rate of 83.6% across all tested matrix configurations, enabling efficient exploitation of available computing resources. Meanwhile, we observe a measured average speedup of 1.14× relative to the AscendSipBoost implementation tested on an identical Ascend NPU, alongside a measured 3.17× speedup compared with cuBLAS running on the Nvidia GPU platform adopted in our experiments across all evaluated matrix sizes. These results reflect the promising capability of Ascend NPUs for high-precision complex-valued computing workloads within the tested experimental setup. [ABSTRACT FROM AUTHOR]
– Name: Abstract
  Label:
  Group: Ab
  Data: <i>Copyright of Computers (2073-431X) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=195794706
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3390/computers15070407
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 17
        StartPage: 407
    Subjects:
      – SubjectFull: Matrix multiplications
        Type: general
      – SubjectFull: Complex variables
        Type: general
      – SubjectFull: Numerical solutions for linear algebra
        Type: general
      – SubjectFull: Signal processing
        Type: general
      – SubjectFull: Neural computers
        Type: general
      – SubjectFull: Scheduling
        Type: general
      – SubjectFull: Mathematical optimization
        Type: general
      – SubjectFull: Optimization algorithms
        Type: general
    Titles:
      – TitleFull: Design and Optimization of GEMM for Complex Numbers on Ascend NPU.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zhang, Erkun
      – PersonEntity:
          Name:
            NameFull: Zhang, Yu
      – PersonEntity:
          Name:
            NameFull: Xu, Pengxiang
      – PersonEntity:
          Name:
            NameFull: Lu, Lu
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 2073431X
          Numbering:
            – Type: volume
              Value: 15
            – Type: issue
              Value: 7
          Titles:
            – TitleFull: Computers (2073-431X)
              Type: main
ResultId 1