Academic Journal
Design and Optimization of GEMM for Complex Numbers on Ascend NPU.
| Τίτλος: | Design and Optimization of GEMM for Complex Numbers on Ascend NPU. |
|---|---|
| Συγγραφείς: | Zhang, Erkun, Zhang, Yu, Xu, Pengxiang, Lu, Lu |
| Πηγή: | Computers (2073-431X); Jul2026, Vol. 15 Issue 7, p407, 17p |
| Θεματικοί όροι: | Matrix multiplications, Complex variables, Numerical solutions for linear algebra, Signal processing, Neural computers, Scheduling, Mathematical optimization, Optimization algorithms |
| Περίληψη: | It is widely acknowledged that General Matrix Multiplication (GEMM) serves as a foundational kernel across numerous application domains. Complex numbers exhibit distinctive mathematical properties that enable their widespread adoption across engineering computing scenarios, including signal processing and signal transformation. This study investigates high-efficiency CGEMM, namely, complex-valued GEMM, for NPU hardware, broadening the application scope of NPUs beyond mainstream low-precision AI computation workloads. The major contributions of this study are as follows: (i) numerical precision and hardware utilization of the 3M and 4M decomposition schemes on Ascend NPUs are analyzed, and the 4M method is selected as the preferred CGEMM implementation under our tested hardware constraints to fit the bandwidth limitations of modern accelerators for both precision-sensitive and performance-critical matrix computation scenarios; (ii) a complete high-performance CGEMM design based on the 4M scheme tailored for Ascend NPUs is proposed, with an AIC/AIV dual-stream pipeline scheduling strategy equipped to coordinate padding operations, matrix–matrix multiplications, and element-wise instructions across multi-level memory hierarchies and compute units; (iii) a fine-grained task scheduling and assignment mechanism is implemented to maximize Cube core occupancy across diverse matrix dimensions, improving hardware utilization for various computation workloads. Our experimental measurements show that the proposed CGEMM achieves a competitive hardware utilization rate of 83.6% across all tested matrix configurations, enabling efficient exploitation of available computing resources. Meanwhile, we observe a measured average speedup of 1.14× relative to the AscendSipBoost implementation tested on an identical Ascend NPU, alongside a measured 3.17× speedup compared with cuBLAS running on the Nvidia GPU platform adopted in our experiments across all evaluated matrix sizes. These results reflect the promising capability of Ascend NPUs for high-precision complex-valued computing workloads within the tested experimental setup. [ABSTRACT FROM AUTHOR] |
| Copyright of Computers (2073-431X) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Βάση Δεδομένων: | Complementary Index |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edb&genre=article&issn=2073431X&ISBN=&volume=15&issue=7&date=20260701&spage=407&pages=407-423&title=Computers (2073-431X)&atitle=Design%20and%20Optimization%20of%20GEMM%20for%20Complex%20Numbers%20on%20Ascend%20NPU.&aulast=Zhang%2C%20Erkun&id=DOI:10.3390/computers15070407 Name: Full Text Finder (for New FTF UI) (ns324271) Category: fullText Text: Full Text Finder MouseOverText: Full Text Finder |
|---|---|
| Header | DbId: edb DbLabel: Complementary Index An: 195794706 RelevancyScore: 1082 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 1082.42175292969 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Design and Optimization of GEMM for Complex Numbers on Ascend NPU. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Zhang%2C+Erkun%22">Zhang, Erkun</searchLink><br /><searchLink fieldCode="AR" term="%22Zhang%2C+Yu%22">Zhang, Yu</searchLink><br /><searchLink fieldCode="AR" term="%22Xu%2C+Pengxiang%22">Xu, Pengxiang</searchLink><br /><searchLink fieldCode="AR" term="%22Lu%2C+Lu%22">Lu, Lu</searchLink> – Name: TitleSource Label: Source Group: Src Data: Computers (2073-431X); Jul2026, Vol. 15 Issue 7, p407, 17p – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Matrix+multiplications%22">Matrix multiplications</searchLink><br /><searchLink fieldCode="DE" term="%22Complex+variables%22">Complex variables</searchLink><br /><searchLink fieldCode="DE" term="%22Numerical+solutions+for+linear+algebra%22">Numerical solutions for linear algebra</searchLink><br /><searchLink fieldCode="DE" term="%22Signal+processing%22">Signal processing</searchLink><br /><searchLink fieldCode="DE" term="%22Neural+computers%22">Neural computers</searchLink><br /><searchLink fieldCode="DE" term="%22Scheduling%22">Scheduling</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematical+optimization%22">Mathematical optimization</searchLink><br /><searchLink fieldCode="DE" term="%22Optimization+algorithms%22">Optimization algorithms</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: It is widely acknowledged that General Matrix Multiplication (GEMM) serves as a foundational kernel across numerous application domains. Complex numbers exhibit distinctive mathematical properties that enable their widespread adoption across engineering computing scenarios, including signal processing and signal transformation. This study investigates high-efficiency CGEMM, namely, complex-valued GEMM, for NPU hardware, broadening the application scope of NPUs beyond mainstream low-precision AI computation workloads. The major contributions of this study are as follows: (i) numerical precision and hardware utilization of the 3M and 4M decomposition schemes on Ascend NPUs are analyzed, and the 4M method is selected as the preferred CGEMM implementation under our tested hardware constraints to fit the bandwidth limitations of modern accelerators for both precision-sensitive and performance-critical matrix computation scenarios; (ii) a complete high-performance CGEMM design based on the 4M scheme tailored for Ascend NPUs is proposed, with an AIC/AIV dual-stream pipeline scheduling strategy equipped to coordinate padding operations, matrix–matrix multiplications, and element-wise instructions across multi-level memory hierarchies and compute units; (iii) a fine-grained task scheduling and assignment mechanism is implemented to maximize Cube core occupancy across diverse matrix dimensions, improving hardware utilization for various computation workloads. Our experimental measurements show that the proposed CGEMM achieves a competitive hardware utilization rate of 83.6% across all tested matrix configurations, enabling efficient exploitation of available computing resources. Meanwhile, we observe a measured average speedup of 1.14× relative to the AscendSipBoost implementation tested on an identical Ascend NPU, alongside a measured 3.17× speedup compared with cuBLAS running on the Nvidia GPU platform adopted in our experiments across all evaluated matrix sizes. These results reflect the promising capability of Ascend NPUs for high-precision complex-valued computing workloads within the tested experimental setup. [ABSTRACT FROM AUTHOR] – Name: Abstract Label: Group: Ab Data: <i>Copyright of Computers (2073-431X) is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edb&AN=195794706 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.3390/computers15070407 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 17 StartPage: 407 Subjects: – SubjectFull: Matrix multiplications Type: general – SubjectFull: Complex variables Type: general – SubjectFull: Numerical solutions for linear algebra Type: general – SubjectFull: Signal processing Type: general – SubjectFull: Neural computers Type: general – SubjectFull: Scheduling Type: general – SubjectFull: Mathematical optimization Type: general – SubjectFull: Optimization algorithms Type: general Titles: – TitleFull: Design and Optimization of GEMM for Complex Numbers on Ascend NPU. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Zhang, Erkun – PersonEntity: Name: NameFull: Zhang, Yu – PersonEntity: Name: NameFull: Xu, Pengxiang – PersonEntity: Name: NameFull: Lu, Lu IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 2073431X Numbering: – Type: volume Value: 15 – Type: issue Value: 7 Titles: – TitleFull: Computers (2073-431X) Type: main |
| ResultId | 1 |