{"id":48083,"date":"2023-10-03T10:05:00","date_gmt":"2023-10-03T17:05:00","guid":{"rendered":"https:\/\/devblogs.microsoft.com\/dotnet\/?p=48083"},"modified":"2023-10-03T10:05:00","modified_gmt":"2023-10-03T17:05:00","slug":"this-arm64-performance-in-dotnet-8","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/dotnet\/this-arm64-performance-in-dotnet-8\/","title":{"rendered":"Arm64 Performance Improvements in .NET 8"},"content":{"rendered":"<p>.NET 8 comes with a host of powerful features, enhanced support for cutting-edge architectural capabilities, and a significant boost in performance. For a comprehensive look at the general performance enhancements in .NET 8, be sure to delve into <a href=\"https:\/\/devblogs.microsoft.com\/dotnet\/performance-improvements-in-net-8\/\">Stephen Toub&#8217;s insightful blog post<\/a> on the subject.<\/p>\n<p>Building on the foundation laid by previous releases such as <a href=\"https:\/\/devblogs.microsoft.com\/dotnet\/arm64-performance-in-net-5\/\">ARM64 Performance in .NET 5<\/a> and <a href=\"https:\/\/devblogs.microsoft.com\/dotnet\/arm64-performance-improvements-in-dotnet-7\/\">ARM64 Performance in .NET 7<\/a>, this article will focus on the specific feature enhancements and performance optimizations tailored to ARM64 architecture in .NET 8.<\/p>\n<p>A key objective for .NET 8 was to enhance the performance of the platform on Arm64 systems. However, we also set our sights on incorporating support for advanced features offered by the Arm architecture, thereby elevating the overall code quality for Arm-based platforms. In this blog post, we&#8217;ll examine some of these noteworthy feature additions before diving into the extensive performance improvements that have been implemented. Lastly, we&#8217;ll provide insights into the results of our performance analysis on real-world applications designed for Arm64 devices.<\/p>\n<h2>Conditional Selection<\/h2>\n<p>Branches within the code can significantly impact the efficiency of the processor pipeline, potentially causing execution delays. While modern processors have made strides in branch prediction to maximize accuracy, they still mispredict branches, resulting in wasted processing cycles. Conditional selection instructions offer an alternative by encoding the necessary conditional flags directly within the instruction, eliminating the need for branch instructions. <a href=\"https:\/\/eclecticlight.co\/2021\/07\/20\/code-in-arm-assembly-conditions-without-branches\/\">This article<\/a> offers a comprehensive overview of the various formats of these instructions.<\/p>\n<p>In .NET 8, we&#8217;ve taken on the challenge of addressing all the scenarios outlined in <a href=\"https:\/\/github.com\/dotnet\/runtime\/issues\/55364\">dotnet\/runtime#55364<\/a> to implement conditional selection instructions effectively. In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/73472\">dotnet\/runtime#73472<\/a>, <a href=\"https:\/\/github.com\/a74nh\">@a74nh<\/a> introduced a new &#8220;if-conversion&#8221; phase designed to generate conditional selection instructions. Let&#8217;s dig into an illustrative example.<\/p>\n<pre><code class=\"language-csharp\">if (op1 &gt; 0) { op1 = 5; }<\/code><\/pre>\n<p>In .NET 7, we generated <code>branch<\/code> for such code:<\/p>\n<pre><code class=\"language-asm\">G_M12157_IG02:\n            cmp     w0, #0\n            ble     G_M12157_IG04\n\nG_M12157_IG03:\n            mov     w0, #5\n\nG_M12157_IG04:\n            ...<\/code><\/pre>\n<p>In .NET 8, we introduced the generation of <code>csel<\/code> instructions, illustrated below. The <code>cmp<\/code> instruction sets the flags based on the value in <code>w0<\/code>, determining whether it&#8217;s greater than, equal to, or less than zero. The <code>csel<\/code> instruction then evaluates the &#8220;le&#8221; condition (the 4th operand of <code>csel<\/code>) triggered by the <code>cmp<\/code>. Essentially, it checks if <code>w0 &lt;= 0<\/code>. If this condition holds true, it preserves the current value of <code>w0<\/code> (the 2nd operand of <code>csel<\/code>) as the result (the 1st operand of <code>csel<\/code>). However, if <code>w0 &gt; 0<\/code>, it updates <code>w0<\/code> with the value in <code>w1<\/code> (the 3rd operand of <code>csel<\/code>).<\/p>\n<pre><code class=\"language-asm\">G_M12157_IG02:\n            mov     w1, #5\n            cmp     w0, #0\n            csel    w0, w0, w1, le<\/code><\/pre>\n<p><a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/77728\">dotnet\/runtime#77728<\/a> extended this work to handle <code>else<\/code> conditions as seen in following example:<\/p>\n<pre><code class=\"language-csharp\"> if (op1 &lt; 7) { op2 = 5; } else { op2 = 9; }<\/code><\/pre>\n<p>Previously, we generated two branch instructions for <code>if-else<\/code> style code:<\/p>\n<pre><code class=\"language-asm\"> G_M52833_IG02:\n            cmp     w0, #7\n            bge     G_M52833_IG04\n\nG_M52833_IG03:\n            mov     w1, #5\n            b       G_M52833_IG05\n\nG_M52833_IG04:\n            mov     w1, #9\n\nG_M52833_IG05:\n            ...<\/code><\/pre>\n<p>With the introduction of <code>csel<\/code>, there is no longer a need for branches in this example. Instead, we can set the appropriate result value based on the conditional flag encoded within the instruction itself.<\/p>\n<pre><code class=\"language-asm\"> G_M52833_IG02:\n            mov     w2, #9\n            mov     w3, #5\n            cmp     w0, #7\n            csel    w0, w2, w3, ge<\/code><\/pre>\n<p>As indicated in <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/9677\">this<\/a>, <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/9496\">this<\/a> and <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/9496\">this<\/a> reports, we have observed nearly a 50% improvement in such scenarios.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/pr_csel_1.png\" alt=\"pr_csel_1.png\" \/><\/p>\n<h3>Conditional Comparison<\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/83089\">dotnet\/runtime#83089<\/a>, we enhanced the code even further by consolidating multiple conditions separated by the <code>||<\/code> operation into a conditional comparison <code>ccmp<\/code> instruction.<\/p>\n<pre><code class=\"language-csharp\">[MethodImpl(MethodImplOptions.NoInlining)]\npublic static int Test3(int a, int b, int c)\n{\n    if (a &lt; b || b &lt; c || c &lt; 10)\n        return 42;\n\n    return 13;\n}<\/code><\/pre>\n<p>In the code snippet above, we observe several conditions linked together using the logical OR <code>||<\/code> operator. As a quick reminder, when using <code>||<\/code>, we evaluate the first condition, and if it&#8217;s true, we don&#8217;t check the remaining conditions. We only evaluate the subsequent conditions if the preceding ones are <code>false<\/code>.<\/p>\n<p>Prior to .NET 8, our approach involved generating numerous instructions to conduct separate comparisons using <code>cmp<\/code> and storing the result of each comparison in a register. Even if the first condition evaluated to <code>true<\/code>, we still processed all the conditions joined by the <code>||<\/code> operator. The outcomes of these individual comparisons were then combined using the <code>orr<\/code> instruction, and a branch instruction <code>cbz<\/code> was employed to determine whether to jump based on the final result of all the conditions collectively.<\/p>\n<pre><code class=\"language-asm\">G_M35702_IG02:\n            cmp     w0, w1\n            cset    x0, lt\n            cmp     w1, w2\n            cset    x1, lt\n            orr     w0, w0, w1\n            cmp     w2, #10\n            cset    x1, lt\n            orr     w0, w0, w1\n            cbz     w0, G_M35702_IG05\n\nG_M35702_IG03:\n            mov     w0, #42\n\nG_M35702_IG04:\n            ldp     fp, lr, [sp],#0x10\n            ret     lr\n\nG_M35702_IG05:\n            mov     w0, #13\n\nG_M35702_IG06:\n            ...<\/code><\/pre>\n<p>In .NET 8, we&#8217;ve introduced a more efficient approach. We initiate the process by comparing the values of <code>w0<\/code> and <code>w1<\/code> (as indicated by the <code>a &lt; b<\/code> condition in our code snippet) using the <code>cmp<\/code> instruction. Subsequently, the <code>ccmp w1, w2, nc, ge<\/code> instruction comes into play. This instruction will only compare the values of <code>w1<\/code> and <code>w2<\/code> if the previous <code>cmp<\/code> instruction determined that the condition <code>ge<\/code> (greater than or equal) was met. In simpler terms, <code>w1<\/code> and <code>w2<\/code> are compared only when the <code>cmp<\/code> instruction finds that <code>w0<\/code> is greater than or equal to <code>w1<\/code>, meaning our condition <code>a &lt; b<\/code> is false. If the first condition evaluates to true, there&#8217;s no need for further condition checks, so the <code>ccmp<\/code> instruction simply sets the <code>nc<\/code> flags. This process is repeated for the subsequent <code>ccmp<\/code> instruction. Finally, the <code>csel<\/code> instruction determines the value (<code>w3<\/code> or <code>w4<\/code>) to place in the result (<code>w0<\/code>) based on the outcome of the <code>ge<\/code> condition. This example illustrates how processor cycles are conserved through the use of the <code>ccmp<\/code> instruction.<\/p>\n<pre><code class=\"language-asm\">G_M35702_IG02:\n            mov     w3, #13\n            mov     w4, #42\n            cmp     w0, w1\n            ccmp    w1, w2, nc, ge\n            ccmp    w2, #10, nc, ge\n            csel    w0, w3, w4, ge<\/code><\/pre>\n<h3>Conditional Increment, Negation and Inversion<\/h3>\n<p>The <code>cinc<\/code> instruction is part of the conditional selection family of instructions, and it serves to increment the value of the source register by <code>1<\/code> if a specific condition is satisfied. <\/p>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/82031\">dotnet\/runtime#82031<\/a>, the code sequence was optimized by <a href=\"https:\/\/github.com\/SwapnilGaikwad\">@SwapnilGaikwad<\/a> to generate <code>cinc<\/code> instructions instead of <code>csel<\/code> instructions whenever the code conditionally needed to increase a value by 1. This optimization helps streamline the code and improves its efficiency by directly incrementing the value when necessary, rather than relying on conditional selection for this specific operation.<\/p>\n<pre><code class=\"language-csharp\">public static int Test4(bool c)\n{\n    return c ? 100 : 101;\n}<\/code><\/pre>\n<p>Before .NET 8, two registers were required to store the values of both branches. After the <code>cmp<\/code> instruction, a selection would be made to determine which of the two values (<code>w1<\/code> and <code>w2<\/code>) should be stored in the result register <code>w0<\/code>. This approach used additional registers and instructions, which could impact performance and code size.<\/p>\n<p>Here&#8217;s how the code looked in older releases:<\/p>\n<pre><code class=\"language-asm\">            mov     w1, #100\n            mov     w2, #101\n            cmp     w0, #0\n            csel    w0, w1, w2, ne<\/code><\/pre>\n<p>With this structure, <code>w0<\/code> would either contain the value of <code>w1<\/code> or <code>w2<\/code> based on the condition, leading to potential register usage inefficiencies.<\/p>\n<p>In .NET 8, we have eliminated the need to maintain an additional register. Instead, we use the <code>cinc<\/code> instruction to increment the value if the <code>tst<\/code> instruction succeeds. This enhancement simplifies the code and reduces the reliance on extra registers, potentially leading to improved performance and more efficient code execution.<\/p>\n<pre><code class=\"language-asm\">            mov     w1, #100\n            tst     w0, #255\n            cinc    w0, w1, eq<\/code><\/pre>\n<p>Similarly, in <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/84926\">dotnet\/runtime#84926<\/a>, we introduced support for <code>cinv<\/code> and <code>cneg<\/code> instructions, replacing the use of <code>csel<\/code>. This is demonstrated in the following example:<\/p>\n<pre><code class=\"language-csharp\">public static int cinv_test(int x, bool c)\n{\n    return c ? x : ~x;\n}\n\npublic static int cneg_test(int x, bool c)\n{\n    return c ? x : -x;\n}<\/code><\/pre>\n<pre><code class=\"language-asm\">            tst     w1, #255\n            csinv   w0, w0, w0, ne\n            ...\n            tst     w1, #255\n            csneg   w0, w0, w0, ne<\/code><\/pre>\n<p><a href=\"https:\/\/github.com\/a74nh\">@a74nh<\/a> has published a three part deep dive blog series explaining this work that was done in .NET. Read <a href=\"https:\/\/community.arm.com\/arm-community-blogs\/b\/architectures-and-processors-blog\/posts\/if-conversion-within-dotnet-part-1\">part 1<\/a>, <a href=\"https:\/\/community.arm.com\/arm-community-blogs\/b\/architectures-and-processors-blog\/posts\/if-conversion-within-dotnet-part-2\">part 2<\/a> and <a href=\"https:\/\/community.arm.com\/arm-community-blogs\/b\/architectures-and-processors-blog\/posts\/if-conversion-within-dotnet-part-3\">part 3<\/a> for more details.<\/p>\n<h2>VectorTableLookup and VectorTableLookupExtension<\/h2>\n<p>In .NET 8, we added two new set of APIs under <code>System.Runtime.Intrinsics.Arm<\/code> namespace: <code>VectorTableLookup<\/code> and <code>VectorTableLookupExtension<\/code>. <\/p>\n<pre><code class=\"language-csharp\">      public static Vector64&lt;byte&gt; VectorTableLookup((Vector128&lt;byte&gt;, Vector128&lt;byte&gt;) table, Vector64&lt;byte&gt; byteIndexes);\n      public static Vector64&lt;byte&gt; VectorTableLookup(Vector64&lt;byte&gt; defaultValues, (Vector128&lt;byte&gt;, Vector128&lt;byte&gt;) table, Vector64&lt;byte&gt; byteIndexes);\n<\/code><\/pre>\n<p>Let us see an example of each these APIs.<\/p>\n<pre><code class=\"language-c#\">\/\/ Vector128&lt;byte&gt; a = 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16\n\/\/ Vector128&lt;byte&gt; b = 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160\n\/\/ Vector64&lt;byte&gt; index = 3, 31, 4, 40, 18, 19, 30, 1\n\nVector64&lt;byte&gt; ans = VectorTableLookup((a, b), index);\n\n\/\/ ans = 4, 160, 5, 0, 30, 40, 150, 2<\/code><\/pre>\n<p>In the example above, the vectors <code>a<\/code> and <code>b<\/code> are treated as a single table with a total of 32 entries (16 from <code>a<\/code> and 16 from <code>b<\/code>), indexed starting from 0. The <code>index<\/code> parameter allows you to retrieve values from this table at specific indices. If an index is out of bounds, such as attempting to access index <code>40<\/code> in our example, the API will return a value of 0 for that out-of-bounds index.<\/p>\n<pre><code class=\"language-c#\">\/\/ Vector64&lt;byte&gt; d = 100, 200, 300, 400, 500, 600, 700, 800\n\/\/ Vector128&lt;byte&gt; a = 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16\n\/\/ Vector128&lt;byte&gt; b = 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160\n\/\/ Vector64&lt;byte&gt; index = 3, 31, 4, 40, 18, 19, 30, 1\n\nVector64&lt;byte&gt; ans = VectorTableLookupExtension(d, (a, b), index);\n\n\/\/ ans = 4, 160, 5, 400, 30, 40, 150, 2<\/code><\/pre>\n<p>In contrast to the <code>VectorTableLookup<\/code>, when using <code>VectorTableLookupExtension<\/code> method, if an index falls outside the valid range, the corresponding element in the result will be determined by the values provided in the <code>defaultValues<\/code> parameter. It&#8217;s worth noting that there are other variations of these APIs that operate on 3-entity and 4-entity tuples as well, providing flexibility for various use cases.<\/p>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/85189\">dotnet\/runtime#85189<\/a>, <a href=\"https:\/\/github.com\/MihaZupan\">@MihaZupan<\/a> leveraged this API to optimize <code>IndexOfAny<\/code>, resulting in a remarkable 30% improvement in performance. Similarly, in <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/87126\">dotnet\/runtime#87126<\/a>, <a href=\"https:\/\/github.com\/SwapnilGaikwad\">@SwapnilGaikwad<\/a> significantly enhanced the performance of the Guid formatter, achieving up to a 40% performance boost. These optimizations demonstrate the substantial performance gains that can be achieved by harnessing this powerful API.<\/p>\n<h2>Consecutive Register Allocation<\/h2>\n<p>The <code>VectorTableLookup<\/code> and <code>VectorTableLookupExtension<\/code> APIs utilize the <code>tbl<\/code> and <code>tbx<\/code> instructions, respectively, for table lookup operations. These instructions work with input vectors contained in the <code>table<\/code> parameter, which can consist of 1, 2, 3, or 4 vector entities. It&#8217;s important to note that these instructions require the input vectors to be stored in consecutive vector registers on the processor. For example, consider the code <code>VectorTableLookup((a, b, c, d), index)<\/code>, which generates the <code>tbl v11.16b, {v16.16b, v17.16b, v18.16b, v19.16b}, v20.16b<\/code> instruction. In this instruction, <code>v11<\/code> serves as the result vector, <code>v20<\/code> holds the value of the <code>index<\/code> variable, and the <code>table<\/code> parameter, represented as the tuple <code>(a, b, c, d)<\/code>, is stored in consecutive registers <code>v16<\/code> through <code>v19<\/code>.<\/p>\n<p>Before .NET 8, RyuJIT&#8217;s register allocator could only allocate a single register for each variable. However, in <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/80297\">dotnet\/runtime#80297<\/a>, we introduced a new feature in the register allocator to allocate multiple consecutive registers for these types of instructions, enabling more efficient code generation.<\/p>\n<p>To implement this feature, we made adjustments to the existing allocator algorithm and certain data structures to effectively manage variables that belong to a tuple series requiring consecutive registers. Here&#8217;s an overview of how it works:<\/p>\n<ol>\n<li>\n<p>Tracking Consecutive Register Requirement: The algorithm now identifies the first variable in a tuple series that requires consecutive registers.<\/p>\n<\/li>\n<li>\n<p>Register Allocation: Before assigning registers to variables, the algorithm checks for the availability of X consecutive free registers, where X is the number of variables in the <code>table<\/code> tuple.<\/p>\n<\/li>\n<li>\n<p>Handling Register Shortages: If there aren&#8217;t X consecutive free registers available, the algorithm looks for busy registers that are adjacent to some of the free registers. It then frees up these busy registers to meet the requirement of assigning X consecutive registers to the X variables in the <code>table<\/code> tuple.<\/p>\n<\/li>\n<li>\n<p>Versatility: The allocation of consecutive registers isn&#8217;t limited to just <code>tbl<\/code> and <code>tbx<\/code> instructions. It&#8217;s also used for various load and store instructions that will be implemented in .NET 9. For instance, <code>ld2<\/code>, <code>ld3<\/code>, and <code>ld4<\/code> are load instructions capable of loading 2, 3, or 4 vector registers from memory, respectively. These instructions, as well as store instructions like <code>st2<\/code>, <code>st3<\/code>, and <code>st4<\/code>, all require consecutive registers.<\/p>\n<\/li>\n<\/ol>\n<p>This enhancement in the register allocator ensures that variables involved in tuple-based operations are allocated consecutive registers when necessary, improving code generation efficiency not only for existing instructions but also for upcoming load and store instructions in future .NET versions.<\/p>\n<h2>Peephole optimizations<\/h2>\n<p>In .NET 5, we encountered several issues, as documented in <a href=\"https:\/\/github.com\/dotnet\/runtime\/issues\/55365\">dotnet\/runtime#55365<\/a>, where the application of peephole optimizations could significantly enhance the generated .NET code. Collaborating closely with engineers from Arm Corp., including <a href=\"https:\/\/github.com\/a74nh\">@a74nh<\/a>, <a href=\"https:\/\/github.com\/SwapnilGaikwad\">@SwapnilGaikwad<\/a>, and <a href=\"https:\/\/github.com\/AndyJGraham\">@AndyJGraham<\/a>, we successfully addressed all these issues in .NET 8. In the sections below, I will provide an overview of each of these improvements, along with illustrative examples.<\/p>\n<h3>Replace successive <code>ldr<\/code> and <code>str<\/code> with <code>ldp<\/code> and <code>stp<\/code><\/h3>\n<p>Through the combined efforts of <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/77540\">dotnet\/runtime#77540<\/a>, <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/85032\">dotnet\/runtime#85032<\/a>, and <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/84399\">dotnet\/runtime#84399<\/a>, we&#8217;ve introduced an optimization that replaces successive pairs of load (<code>ldr<\/code>) and store (<code>str<\/code>) instructions with load pair (<code>ldp<\/code>) and store pair (<code>stp<\/code>) instructions, respectively. This enhancement has led to a remarkable reduction in the code size of several of our libraries and benchmark methods, amounting to a decrease of nearly 700KB in bytes.<\/p>\n<p>As a result of this optimization, the following sequence of two <code>ldr<\/code> instructions and two <code>str<\/code> instructions has been streamlined to generate a single <code>ldp<\/code> instruction and a single <code>stp<\/code> instruction:<\/p>\n<pre><code class=\"language-diff\">- ldr x0, [fp, #0x10]\n- ldr x1, [fp, #0x16]\n+ ldp x0, x1, [fp, #0x10]\n...\n- str x12, [x14]\n- str x12, [x14, #8]\n+ stp x12, [x14]<\/code><\/pre>\n<p>This transformation contributes to improved code efficiency and a reduction in overall binary size.<\/p>\n<h3>Use ldp\/stp for SIMD registers<\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/84135\">dotnet\/runtime#84135<\/a>, we introduced an optimization where a pair of load and store instructions involving SIMD (Single Instruction, Multiple Data) registers were replaced with a single <code>ldp<\/code> (Load Pair) and <code>stp<\/code> (Store Pair) instruction. This optimization is demonstrated in the following example:<\/p>\n<pre><code class=\"language-diff\">- ldr q1, [x0, #0x20]\n- ldr q2, [x0, #0x30]\n+ ldp q1, q2, [x0, #0x20]<\/code><\/pre>\n<h3>Replace pair of <code>str wzr<\/code> with <code>str xzr<\/code><\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/84350\">dotnet\/runtime#84350<\/a>, we introduced an optimization that consolidated a pair of stores involving 4-byte zero registers into a single store instruction that utilized an 8-byte zero register. This optimization enhances code efficiency by reducing the number of instructions required, resulting in improved performance and potentially smaller code size.<\/p>\n<pre><code class=\"language-diff\">- stp wzr, wzr, [x2, #0x08]\n+ str xzr, [x2, #0x08]\n...\n...\n- stp wzr, wzr, [x14, #0x20]\n- str wzr, [x14, #0x18]\n+ stp xzr, xzr, [x14, #0x18]<\/code><\/pre>\n<h3>Replace load with cheaper <code>mov<\/code><\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/83458\">dotnet\/runtime#83458<\/a>, we implemented an optimization that replaced a heavier load instruction with a more efficient <code>mov<\/code> instruction in certain scenarios. For example, when the content of memory loaded into register <code>w1<\/code> was subsequently loaded into <code>w0<\/code>, we used the <code>mov<\/code> instruction instead of performing a redundant load operation.<\/p>\n<pre><code class=\"language-diff\">ldr w1, [fp, #0x28]\n- ldr w0, [fp, #0x28]\n+ mov w0, w1<\/code><\/pre>\n<p>The same PR also eliminated redundant load instructions when the contents were already present in a register. For instance, in the following example, if <code>w1<\/code> had already been loaded from <code>[fp + 0x20]<\/code>, the second load instruction could be removed.<\/p>\n<pre><code class=\"language-diff\">ldr w1, [fp, #0x20]\n- ldr w1, [fp, #0x20]<\/code><\/pre>\n<h3>Convert mul + neg -&gt; mneg<\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/79550\">dotnet\/runtime#79550<\/a>, we transformed two operations involving multiplication and negation instruction to a single <code>mneg<\/code> instruction.<\/p>\n<h2>Code quality improvements<\/h2>\n<p>In .NET 8, we also implemented various enhancements in other aspects of RyuJIT to enhance the performance of Arm64 code.<\/p>\n<h3>Faster Vector128\/Vector64 compare<\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/75864\">dotnet\/runtime#75864<\/a>, we optimized vector comparisons in commonly used algorithms like <code>SequenceEqual<\/code> or <code>IndexOf<\/code>. Let&#8217;s take a look at the code generated before these optimizations.<\/p>\n<pre><code class=\"language-c#\">bool Test1(Vector128&lt;int&gt; a, Vector128&lt;int&gt; b) =&gt; a == b;<\/code><\/pre>\n<p>In the generated code, the first step involves comparing two vectors, <code>a<\/code> and <code>b<\/code>, which are stored in vector registers <code>v0<\/code> and <code>v1<\/code>. This comparison is performed using the <code>cmeq<\/code> instruction that compares two vectors bitwise. This instruction compares bytes in each lane of the two vectors, <code>v0<\/code> and <code>v1<\/code>. If the bytes are equal, it sets the corresponding lane to <code>0xFF<\/code>, otherwise, it sets it to <code>0x00<\/code>. Following the comparison, the <code>uminv<\/code> instruction is used to find the minimum byte among all the lanes. It&#8217;s important to note that the <code>uminv<\/code> instruction has a higher latency because it operates on all lanes after the data for all the lanes is available, and therefore, it does not operate in parallel.<\/p>\n<pre><code class=\"language-asm\">          cmeq    v16.4s, v0.4s, v1.4s\n          uminv   b16, v16.16b\n          umov    w0, v16.b[0]\n          cmp     w0, #0\n          cset    x0, ne<\/code><\/pre>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/75864\">dotnet\/runtime#75864<\/a>, we made a significant improvement by changing the instruction from <code>uminv<\/code> to <code>uminp<\/code> in the generated code. The <code>uminp<\/code> instruction finds the minimum bytes in pairwise lanes, which allows it to operate more efficiently. After that, the <code>umov<\/code> instruction gathers the combined bytes from the first half of the vector. Finally, the <code>cmn<\/code> instruction checks if those bytes are zero or not. This change results in better performance compared to the previous <code>uminv<\/code> instruction because <code>uminp<\/code> can find the minimum value in parallel.<\/p>\n<pre><code class=\"language-asm\">          uminp   v16.4s, v16.4s, v16.4s\n          umov    x0, v16.d[0]\n          cmn     x0, #1\n          cset    x0, eq<\/code><\/pre>\n<p>We observed significant improvements in various core .NET library methods as a result of these optimizations, with performance gains of up to 40%. You can find more details about these improvements in <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/8660\">this report<\/a>.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/pr_75864_1.png\" alt=\"pr_75864_1.png\" \/><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/pr_75864_3.png\" alt=\"pr_75864_3.png\" \/><\/p>\n<h3>Improve vector == Vector128&lt;&gt;.Zero<\/h3>\n<p>Thanks to the suggestion from <a href=\"https:\/\/github.com\/TamarChristinaArm\">@TamarChristinaArm<\/a>, we made further improvements to vector comparisons with the <code>Zero<\/code> vector in <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/75999\">dotnet\/runtime#75999<\/a>. These optimizations resulted in <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/8767\">significant performance improvements<\/a>. In this optimization, we replaced the <code>umaxv<\/code> instruction, which compares the maximum value across all lanes, with the more efficient <code>umaxp<\/code> instruction that performs pairwise comparisons to find the maximum value.<\/p>\n<pre><code class=\"language-csharp\">static bool IsZero1(Vector128&lt;int&gt; v) =&gt; v == Vector128&lt;int&gt;.Zero;<\/code><\/pre>\n<p>In the previous code, we observed that the <code>umaxv<\/code> instruction was generated, which found the maximum value across all lanes and stored it in the 0th lane. Then, the value from the 0th lane was extracted and compared with <code>0<\/code>.<\/p>\n<pre><code class=\"language-asm\">          umaxv   s16, v0.4s\n          umov    w0, v16.s[0]\n          cmp     w0, #0<\/code><\/pre>\n<p>Now, we use <code>umaxp<\/code> instruction which has better latency.<\/p>\n<pre><code class=\"language-asm\">          umaxp   v16.4s, v0.4s, v0.4s\n          umov    x0, v16.d[0]\n          cmp     x0, #0<\/code><\/pre>\n<p>Here is an example of performance improvement of <code>SequenceEqual<\/code> benchmark.\n<img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/pr_75999_1.png\" alt=\"pr_75999_1.png\" \/><\/p>\n<h3>Unroll Memmove<\/h3>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/83740\">dotnet\/runtime#83740<\/a>, we introduced a feature to unroll the memory move operation in the <code>Buffer.Memmove<\/code> API. Without unrolling, the memory move operations need to be done in a loop that operates on smaller chunks of the memory data to be moved. Unrolling optimizes the operation by reducing the loop overhead and enhancing memory access patterns. By copying larger chunks of data in each iteration, it minimizes the number of loop control instructions and leverages modern processors&#8217; ability to perform multiple memory transfers in parallel. This can result in faster and more efficient code execution, especially when working with large data sets or when optimizing critical performance-sensitive code paths. This enhancement is utilized in various APIs such as <code>TryCopyTo<\/code>, <code>ToArray<\/code>, and others.<\/p>\n<p>As demonstrated in <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/14619\">this<\/a> and <a href=\"https:\/\/github.com\/dotnet\/perf-autofiling-issues\/issues\/14647\">this<\/a> performance reports, these optimizations have led to significant performance improvements, with speedups of up to 20%.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/pr_83740.png\" alt=\"pr_83740.png\" \/><\/p>\n<h2>Throughput improvements<\/h2>\n<p>In our continuous efforts to enhance the quality and performance of our code, we are equally dedicated to ensuring that the time taken to produce code remains efficient. This aspect is particularly crucial for applications where startup time is of paramount importance, such as user interface applications. As a result, we have integrated measures to evaluate the Just-In-Time (JIT) compiler&#8217;s throughput in our development process. For those unfamiliar, <a href=\"https:\/\/github.com\/dotnet\/runtime\/blob\/main\/src\/coreclr\/tools\/superpmi\/readme.md\">superpmi<\/a> is an internally developed tool used by RyuJIT for validating the compilation of millions of methods. The <code>superpmi-diff<\/code> functionality is employed to validate changes made to the codebase by comparing the assembly code produced before and after these changes. This comparison not only checks the generated code but also evaluates the time it takes for code compilation. With the recent introduction of our <a href=\"https:\/\/dev.azure.com\/dnceng-public\/public\/_build?definitionId=152&amp;_a=summary\">superpmi-diffs CI pipeline<\/a>, we now systematically assess the JIT throughput impact of every pull request that involves the JIT codebase. This rigorous process ensures that any changes made do not compromise the efficiency of code production, ultimately benefiting the performance of .NET applications.<\/p>\n<p>Efficient throughput is essential for applications with strict startup time requirements. To optimize this aspect, we have implemented several improvements, with a focus on the JIT&#8217;s register allocation process for generating Arm64 code.<\/p>\n<p>In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/85744\">dotnet\/runtime#85744<\/a>, we introduced a mechanism to detect whether a method utilizes floating-point variables. If no floating-point variables are present, we skip the iteration over floating-point registers during the register allocation phase. This seemingly minor change resulted in a throughput gain of up to 0.5%.<\/p>\n<p>During register allocation, the algorithm iterates through various access points of both user-defined variables and internally created variables used for storing temporary results. When determining which register to assign at a given access point, the algorithm traditionally iterated through all possible registers. However, it is more efficient to iterate only over the pre-determined set of allocatable registers at each access point. In <a href=\"https:\/\/github.com\/dotnet\/runtime\/pull\/87424\">dotnet\/runtime#87424<\/a>, we addressed this issue, leading to significant throughput improvements of up to 5%, as demonstrated below:<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/lsra_87424.png\" alt=\"lsra_87424.png\" \/><\/p>\n<p>It&#8217;s important to note that due to the larger number of registers available in Arm64 compared to x64 architecture, these changes resulted in more substantial throughput improvements for Arm64 code generation as compared to x64 targets.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/devblogs.microsoft.com\/dotnet\/wp-content\/uploads\/sites\/10\/2023\/10\/lsra_87424_x64.png\" alt=\"lsra_87424_x64.png\" \/><\/p>\n<h2>Conclusion<\/h2>\n<p>In .NET 8, our collaboration with Arm Holdings engineers resulted in significant feature enhancements and improved code quality for Arm64 platforms. We addressed longstanding issues, introduced peephole optimizations, and adopted advanced instructions such as conditional selection. The addition of consecutive register allocation was a crucial feature that not only enabled instructions like <code>VectorTableLookup<\/code> but also paved the way for future instructions like those capable of loading and storing multiple vectors as seen in the API proposal <a href=\"https:\/\/github.com\/dotnet\/runtime\/issues\/84510\">dotnet\/runtime#84510<\/a>. Looking ahead, our goals also include adding support for advanced architecture features like SVE and SVE2.<\/p>\n<p>We extend our gratitude to the many contributors who have helped us deliver a faster .NET 8 on Arm64 devices.<\/p>\n<p>Thank you for taking the time to explore .NET on Arm64, and please share your feedback with us. Happy coding on Arm64!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>.NET 8 added some key features for new functionality as well as performance improvements for developers including developers targeting Arm64 devices. In this blog I break down everything you need to know about the improvements in .NET 8.<\/p>\n","protected":false},"author":38211,"featured_media":48084,"comment_status":"open","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[685],"tags":[7701,7756,7173,108],"class_list":["post-48083","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-dotnet","tag-dotnet-8","tag-aarch64","tag-arm","tag-performance"],"acf":[],"blog_post_summary":"<p>.NET 8 added some key features for new functionality as well as performance improvements for developers including developers targeting Arm64 devices. In this blog I break down everything you need to know about the improvements in .NET 8.<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/posts\/48083","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/users\/38211"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/comments?post=48083"}],"version-history":[{"count":0,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/posts\/48083\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/media\/48084"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/media?parent=48083"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/categories?post=48083"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/dotnet\/wp-json\/wp\/v2\/tags?post=48083"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}