{"id":36915,"date":"2026-09-16T22:09:50","date_gmt":"2026-09-16T22:09:50","guid":{"rendered":"https:\/\/devblogs.microsoft.com\/cppblog\/?p=36915"},"modified":"2026-09-16T22:09:50","modified_gmt":"2026-09-16T22:09:50","slug":"bringing-correctly-rounded-math-to-production-with-llvm-libc","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/cppblog\/bringing-correctly-rounded-math-to-production-with-llvm-libc\/","title":{"rendered":"Bringing Correctly Rounded Math to Production with LLVM-libc"},"content":{"rendered":"<blockquote><p>As we mentioned in the <a href=\"https:\/\/devblogs.microsoft.com\/cppblog\/msvc-c23-constexpr-cmath-with-llvm-libc\/\">last cmath blog post<\/a>, MSVC is using LLVM-libc for compile-time evaluation and runtime execution of math functions when <code>\/Zc:cmath<\/code> is enabled. I invited LLVM-libc contributors Michael Jones and Tue Ly to write a guest blog post about how they\u2019ve achieved what they\u2019ve achieved. This is that blog post! I hope you enjoy it. I learned a lot from reading it, even after having spent a lot of time in this area.<\/p>\n<p>&#8211; Cody, Microsoft C++ Compiler Frontend Engineer<\/p><\/blockquote>\n<p>Hi, I&#8217;m Michael Jones, the lead maintainer for LLVM-libc. Cody invited me and our math lead, Tue Ly, to write a guest post about LLVM-libc&#8217;s math functions. Tue and I work on LLVM-libc at Google, where we&#8217;ve been working on the project for about five years now.<\/p>\n<p>I&#8217;m going to talk a bit about the history and philosophy of LLVM-libc, then Tue will do a deeper dive into how and why our math functions are so good. I&#8217;m really excited for you all to use the code we&#8217;ve been working on, and thanks to Cody for putting all of this together!<\/p>\n<h2>1. libc as a Library<\/h2>\n<p>Generally, implementations of the C standard library (<code>libc<\/code>) are written for a specific target. They are usually monolithic libraries that are closely tied to the OS interface they&#8217;re based on. They also don&#8217;t make any effort to keep individual functions separate, which makes them difficult to break into smaller chunks.<\/p>\n<p>In 2019, LLVM-libc was started with a different design philosophy. The focus was on portability, modularity, and writing clean C++ that could be reused. The original design describes this as writing &#8220;libc as a library&#8221;, meaning that it would be much closer to an idiomatic C\/C++ library. From the beginning, there has been a separation between the OS interface layer and the public interface, as well as making each individual function independent. Portability has also involved avoiding assembly except where absolutely necessary, which has the added benefit of enabling the compiler to make deeper optimizations.<\/p>\n<p>Building on the modular design, I started Project Hand-in-Hand which allowed LLVM&#8217;s libc++ to use the same float-to-string conversion code as LLVM-libc. Other LLVM projects have also started to integrate LLVM-libc&#8217;s code, such as OpenMP using the printf core and compiler-rt using the software floating point support. Soon both clang and MSVC will be using LLVM-libc&#8217;s math library for constexpr evaluation, but what makes LLVM-libc&#8217;s math suitable for compiler use?<\/p>\n<h2>2. libm as an Opportunity<\/h2>\n<p>I feel like I know more than the average programmer about floating point numbers, but I would not consider myself an expert. Tue Ly is definitely an expert, with the experience and credentials to prove it. He was the one who designed LLVM-libc&#8217;s math library from the ground up, and it was his idea to make it correctly rounded for all rounding modes. This turned out to be hugely important to Google for much the same reason it&#8217;s important to Microsoft: different versions of your math library returning different results can have disastrous consequences when working with heterogeneous distributed systems. I&#8217;ll let Tue go into the details of exactly how the math functions are still performant while being so accurate, as well as the path to get there.<\/p>\n<h2>3. The &#8220;Dark Magic&#8221; of libm and Why Consistency Demands Correct Rounding<\/h2>\n<p>Thanks, Michael! And thanks to Cody for inviting us to contribute to this series.<\/p>\n<p>In the previous blog post, Cody laid out the four dimensions users care about when it comes to math libraries: <strong>Stability<\/strong>,\u00a0<strong>Accuracy<\/strong>, <strong>Consistency<\/strong>, and\u00a0<strong>Performance<\/strong>\u00a0(or, in Cody&#8217;s words,\u00a0<em>&#8220;Numbers go brrr&#8221;<\/em>).<\/p>\n<p>Floating-point math in software is often treated like dark magic. When developers see different results across platforms or compilers, they often shrug and say,\u00a0<em>&#8220;Well, floating-point is inexact.&#8221;<\/em>\u00a0But inexact representation does not mean mathematical operations have to be non-deterministic.<\/p>\n<p>Traditional math libraries (<code>libm<\/code>) frequently produce different results across platforms and versions. In production systems, these differences cause real issues:<\/p>\n<ul>\n<li><strong>Instruction-Level Divergence<\/strong>:\u00a0When hardware designers try to speed up math operations by implementing them directly in silicon (like fast reciprocal square roots or transcendentals), different CPU vendors often make different trade-offs. As a result, the exact same C++ code running on Intel, AMD, ARM, or a GPU can produce slightly different numbers.<\/li>\n<li><strong>The Trap of Bug-for-Bug Compatibility<\/strong>:\u00a0Once an inaccurate math routine ships in a major platform or OS, downstream software inevitably starts relying on those exact bits. A classic example is the legacy x87 hardware transcendental instructions (<code>FSIN<\/code>,\u00a0<code>FCOS<\/code>):\n<ul>\n<li>Intel\u2019s original hardware implementation used a truncated 66-bit approximation of \u03c0 for argument reduction. As Bruce Dawson documented in <a href=\"https:\/\/randomascii.wordpress.com\/2014\/10\/09\/intel-underestimates-error-bounds-by-1-3-quintillion\/\"><strong><em>Intel Underestimates Error Bounds by 1.3 quintillion<\/em><\/strong><\/a>, for certain inputs this causes huge errors, producing inaccuracies of up to <strong><span class=\"katex\"><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"mord\">1.37<\/span><span class=\"mbin\">\u00d7<\/span><\/span><span class=\"base\"><span class=\"mord\">1<\/span><span class=\"mord\">0<\/span><\/span><\/span><\/span><\/strong>\u00a0ULPs (Units in the Last Place)!<\/li>\n<li>To preserve bug-for-bug compatibility with software expecting Intel x87 outputs, AMD adopted the same 66-bit reduction behavior in their\u00a0<a href=\"https:\/\/www.felixcloutier.com\/x86\/fsin\"><strong>x87 instruction set<\/strong><\/a>, locking the x86 ecosystem into the same legacy behavior.<\/li>\n<li>On top of that, because these legacy instructions are microcoded in modern processors, they are actually much\u00a0<em>slower<\/em>\u00a0than optimized software implementations using modern range-reduction algorithms like\u00a0<a href=\"https:\/\/dl.acm.org\/doi\/10.1145\/1057600.1057602\"><strong>Payne-Hanek<\/strong><\/a>.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Breaking Golden Tests<\/strong>:\u00a0Golden tests check that a specific input gives a consistent output, usually by comparing against a fixed &#8220;golden&#8221; result. These are common for image processing pipelines where stability of results is important. Even in software libraries, if a library only guarantees a loose error bound, internal optimizations or bug fixes can shift output bits. A tiny 1-ULP shift in\u00a0<code>std::sin<\/code>\u00a0or\u00a0<code>std::pow<\/code>\u00a0can invalidate hundreds of golden tests requiring engineering teams to spend time updating test baselines.<\/li>\n<\/ul>\n<h3>The Limits of &#8220;Flag-Based&#8221; and Environment-Freezing Workarounds<\/h3>\n<p>Developers have dealt with floating-point differences across platforms for a long time, as seen across game development, physics simulations, and language standardization:<\/p>\n<ul>\n<li><strong>Game Engines &amp; Lockstep Simulations<\/strong>:\u00a0Analyses by\u00a0<a href=\"https:\/\/randomascii.wordpress.com\/2013\/07\/16\/floating-point-determinism\/\"><strong>Bruce Dawson<\/strong><\/a>\u00a0(<em>Random ASCII<\/em>, including articles on\u00a0<a href=\"https:\/\/randomascii.wordpress.com\/2012\/03\/21\/intermediate-floating-point-precision\/\"><strong>intermediate precision<\/strong><\/a>),\u00a0<a href=\"https:\/\/gafferongames.com\/post\/floating_point_determinism\/\"><strong>Glenn Fiedler<\/strong><\/a>\u00a0(<em>Gaffer on Games<\/em>), and\u00a0<a href=\"https:\/\/www.forrestthewoods.com\/blog\/synchronous_rts_engines_and_a_tale_of_desyncs\/\"><strong>Forrest Smith<\/strong><\/a>\u00a0(<em>Planetary Annihilation<\/em>) show how differences in instructions, compiler optimizations, and math libraries routinely desynchronize multiplayer physics engines.<\/li>\n<li><strong>Custom Math Workarounds<\/strong>:\u00a0When building\u00a0<em>Factorio<\/em>, Wube Software found that standard math functions (<code>sin<\/code>,\u00a0<code>cos<\/code>) returned slightly different results between Windows, Linux, and macOS. Without a correctly rounded math library available,\u00a0<a href=\"https:\/\/factorio.com\/blog\/post\/fff-52\"><strong>they had to write their own custom software trigonometry routines<\/strong><\/a>\u00a0to keep multiplayer games in sync. Similarly, open-source RTS games bundled custom math libraries like\u00a0<a href=\"http:\/\/nicolas.brodu.net\/en\/programmation\/streflop\/\"><strong>STREFLOP<\/strong><\/a>\u00a0(<em>Spring RTS \/ Beyond All Reason<\/em>) to get identical calculations across operating systems.<\/li>\n<li><strong>C++ Standardization Proposals<\/strong>:\u00a0The ISO C++ Committee is actively discussing these reproducibility issues across compilers and targets in proposals like\u00a0<a href=\"https:\/\/www.open-std.org\/jtc1\/sc22\/wg21\/docs\/papers\/2025\/p3375r2.html\"><strong>P3375<\/strong><\/a>\u00a0(<em>&#8220;Reproducible floating-point results&#8221;<\/em>,\u00a0<em>Davidson et al.<\/em>).<\/li>\n<\/ul>\n<p>However, all of these traditional workarounds share an inherent limitation:\u00a0<strong>while compiler flags and hardware normalization can achieve reproducibility for basic arithmetic operations (<span class=\"katex\"><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"mord\">+<\/span><\/span><\/span><\/span>,\u00a0<span class=\"katex\"><span class=\"katex-mathml\">\u2212<\/span><\/span>,\u00a0<span class=\"katex\"><span class=\"katex-mathml\">\u00d7<\/span><\/span>,\u00a0<span class=\"katex\"><span class=\"katex-mathml\">\/<\/span><\/span>, \u221a), they fall short for transcendental functions (<code>sin<\/code>,\u00a0<code>cos<\/code>,\u00a0<code>exp<\/code>,\u00a0<code>log<\/code>, etc.), which form a substantial part of standard math.h.<\/strong><\/p>\n<p>Basic arithmetic operations are strictly specified and hardware-mandated to be correctly rounded under IEEE-754. But transcendental functions in standard\u00a0math.h\u00a0are almost always software approximations where IEEE-754 has historically not mandated bit-level exactness. Consequently, their output bits are simply an artifact of a specific library&#8217;s polynomial approximations, range-reduction splits, or lookup table values.<\/p>\n<p>The moment a library maintainer updates a function to make it faster, reduce binary size, or vectorize it, the least significant bits will inevitably shift for many inputs.<\/p>\n<p>Freezing compiler flags or using custom lookup tables cannot protect you from future library updates. You are left choosing between never updating dependencies or frequently updating downstream tests.<\/p>\n<h3>Why Correct Rounding Solves This<\/h3>\n<p>This is where\u00a0<strong>correct rounding<\/strong>\u00a0fixes the problem.<\/p>\n<p>Under IEEE-754, if a math function is guaranteed to be\u00a0<strong>correctly rounded<\/strong>\u00a0across all standard rounding modes (round-to-nearest, round-up, round-down, and round-to-zero),\u00a0<strong>the output bit pattern for any given input is mathematically unique<\/strong>, meaning there is only one correct answer.<\/p>\n<p>By solving for\u00a0<strong>Accuracy<\/strong>\u00a0through correct rounding, you automatically get\u00a0<strong>Stability<\/strong>\u00a0and\u00a0<strong>Consistency<\/strong>\u00a0for free. Because the output is mathematically determined rather than an implementation detail, library authors can completely rewrite, optimize, and vectorize their algorithms across future releases without shifting output bits. Your math functions return the exact same bits on x86-64, ARM64, Windows, Linux, across datacenters, across library updates, and, with C++23 (<a href=\"https:\/\/www.open-std.org\/jtc1\/sc22\/wg21\/docs\/papers\/2021\/p0533r9.pdf\"><strong>P0533R9<\/strong><\/a>) and C++26 (<a href=\"https:\/\/www.open-std.org\/jtc1\/sc22\/wg21\/docs\/papers\/2023\/p1383r2.pdf\"><strong>P1383R2<\/strong><\/a>) adding\u00a0constexpr\u00a0math, between compile-time evaluation and runtime execution.<\/p>\n<h2>4. Standing on the Shoulders of Giants: The Road to Production Correct Rounding<\/h2>\n<p>In his blog post, Cody joked that writing a complete, accurate math library seemed to require\u00a0<em>&#8220;years of monk-like study.&#8221;<\/em>\u00a0Cody wasn&#8217;t wrong, except it was decades of research across the computer arithmetic community.<\/p>\n<h3>The Table-Maker&#8217;s Dilemma<\/h3>\n<p>For decades, correct rounding for transcendental functions (<em>e<span class=\"katex\"><span class=\"katex-mathml\">\u02e3<\/span><\/span><\/em>, log\u2009<i>x<\/i>, <span class=\"katex\"><span class=\"katex-html\" aria-hidden=\"true\"><span class=\"base\"><span class=\"mop\">sin\u2009<i>x<\/i><\/span><\/span><\/span><\/span>, etc.) was considered too slow for production systems due to the <a href=\"https:\/\/en.wikipedia.org\/wiki\/Rounding#Table-maker's_dilemma\"><strong>Table-Maker&#8217;s Dilemma<\/strong><\/a>.<\/p>\n<p>The dilemma is simple to state but hard to solve: when computing an approximation <em>\u0177<\/em> of a transcendental function <em><span class=\"katex-mathml\">y = f(x)<\/span><\/em>, how many extra bits of intermediate precision do you need to guarantee that rounding <em>\u0177 <\/em>produces the exact same floating-point value as rounding the exact mathematical value <em><span class=\"katex-mathml\">f(x) <\/span><\/em>?<\/p>\n<p><a href=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2.webp\"><img decoding=\"async\" class=\"aligncenter wp-image-36978 size-large\" src=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-1024x788.webp\" alt=\"Depicts a number line centered on the rounding midpoint between two representable floating-point values. The approximation interval contains the rounding midpoint.\" width=\"1024\" height=\"788\" srcset=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-1024x788.webp 1024w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-300x231.webp 300w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-768x591.webp 768w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-1536x1182.webp 1536w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/round2-2048x1577.webp 2048w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/p>\n<p>If the true mathematical result falls very close to a rounding midpoint (for round-to-nearest) or a representable floating-point boundary (for directed rounding), an approximation with standard precision cannot tell whether to round up or down.<\/p>\n<p>The first major attempt to bring correctly rounded transcendentals to production was Abraham Ziv\u2019s work on\u00a0<strong>IBM\u00a0<code>libultim<\/code><\/strong>\u00a0in the 1990s (later used in\u00a0<code>glibc<\/code>). Ziv introduced the multi-stage evaluation strategy: compute a fast approximation; if the result is too close to a rounding boundary (Ziv&#8217;s rounding test), fall back to a higher-precision path.<\/p>\n<p>However, because the maximum required precision was unknown at the time,\u00a0<code>libultim<\/code>&#8216;s fallback path used arbitrary-precision arithmetic scaling up to\u00a0<strong>768 bits of precision<\/strong>\u00a0(via internal helpers like\u00a0slowpow.c\u00a0and\u00a0mppow.c). When programs hit one of these hard-to-round inputs, performance dropped sharply: while a normal fast-path\u00a0pow\u00a0took ~70 cycles, the 768-bit fallback took\u00a0up to <strong>440,000 cycles<\/strong>\u00a0(as documented in\u00a0<strong><a href=\"https:\/\/sourceware.org\/bugzilla\/show_bug.cgi?id=13932\">glibc Bug 13932<\/a><\/strong>\u00a0and\u00a0<a href=\"https:\/\/sourceware.org\/bugzilla\/show_bug.cgi?id=16898\"><strong>Bug 16898<\/strong><\/a>) which is a\u00a0<strong>~6,000\u00d7<\/strong> slowdown\u00a0on a single function call!<\/p>\n<p>These unpredictable slowdowns caused performance issues in production workloads, eventually leading\u00a0glibc\u00a0maintainers to remove\u00a0<code>libultim<\/code>\u00a0in favor of faster, non-correctly rounded routines with bounded latency. As a result, correct rounding got a reputation for being &#8220;too slow for production.&#8221;<\/p>\n<h3>Hunting the Worst Cases<\/h3>\n<p>To avoid these slowdowns, the arithmetic research community needed to answer a fundamental question:\u00a0<strong>What is the worst-case hard-to-round input across the entire floating-point domain?<\/strong>\u00a0If you know the worst case, you know the maximum precision ever required, and you can build a bounded, non-allocating fallback path.<\/p>\n<p>Finding these worst cases is very difficult. In double precision, searching the <span class=\"mord\">2\u2076\u2074<\/span>\u00a0input space by brute force was long considered computationally prohibitive.<\/p>\n<p>The breakthrough began with the foundational work of\u00a0<strong>Vincent Lef\u00e8vre<\/strong>\u00a0and\u00a0<strong>Jean-Michel Muller<\/strong>\u00a0(Ar\u00e9naire \/ AriC at ENS Lyon \/ CNRS \/ Inria). In pioneering publications like\u00a0<a href=\"https:\/\/doi.org\/10.1109\/12.736435\"><strong><em>Toward correctly rounded transcendentals<\/em><\/strong><\/a>\u00a0(<em>Lef\u00e8vre, Muller, &amp; Tisserand, 1998<\/em>) and\u00a0<a href=\"https:\/\/inria.hal.science\/inria-00072594\/\"><strong><em>Worst Cases for Correct Rounding of the Elementary Functions in Double Precision<\/em><\/strong><\/a>\u00a0(<em>Lef\u00e8vre &amp; Muller, 2001<\/em>), they proved that the worst-case precision required for double-precision functions could be systematically determined. Their work culminated in the release of\u00a0<a href=\"https:\/\/ens-lyon.hal.science\/ensl-01529804\/file\/crlibm.pdf\"><strong>CR-LIBM<\/strong><\/a>, demonstrating for the first time that double-precision elementary functions could be correctly rounded with reasonably bounded intermediate precision.<\/p>\n<p>Over the last several years, this search was completed. Through our close collaboration with\u00a0<strong>Paul Zimmermann<\/strong>\u00a0and\u00a0<strong>Vincent Lef\u00e8vre<\/strong>, culminating in our recent paper\u00a0<a href=\"https:\/\/inria.hal.science\/hal-05593313\"><strong><em>Computing hard-to-round cases of sin, cos, tan in double precision<\/em><\/strong><\/a>\u00a0(<a href=\"https:\/\/arith2026.org\/program.html\"><strong>ARITH 2026<\/strong><\/a>), we have now systematically searched and proven the worst-case bounds for\u00a0<strong>all standard univariate double-precision math functions<\/strong>.<\/p>\n<p>Most importantly,\u00a0<strong>all of these worst-case inputs and mathematical bounds are openly hosted and maintained in the<\/strong>\u00a0<a href=\"https:\/\/gitlab.inria.fr\/core-math\/core-math\"><strong>CORE-MATH project repository<\/strong><\/a>\u00a0(and documented at\u00a0<a href=\"https:\/\/core-math.gitlabpages.inria.fr\/\"><strong>core-math.gitlabpages.inria.fr<\/strong><\/a>), serving as an open, shared ground truth for library implementers worldwide.<\/p>\n<h3>Bringing Correct Rounding to Production<\/h3>\n<p>Around 2020, when we started architecting the math library for LLVM-libc, several independent developments were converging:<\/p>\n<ol>\n<li><strong>Proven Mathematical Bounds<\/strong>:\u00a0Decades of search algorithms had established the hardest-to-round cases, proving that\u00a0<strong>128-bit intermediate precision is sufficient<\/strong>\u00a0for most univariate double-precision functions.<\/li>\n<li><strong>Modern 64-bit Hardware<\/strong>:\u00a0Modern processors had native hardware Fused Multiply-Add (FMA), wide vector pipelines, and rich register files capable of executing branchless 128-bit integer arithmetic in a handful of cycles.<\/li>\n<\/ol>\n<p>The realization that the timing was finally right wasn&#8217;t unique to LLVM-libc. Around the exact same time, two other major research efforts started with the same objective:<\/p>\n<ul>\n<li><strong>The\u00a0<\/strong><a href=\"https:\/\/core-math.gitlabpages.inria.fr\/\"><strong>CORE-MATH project<\/strong><\/a>\u00a0led by\u00a0<strong>Paul Zimmermann<\/strong>\u00a0at Inria, focusing on open-source, correctly rounded reference implementations.<\/li>\n<li><strong>The\u00a0<\/strong><a href=\"https:\/\/people.cs.rutgers.edu\/~sn349\/rlibm\/\"><strong>RLIBM project<\/strong><\/a>\u00a0led by\u00a0<strong>Santosh Nagarakatte<\/strong>\u00a0at Rutgers University, exploring novel polynomial generation techniques for correct rounding.<\/li>\n<\/ul>\n<p>From the very beginning, the LLVM-libc math team established an active collaboration with CORE-MATH and RLIBM: discussing algorithm approaches, sharing hard-to-round test vectors, developing parallel implementations, and cross-validating each other&#8217;s implementations.<\/p>\n<p>This collaboration helped us verify correctness, optimize performance, and prove that correctly rounded math is fast enough for production use.<\/p>\n<h2>5. Under the Hood: Making Numbers Go Brrr with Correct Rounding<\/h2>\n<p>When developers hear that LLVM-libc guarantees correct rounding, their immediate reaction is often skepticism:\u00a0<em>&#8220;If you&#8217;re checking every last bit against the exact mathematical value, how is your library not slow?&#8221;<\/em><\/p>\n<p>The answer lies in how we structure our execution pipeline around modern hardware strengths:<\/p>\n<p><a href=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow.webp\"><img decoding=\"async\" class=\"aligncenter wp-image-36982 size-large\" src=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow-1024x926.webp\" alt=\"Depicts the control flow of mathematical functions. All functions execute a fast computation. Greater than 99.99% of inputs pass Ziv's rounding test and return immediately. Those that don't go through a bounded accuracy pass.\" width=\"1024\" height=\"926\" srcset=\"https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow-1024x926.webp 1024w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow-300x271.webp 300w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow-768x694.webp 768w, https:\/\/devblogs.microsoft.com\/cppblog\/wp-content\/uploads\/sites\/9\/2026\/09\/flow.webp 1535w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/p>\n<p>When an input comes in, we don&#8217;t immediately jump into high-precision math. Instead, our first stage evaluates a fast, carefully tuned Taylor or minimax polynomial approximation using double-double or truncated double-double arithmetic to achieve 70-106 bits of precision, heavily accelerated by native hardware FMA instructions. We break the function&#8217;s domain into small intervals such that lookup tables (if needed) comfortably reside in the CPU&#8217;s L1 cache.<\/p>\n<p>Once the fast path computes this higher-than-double approximation, we perform Ziv\u2019s rounding test: we check whether the error interval contains a rounding boundary. If the answer is clear (which happens for\u00a0<strong>over 99.99% of all possible inputs<\/strong>), we round to the target precision and return immediately. In terms of latency and throughput, this fast path is very similar to fast, non-correctly rounded math libraries.<\/p>\n<p>What about that rare fraction of a percent (&lt;0.01%) of tricky, hard-to-round inputs where the fast path is inconclusive?<\/p>\n<p>Because the maximum intermediate precision is mathematically proven to never exceed 128 bits, our second stage never needs to invoke an arbitrary-precision library or allocate memory on the heap. Instead, it executes a fixed, branchless 128-bit or 256-bit integer routine with a predictable, bounded cycle count. There are no latency cliffs.<\/p>\n<h3>The Freedom to Optimize<\/h3>\n<p>Guaranteed correct rounding completely changes the economics of math library development. Historically, maintainers avoided modifying math routines because any tweak, such as changing a polynomial degree or shrinking a lookup table, would shift output bits and break downstream regression tests. With correct rounding, that concern disappears. Because the output is mathematically determined, we are free to explore all kinds of performance improvements: tuning polynomial degrees, optimizing cache footprints, leveraging target-specific hardware instructions (like FMA and SIMD), or rewriting entire fast paths without worrying about breaking downstream software.<\/p>\n<p>Additionally, implementing the entire library in clean, standard C++ gives the compiler optimizers full visibility into the code. Rather than relying on inline assembly or platform-specific constructs that obscure dataflow, modern compilers can freely inline helper functions, schedule instructions to increase instruction-level parallelism, and auto-vectorize loops across different architectures.<\/p>\n<h2>Conclusion &amp; Joining the Community<\/h2>\n<p>Bringing correctly rounded math into production has been a multi-year effort spanning academic pioneers, open-source maintainers, and industry partners.<\/p>\n<p>Today, software stacks are becoming increasingly complex and distributed. With machine learning models and mixed-precision pipelines already introducing non-determinism, having standard math functions return identical results across platforms and versions helps keep software reliable and predictable.<\/p>\n<p>We are very excited to see MSVC adopting LLVM-libc\u2019s math library for constant evaluation and more in C++23, bringing deterministic floating-point math to all the Windows developers worldwide.<\/p>\n<p>If you&#8217;re interested in contributing to LLVM-libc, either for the floating-point math effort or the overall standard library development, we&#8217;re always looking for new contributors! You can find more information including how to set up a local build at\u00a0<a href=\"https:\/\/libc.llvm.org\/\"><strong>https:\/\/libc.llvm.org\/<\/strong><\/a>. You can also find the community on\u00a0<a href=\"https:\/\/discord.gg\/xS7Z362\"><strong>the LLVM discord<\/strong><\/a>\u00a0in the\u00a0#libc\u00a0channel or join one of our regular public meetings. We hope to see you there!<\/p>\n<p>Michael Jones<\/p>\n<p>Tue Ly<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>As we mentioned in the last cmath blog post, MSVC is using LLVM-libc for compile-time evaluation and runtime execution of math functions when \/Zc:cmath is enabled. I invited LLVM-libc contributors Michael Jones and Tue Ly to write a guest blog post about how they\u2019ve achieved what they\u2019ve achieved. This is that blog post! I hope [&hellip;]<\/p>\n","protected":false},"author":220230,"featured_media":35994,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-36915","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-cplusplus"],"acf":[],"blog_post_summary":"<p>As we mentioned in the last cmath blog post, MSVC is using LLVM-libc for compile-time evaluation and runtime execution of math functions when \/Zc:cmath is enabled. I invited LLVM-libc contributors Michael Jones and Tue Ly to write a guest blog post about how they\u2019ve achieved what they\u2019ve achieved. This is that blog post! I hope [&hellip;]<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts\/36915","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/users\/220230"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/comments?post=36915"}],"version-history":[{"count":1,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts\/36915\/revisions"}],"predecessor-version":[{"id":36984,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts\/36915\/revisions\/36984"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/media\/35994"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/media?parent=36915"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/categories?post=36915"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/tags?post=36915"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}