{"id":3803,"date":"2010-07-07T13:39:00","date_gmt":"2010-07-07T13:39:00","guid":{"rendered":"https:\/\/blogs.msdn.microsoft.com\/vcblog\/2010\/07\/07\/how-we-test-the-compiler-performance\/"},"modified":"2019-02-18T18:45:29","modified_gmt":"2019-02-18T18:45:29","slug":"how-we-test-the-compiler-performance","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/cppblog\/how-we-test-the-compiler-performance\/","title":{"rendered":"How we test the compiler performance"},"content":{"rendered":"<p class=\"MsoNormal\"><span style=\"font-family: Calibri;font-size: small\">The C++ back-end team is very conscious of the performance of our product.&nbsp; Today I will present to you an overview of how we define &ldquo;performance of our product&rdquo; and the way we measure it.&nbsp; Along the way I hope to introduce you to some new ideas that you can use to test your product&rsquo;s performance as well. You can read Alex Thaman&rsquo;s blog post on <\/span><a href=\"http:\/\/blogs.msdn.com\/b\/vcblog\/archive\/2010\/06\/01\/how-we-test-the-compiler-backend.aspx\"><span style=\"font-family: Calibri;color: #0000ff;font-size: small\">&ldquo;How we test the compiler backend&rdquo;<\/span><\/a><span style=\"font-family: Calibri;font-size: small\"> for some background on compiler testing.<span>&nbsp; <\/span>Additionally, you can read Asmaa Taha&rsquo;s blog post on <\/span><a href=\"http:\/\/blogs.msdn.com\/b\/vcblog\/archive\/2009\/03\/16\/performance-tests.aspx\"><span style=\"font-family: Calibri;color: #0000ff;font-size: small\">VCBench, our performance test automation system<\/span><\/a><span style=\"font-size: small\"><span style=\"font-family: Calibri\">.<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"color: #1f497d\"><span><\/p>\n<p><span style=\"font-family: Calibri;font-size: small\">&nbsp;<\/span><\/p>\n<p><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span>Questions before answers<\/span><span style=\"line-height: 115%;font-size: 11pt\"><\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Some of the questions that you should answer before you start measuring the performance of your product are:<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">What features of my product will customers need to be performant?<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">What scenarios can I run to reliably measure performance?<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">What metrics can I measure that will produce actionable results?<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">How can I execute the tests to minimize variability?<\/p>\n<p><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\">Picking what features to measure<\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Understanding customer needs is the first step towards figuring out what feature areas they will need to be performant.&nbsp; Some of the avenues that we identify performance scenarios through are:<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Conversations with customers through conferences, meetings, surveys, blogs, etc.<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Connect bugs reported by you (a lot of attention given to connect bugs)<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Feedback from internal Microsoft teams<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Common historical performance bug reports<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\">\n<p><span style=\"font-family: Calibri;font-size: small\">&nbsp;<\/span><\/p>\n<\/p>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Based on this data we break down performance of the Visual C++ compiler into three areas:<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Compiler<b><span style=\"color: #00b050\"> Throughput<\/span><\/b> &ndash; seconds to compile a set of sources with a set of compiler parameters<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: 'Courier New'\"><span><span style=\"font-size: small\">&nbsp;&nbsp;&nbsp; o<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Especially important for customers who have to wait for builds to finish<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Generated <b><span style=\"color: #00b050\">Code Quality<\/span><\/b> &ndash; seconds to run generated executable with a fixed workload<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: 'Courier New'\"><span><span style=\"font-size: small\">&nbsp;&nbsp;&nbsp; o<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Especially important for technical computing, graphics, gaming and other performance intensive applications<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Generated <b><span style=\"color: #00b050\">Code Size<\/span><\/b> &ndash; number of bytes in the executable section(s) of the binary<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: 'Courier New'\"><span><span style=\"font-size: small\">&nbsp;&nbsp;&nbsp; o<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">Especially important for applications that need to run on less memory<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"color: #1f497d\"><\/p>\n<p><span style=\"font-family: Calibri;font-size: small\">&nbsp;<\/span><\/p>\n<p><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\">When deciding whether or not to add an optimization algorithm we look at the impact to these three areas on a number of benchmarks.&nbsp; In a perfect world the Visual C++ compiler could take infinite time to compile and generate the perfect binary.&nbsp; Realistically we make tradeoffs to generate the most optimized binary possible in a reasonable amount of time.&nbsp; We try to make sure that the optimization algorithms that we use have an appropriate compilation time cost to code quality &amp; code size benefit ratio.<\/p>\n<p><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\">What scenarios do we run\/what metrics do we measure?<\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Results for correctness tests are easy to report &#8212; pass or fail and in the case of fail give enough information about what failed to dig into the issue.&nbsp; Performance tests are trickier.&nbsp; Instead of a black or white pass\/fail, you have &ldquo;we&rsquo;re fairly sure there is no appreciable change&rdquo;, &ldquo;we&rsquo;re fairly sure there is a significant change&rdquo;, or &ldquo;there is too much variation to tell whether something changed&rdquo;.&nbsp; To make matters worse, some areas of performance are changing while others are not; some are improving and some are regressing.&nbsp; On one extreme you can report thousands of different metrics for each tiny part of the performance that can change.&nbsp; This will drown your developers in numbers so that it takes them forever to interpret them.&nbsp; On the other extreme if you report too few numbers or the wrong numbers you will be hiding important changes and therefore missing regressions.<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">In general you want to report as few metrics as possible that give you the fullest coverage of your identified scenarios.&nbsp; Given that [large] caveat, here are the metrics that we have decided to monitor<span style=\"color: #1f497d\"> <\/span>and report:<\/p>\n<p><\/span><\/span><\/p>\n<h3 style=\"margin: 10pt 0in 0pt\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span style=\"font-size: small\"><span>Throughput<\/span><span style=\"color: #1f497d\"><\/p>\n<p><\/span><\/span><\/span><\/span><\/h3>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">We measure the compilation time of certain critical parts of Windows and SQL as well as smaller projects on a daily basis in order to track our <b><span style=\"color: #00b050\">throughput<\/span><\/b>.&nbsp; On a less regular basis we time how long it takes to build all of Windows.&nbsp; The metrics that we gather are for the front-end, back-end and linker.&nbsp; This sums up to the entire compilation time, however it lets us be more granular in triaging where a performance regression came from.<i><\/p>\n<p><\/i><\/span><\/span><\/p>\n<h3 style=\"margin: 10pt 0in 0pt\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span style=\"font-size: small\"><span>Code Quality \/ Code Size<\/span><span style=\"color: #1f497d\"><\/p>\n<p><\/span><\/span><\/span><\/span><\/h3>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\">We build and run a set of industry standard integer and floating point benchmarks in order to monitor the code size and code quality of our optimized code generation.&nbsp; Each benchmark is real world code, but with the special constraints of being CPU bound (no waiting on user input or heavily reading\/writing to the disk).&nbsp; Some of the benchmarks we run exercise: cryptographic algorithms, XML string processing, compression, mathematical modeling, and artificial intelligence.&nbsp; We also measure NT\/SQL for code size changes.<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri;font-size: small\">For <b><span style=\"color: #00b050\">code size<\/span><\/b> we use &ldquo;dumpbin.exe \/headers &lt;generated_binary&gt;&rdquo; and then accumulate the virtual size of all the sections marked as executable.&nbsp; We do this instead of the simpler &ldquo;how big is the entire binary&rdquo; because if 98% of a binary is data (read: strings, images, etc) and we double code size then it only reads as a 2% code size regression.&nbsp; This is an example of where carefully choosing what metric you report allows us to more accurately see how our optimizations affect the size of the executable code sections of a binary.&nbsp; <\/span><a href=\"http:\/\/msdn.microsoft.com\/en-us\/library\/c1h23y6c(v=VS.100).aspx\"><span style=\"font-family: Calibri;color: #0000ff;font-size: small\">Dumpbin.exe<\/span><\/a><span style=\"font-size: small\"><span style=\"font-family: Calibri\"> is a tool that ships with Visual Studio.<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\">For <b><span style=\"color: #00b050\">code quality<\/span><\/b> we use whatever metric the benchmark deems as appropriate for measuring the quality of the generated code.&nbsp; In many instances this is execution time, but in some cases it may be something like &ldquo;abstraction penalty&rdquo; or &ldquo;iterations per second&rdquo;.&nbsp; It is important to note whether a larger number in the metric is better or worse so that developers know whether they are improving or regressing it!<\/p>\n<p><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\">Minimizing variability<\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\"><span style=\"color: black\">One of the biggest problems in tracking performance is that, with the exception of code size, the results that are <\/span>produced vary from one run to the next<span style=\"color: black\">.&nbsp; Unlike correctness test cases which <\/span><b><span style=\"color: #00b050\">pass<\/span><\/b><span style=\"color: #00b050\"> <\/span><span style=\"color: black\">or <\/span><b><span style=\"color: red\">fail<\/span><\/b><span style=\"color: black\">, performance tests are susceptible to machine variances that can hide a performance change or show a performance change when there is none.&nbsp; To increase the fidelity of our results we run performance tests multiple times and merge the results into an aggregate value.&nbsp; This results in better accuracy of the results and removes outliers.<span class=\"msoDel\"><del datetime=\"2010-06-16T13:34\"><\/p>\n<p><\/del><\/span><\/span><\/span><\/span><\/p>\n<h3 style=\"margin: 10pt 0in 0pt\"><span><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span style=\"font-size: small\">Use an appropriate aggregation method&nbsp;&nbsp;&nbsp; <\/p>\n<p><\/span><\/span><\/span><\/span><\/h3>\n<p class=\"MsoNormal\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\"><span style=\"color: black\">We have found that performance results do not normally have a normal distribution; they are much closer to a <\/span>skew<span style=\"color: black\"> distribution.&nbsp; Because of this using the median to aggregate results is better than using <\/span>a mean <span style=\"color: black\">of all of the results &ndash; it is less <\/span>susceptible to <span style=\"color: black\">large outliers.&nbsp; You can study the characteristic of the results you are getting out of your performance runs and determine the best method of aggregation.&nbsp; It may end up that using the minimum, maximum or a quartile result is best if you only have outliers in one direction.&nbsp; Our team primarily uses median.<\/p>\n<p><\/span><\/span><\/span><\/p>\n<h3 style=\"margin: 10pt 0in 0pt\"><span><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span style=\"font-size: small\">Remove outliers<\/p>\n<p><\/span><\/span><\/span><\/span><\/h3>\n<p class=\"MsoNormal\"><span style=\"color: black\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Here are some simple methods to remove outliers:<\/p>\n<p><\/span><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Remove the top and bottom X% of results<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"font-family: Symbol\"><span><span style=\"font-size: small\">&middot;<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-family: Calibri\"><span style=\"font-size: small\">Remove all results outside of X standard deviations from the sample mean\/median<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"color: #1f497d\"><\/p>\n<p><span style=\"font-family: Calibri;font-size: small\">&nbsp;<\/span><\/p>\n<p><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"color: black\"><span style=\"font-family: Calibri\"><span style=\"font-size: small\">In addition to throwing away general outliers, there can be a significant difference in the performance of a binary when it is not already loaded in the cache.&nbsp; Throwing out the first iteration as a warm up run can counteract this.&nbsp; Note: If your customer usage scenario is that they will primarily be executing it on a cold cache then you need to be concerned with this number!<\/p>\n<p><\/span><\/span><\/span><\/p>\n<h3 style=\"margin: 10pt 0in 0pt\"><span><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span style=\"font-size: small\">Stabilize the machine before running the benchmark<\/p>\n<p><\/span><\/span><\/span><\/span><\/h3>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\"><span style=\"color: black\">The average windows machine has a lot of services running on it from SQL server to IIS to anti-virus.&nbsp; Shutting down as many applications and services as possible before executing your benchmark will help to reduce variability.&nbsp; Anti-virus in particular is important to disable because of all the places it can hook <\/span>into the system and introduce additional overhead.<span style=\"color: #1f497d\">&nbsp; <\/span><span style=\"color: black\">Disabling your network adapters will also help to reduce noise.<\/span><span style=\"color: #1f497d\"><\/p>\n<p><\/span><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\"><span>Making results <span style=\"text-decoration: underline\">Actionable<\/span><\/span><span style=\"color: windowtext\"><\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\">At this point you have a set of benchmarks that run for a number of iterations on a stabilized machine, and you are aggregating the set of result iterations into two numbers that are your baseline results and your changed results.&nbsp; We now need to compare these results and determine whether they are:<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span><span><span style=\"font-family: Calibri;font-size: small\">1.<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">The same (no performance change)<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span><span><span style=\"font-family: Calibri;font-size: small\">2.<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">Different (a statistically significant change)<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span><span><span style=\"font-family: Calibri;font-size: small\">&nbsp;&nbsp;&nbsp; a.<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">Significant improvement<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span><span><span style=\"font-family: Calibri;font-size: small\">&nbsp;&nbsp;&nbsp; b.<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">Significant regression<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span><span><span style=\"font-family: Calibri;font-size: small\">3.<\/span><span style=\"font: 7pt 'Times New Roman'\">&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <\/span><\/span><\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\">Unactionable (we cannot tell if they are the same or different)<\/p>\n<p><\/span><\/span><\/p>\n<p class=\"MsoNoSpacing\"><span style=\"color: #1f497d\"><\/p>\n<p><span style=\"font-family: Calibri;font-size: small\">&nbsp;<\/span><\/p>\n<p><\/span><\/p>\n<p class=\"MsoNormal\"><span style=\"color: black\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\">If the results are accurate enough, or the results are separated enough it can be easy to eyeball the numbers and tell what category the numbers fall in.&nbsp; However, if the numbers are fairly close or you are looking to fully automate this action you will need a more concrete algorithm for categorizing the results.&nbsp; This is where Confidence Intervals are useful &ndash; they allow you to say wi\nth certain confidence levels that two sets of results are either: the same, different, or that you need more iterations to tell one way or another.<\/p>\n<p><\/span><\/span><\/span><\/p>\n<h2 style=\"margin: 10pt 0in 0pt\"><span><span style=\"font-size: medium\"><span style=\"color: #4f81bd\"><span style=\"font-family: Cambria\">Further Reading<\/p>\n<p><\/span><\/span><\/span><\/span><\/h2>\n<p class=\"MsoNormal\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\"><span style=\"color: black\">For more information on confidence intervals, refer to your favorite statistics book.&nbsp; <\/span>A good book that <span style=\"color: black\">covers confidence intervals as well as a lot more about performance analysis read <span style=\"text-decoration: underline\">The Art of Computer Systems Performance Analysis<\/span> by Raj Jain.<\/p>\n<p><\/span><\/span><\/span><\/p>\n<h2 class=\"MsoNormal\" style=\"margin: 0in 0in 10pt\"><span style=\"font-size: small\"><span style=\"font-family: Calibri\"><span style=\"color: black\"><\/p>\n<p><\/span><\/span><\/span><\/h2>\n<p><span style=\"font-family: Calibri;font-size: small\">Thank you for your time<\/span><span style=\"font-size: small\"><span style=\"font-family: Calibri\"><span style=\"color: black\">,<br \/>~Pete Steijn<br \/><\/span>Software Development Engineer in Test<\/span><\/span><span style=\"color: black\"><br \/><span style=\"font-size: small\"><span style=\"font-family: Calibri\">VC++ Code Generation Team &#8211; Performance<\/p>\n<p><\/span><\/span><\/span><\/p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The C++ back-end team is very conscious of the performance of our product.&nbsp; Today I will present to you an overview of how we define &ldquo;performance of our product&rdquo; and the way we measure it.&nbsp; Along the way I hope to introduce you to some new ideas that you can use to test your product&rsquo;s [&hellip;]<\/p>\n","protected":false},"author":289,"featured_media":35994,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1],"tags":[65,66,67,36],"class_list":["post-3803","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-cplusplus","tag-compiler","tag-performance","tag-testing","tag-vc"],"acf":[],"blog_post_summary":"<p>The C++ back-end team is very conscious of the performance of our product.&nbsp; Today I will present to you an overview of how we define &ldquo;performance of our product&rdquo; and the way we measure it.&nbsp; Along the way I hope to introduce you to some new ideas that you can use to test your product&rsquo;s [&hellip;]<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts\/3803","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/users\/289"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/comments?post=3803"}],"version-history":[{"count":0,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/posts\/3803\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/media\/35994"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/media?parent=3803"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/categories?post=3803"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/cppblog\/wp-json\/wp\/v2\/tags?post=3803"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}