{"id":40073,"date":"2004-03-25T17:07:00","date_gmt":"2004-03-25T17:07:00","guid":{"rendered":"https:\/\/blogs.msdn.microsoft.com\/oldnewthing\/2004\/03\/25\/regular-expressions-and-the-dreaded-operator\/"},"modified":"2004-03-25T17:07:00","modified_gmt":"2004-03-25T17:07:00","slug":"regular-expressions-and-the-dreaded-operator","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/oldnewthing\/20040325-00\/?p=40073\/","title":{"rendered":"Regular expressions and the dreaded *? operator"},"content":{"rendered":"<p>The regular expression *? operator means &#8220;Match as few characters  as necessary to make this pattern succeed.&#8221;  But look at what  happens when you mix it up a bit:\n  <code>\".*?\"<\/code>\n  This pattern matches a quoted string containing no embedded quotes.  This works because the first quotation mark starts the string,  the .*? gobbles up everything in between, and then the second   quotation mark eats the close-quote.\n  (Note how this differs from <code>\".*\"<\/code>, which uses a greedy  match.  This time, the .* operator is perfectly happy to gobble  up quotation marks, as long as it leaves one to match the second  quotation mark in the pattern.)\n  Okay, great, now let&#8217;s make a small change to the above pattern:\n  <code>\".*?\"&gt;<\/code>\n  All I did was stick a &gt; at the end of the pattern.  This would  therefore match a quoted string (containing no quotation marks)  followed by a &gt; character, right?\n  Wrong.\n  There&#8217;s nothing in .*? that says &#8220;no quotation marks allowed&#8221;.  It just says &#8220;Don&#8217;t match more than you need to.&#8221;  But there are  strings where it needs to match a quotation mark.  Consider:<\/p>\n<table>\n<col style=\"font-family: monospace;border: solid .5pt black\" align=\"center\">\n<tr>\n<td><code>\"<\/code><\/td>\n<td><code>hello\"world<\/code><\/td>\n<td><code>\"<\/code><\/td>\n<td><code>&gt;<\/code><\/td>\n<\/tr>\n<tr>\n<td style=\"border: solid .5pt black\" align=\"center\"><code>\"<\/code><\/td>\n<td style=\"border: solid .5pt black\" align=\"center\"><code>.*?<\/code><\/td>\n<td style=\"border: solid .5pt black\" align=\"center\"><code>\"<\/code><\/td>\n<td style=\"border: solid .5pt black\" align=\"center\"><code>&gt;<\/code><\/td>\n<\/tr>\n<\/table>\n<p>  Notice that here, the .*? pattern matched the inner quotation mark  because that was the only way to make the pattern match successfully.  (&#8220;I wouldn&#8217;t have done it, but you forced me!&#8221;)\n  <a href=\"http:\/\/blogs.msdn.com\/ericgu\/archive\/2004\/02\/26\/80465.aspx\">Even smart people  make this mistake<\/a>.\n  If you really don&#8217;t want quotation marks to match the .*? then you  need to say so.\n<code>\"[^\"]*\"&gt;<\/code>\n  This means &#8220;Match a quotation mark, then zero or more characters that  aren&#8217;t quotation marks, then another quotation mark, and then a greater-than.&#8221;<\/p>\n<p>  [Raymond is currently on vacation; this message was pre-recorded.]  <\/p>\n","protected":false},"excerpt":{"rendered":"<p>The regular expression *? operator means &#8220;Match as few characters as necessary to make this pattern succeed.&#8221; But look at what happens when you mix it up a bit: &#8220;.*?&#8221; This pattern matches a quoted string containing no embedded quotes. This works because the first quotation mark starts the string, the .*? gobbles up everything [&hellip;]<\/p>\n","protected":false},"author":1069,"featured_media":111744,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1],"tags":[25],"class_list":["post-40073","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-oldnewthing","tag-code"],"acf":[],"blog_post_summary":"<p>The regular expression *? operator means &#8220;Match as few characters as necessary to make this pattern succeed.&#8221; But look at what happens when you mix it up a bit: &#8220;.*?&#8221; This pattern matches a quoted string containing no embedded quotes. This works because the first quotation mark starts the string, the .*? gobbles up everything [&hellip;]<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts\/40073","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/users\/1069"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/comments?post=40073"}],"version-history":[{"count":0,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts\/40073\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/media\/111744"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/media?parent=40073"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/categories?post=40073"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/tags?post=40073"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}