{"id":41953,"date":"2003-11-05T03:20:00","date_gmt":"2003-11-05T03:20:00","guid":{"rendered":"https:\/\/blogs.msdn.microsoft.com\/oldnewthing\/2003\/11\/05\/an-anecdote-about-improper-capitalization\/"},"modified":"2003-11-05T03:20:00","modified_gmt":"2003-11-05T03:20:00","slug":"an-anecdote-about-improper-capitalization","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/oldnewthing\/20031105-00\/?p=41953","title":{"rendered":"An anecdote about improper capitalization"},"content":{"rendered":"\n<p>         I&#8217;ve already discussed <a href=\"http:\/\/blogs.gotdotnet.com\/raymondc\/PermaLink.aspx\/3cfdaedf-2060-4697-b7bf-19205d8448aa\">some         of the strange consequences of case-sensitive comparisons<\/a>.      <\/p>\n<p>         <a href=\"http:\/\/www.eightypercent.net\/Archive\/2003\/10\/14.html\">Joe Beda mentioned         the Internet Explorer capitalization bug that transformed somebody&#8217;s name into a dead         body<\/a>. Allow me to elaborate. You might learn something.      <\/p>\n<p>         This bug occurred because Internet Explorer tried to capitalize the characters in         the name &#8220;Yamada&#8221; but was not mindful of the character-combining rules of the double-byte         932 character set used for Japanese. In this character set, a single glyph can be         represented either by one or two bytes. The Roman character &#8220;A&#8221; is represented by         the single byte 0x41. On the other hand, the characters &#8220;&#12398;&#8221; is represented         by the two bytes 0x82 0xCC. (You will need to have Japanese fonts installed to see         the &#8220;no&#8221; character properly.)      <\/p>\n<p>         When you parse a Japanese string in this character set, you need to maintain state.         If you see a byte that is marked as a &#8220;DBCS lead byte&#8221;, then it and the byte following         must be treated as a single unit. There is no relationship between the character represented         by 0xE8 0x41 (&#37666;) and 0xE8 0x61 (&#37942;) even though the second bytes happen         to be related when taken on their own (0x41 = &#8220;A&#8221; and 0x61 = &#8220;a&#8221;).      <\/p>\n<p>         Internet Explorer forgot this rule and merely inspected and capitalized each byte         independently. So when it came time to capitalize the characters making up the name         &#8220;Yamada&#8221;, the second bytes in the pairs were erroneously treated as if they were Roman         characters and &#8220;capitalized&#8221; accordingly. The result was that the name &#8220;Yamada&#8221; turned         into the characters meaning &#8220;corpse&#8221; and &#8220;field&#8221;. You can imagine how Mr. Yamada felt         about this.      <\/p>\n<p>         Converting the string to Unicode would have helped a little, since the Unicode capitalization         rules would certainly not have connected two unrelated characters in that way. But         there are still risks in character-by-character capitalization: In some languages,         capitalization is itself context-sensitive. <a href=\"http:\/\/www.microsoft.com\/globaldev\/getwr\/steps\/wrg_sort.mspx\">MSDN         gives as an example<\/a> that in Hungarian, &#8220;SC&#8221; and &#8220;Sc&#8221; are not the same thing when         compared case-insensitively.      <\/p>\n","protected":false},"excerpt":{"rendered":"<p>I&#8217;ve already discussed some of the strange consequences of case-sensitive comparisons. Joe Beda mentioned the Internet Explorer capitalization bug that transformed somebody&#8217;s name into a dead body. Allow me to elaborate. You might learn something. This bug occurred because Internet Explorer tried to capitalize the characters in the name &#8220;Yamada&#8221; but was not mindful of [&hellip;]<\/p>\n","protected":false},"author":1069,"featured_media":111744,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1],"tags":[26],"class_list":["post-41953","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-oldnewthing","tag-other"],"acf":[],"blog_post_summary":"<p>I&#8217;ve already discussed some of the strange consequences of case-sensitive comparisons. Joe Beda mentioned the Internet Explorer capitalization bug that transformed somebody&#8217;s name into a dead body. Allow me to elaborate. You might learn something. This bug occurred because Internet Explorer tried to capitalize the characters in the name &#8220;Yamada&#8221; but was not mindful of [&hellip;]<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts\/41953","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/users\/1069"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/comments?post=41953"}],"version-history":[{"count":0,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/posts\/41953\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/media\/111744"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/media?parent=41953"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/categories?post=41953"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/oldnewthing\/wp-json\/wp\/v2\/tags?post=41953"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}