{"id":2933,"date":"2026-09-29T14:30:42","date_gmt":"2026-09-29T21:30:42","guid":{"rendered":"https:\/\/devblogs.microsoft.com\/foundry\/?p=2933"},"modified":"2026-09-30T13:32:32","modified_gmt":"2026-09-30T20:32:32","slug":"why-content-extraction-still-matters-in-the-genai-era","status":"publish","type":"post","link":"https:\/\/devblogs.microsoft.com\/foundry\/why-content-extraction-still-matters-in-the-genai-era\/","title":{"rendered":"Why content extraction still matters in the GenAI era"},"content":{"rendered":"<p>&nbsp;<\/p>\n<p style=\"text-align: left;\"><strong>The future of enterprise AI is not determined by the model you choose. It is <\/strong><strong>determined by whether your agents can trust the content they act on.<\/strong><\/p>\n<p>Five years into the era of large language models, that thesis still holds. If\nanything, it is truer today than when we started. The models are stronger. The\ncontent is messier. Most of the real engineering lies in turning a raw PDF,\nimage, or audio file into grounded, verifiable information an agent can act on.<\/p>\n<p>Over the next few weeks, we will share how the Azure AI team approaches this\nengineering layer: why it remains necessary, what it takes to operate at scale,\nand how advances in document AI and foundation models have changed what is\npossible.<\/p>\n<p><span class=\"TextRun SCXW11987118 BCX8\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW11987118 BCX8\">This opening post starts with a question we hear often: if models can already read and reason over content, why do we still need a dedicated extraction <\/span><span class=\"NormalTextRun SCXW11987118 BCX8\">layer?<\/span><\/span><span class=\"EOP Selected SCXW11987118 BCX8\" data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<h2>Why this question keeps coming back<\/h2>\n<p>A recurring assumption is that smarter models eliminate the need for content\nextraction. The opposite is closer to the truth: the more enterprises depend on\nAI agents, the more the underlying content has to be trustworthy, structured,\nand auditable. Otherwise, every agent decision inherits the limitations and\nambiguities of the source material. Better models raise the ceiling on what is\npossible; they do not remove the floor of what has to be true.<\/p>\n<p>As organizations move LLM-based prototypes to production, a familiar pattern\nemerges. A team wires up a chat interface over its document repository and gets\npromising results. As the collection grows from hundreds to millions of pages,\nnew requirements emerge: predictable costs, consistent latency, reliable\nhandling of complex content, and answers that can be traced to their source.\nThe audit team asks a question a model alone cannot reliably answer:\n<em>Where does that answer come from, and how confident are we?<\/em><\/p>\n<p>The tempting conclusion is &#8220;we need a smarter model.&#8221; The harder truth is that\ncontent extraction is its own engineering discipline: turning raw bytes into\nstructured, grounded, and verifiable inputs. It is also the layer most\nproduction generative AI systems still get wrong.<\/p>\n<h2>What it takes to build extraction directly on an LLM<\/h2>\n<p><span class=\"TrackChangeTextInsertion TrackedChange SCXW96247131 BCX8\"><span class=\"TextRun SCXW96247131 BCX8\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW96247131 BCX8\">C<\/span><\/span><\/span><span class=\"TextRun SCXW96247131 BCX8\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW96247131 BCX8\">onsider what happens when a team builds extraction directly on an LLM. The first prompt is often the easy part; the work expands as real-world <\/span><span class=\"NormalTextRun ContextualSpellingAndGrammarErrorV2Themed SCXW96247131 BCX8\">content<\/span><span class=\"NormalTextRun SCXW96247131 BCX8\"> and production requirements enter the picture.<\/span><\/span><\/p>\n<p>You start with prompting, which is good enough for the first twenty documents.\nThen you encounter PDFs, TIFFs, DOCX files, scans, and other format variations,\neach with different parsing behaviors. Page-count and context-window limits\nforce you to build a chunker and decide how to handle image-heavy files:\nwhether to process embedded images separately or render every page as an image\nwithout driving up token usage.<\/p>\n<p>When the model does not preserve table structure, you add a layout parser.\nPrompt engineering expands into a growing library of instructions and\nexceptions. Multi-page relationships, such as a total on page 12 that refers to\na line item on page 4, need another pass.<\/p>\n<p>Then come per-field confidence and grounding for audits, normalization for\ndates, currencies, and party names, and error handling for pages the model\ncannot process reliably. Before long, you also own token optimization,\nevaluation infrastructure, reliability, scalability, and the security and\ncompliance controls required to run the system in production.<\/p>\n<p>Each challenge is manageable in isolation. Together, they amount to a platform\nyour team must own, operate, and benchmark again whenever a new model ships.<\/p>\n<p>Building it yourself remains a valid choice when your problem is narrow, your\ndocument formats are stable, and you are comfortable owning the entire system.\nModern LLMs and coding agents make that easier than ever. The gap between a\nworking extraction <em>demo<\/em> and an enterprise-grade extraction <em>platform<\/em> is\nstill substantial. Managed extraction services address the concerns that\nemerge after the prototype: scale, reliability, grounding, confidence,\ngovernance, compliance, security, cost, and ongoing model innovation without\nrepeated integration work.<\/p>\n<h2>One portfolio, two complementary approaches<\/h2>\n<p>Our work in this area predates the generative AI wave. Azure Form Recognizer,\nnow <a href=\"https:\/\/learn.microsoft.com\/azure\/ai-services\/document-intelligence\/overview?view=doc-intel-4.0.0\">Azure Document Intelligence in Foundry Tools<\/a>, emerged when cloud-based document processing was still a relatively new category. Our patented Custom Template technology used random forests to learn repeatable document structures from a small set of labeled examples, giving\ndevelopers a practical way to automate extraction from their own forms.<\/p>\n<p>The next major step was Custom Neural in Azure Document Intelligence. This\nevolution was enabled by advances in multimodal document understanding,\nexemplified by <a href=\"https:\/\/www.microsoft.com\/en-us\/research\/publication\/layoutxlm-multimodal-pre-training-for-multilingual-visually-rich-document-understanding\/\">LayoutXLM<\/a>, pioneered by Microsoft Research Asia. LayoutXLM jointly models text, layout, and visual information to understand visually rich documents across languages.\nThese advances made it possible to move beyond repeatable templates and better\ngeneralize across variations in document structure and appearance. Together\nwith Read, Layout, and prebuilt models, they helped establish Azure Document\nIntelligence as a mature platform for high-accuracy, purpose-built document\nprocessing.<\/p>\n<p>The rapid advancement of foundation models opened the next frontier. Large\nlanguage models introduced broad knowledge and reasoning capabilities, creating\nan opportunity to combine high-quality content extraction with generative AI\nand address problems that traditional document-processing models alone could\nnot. <a href=\"https:\/\/learn.microsoft.com\/en-us\/azure\/ai-services\/content-understanding\/overview\">Azure Content Understanding in Foundry Tools<\/a> extends this evolution beyond documents, combining content extraction with generative analysis and reasoning across documents, images,\naudio, and video. It transforms unstructured multimodal content into structured, user-defined outputs for automation, analytics, search, and agentic workflows.<\/p>\n<p>The progression from Azure Document Intelligence to Azure Content Understanding reflects a consistent theme: pioneering new ways to turn unstructured content into structured, actionable information as AI technology evolves. The journey has moved from learning repeatable document templates, to understanding document structure and variation, to reasoning over content across modalities.<\/p>\n<p>Today, Azure Document Intelligence and Azure Content Understanding share parts of this technical foundation, but use different approaches and are optimized for different scenarios:<\/p>\n<ul>\n<li><strong>Azure Document Intelligence<\/strong> provides purpose-trained, structured document\nextraction. It is a mature choice for high-accuracy extraction from known\ndocument types and structures, including tax forms, identity documents,\nreceipts, invoices, and other document-processing scenarios.<\/li>\n<li><strong>Azure Content Understanding<\/strong> builds on high-quality content extraction and\nadds generative AI for schema-based field analysis, multimodal content\nprocessing, and reasoning. It works across documents, images, audio, and\nvideo and supports scenarios ranging from structured extraction to search,\nanalytics, and agentic applications. Its analyzers can combine content\nextraction, contextualization, grounding and confidence signals, and Foundry\nmodels to produce structured outputs defined by the application.<\/li>\n<\/ul>\n<p>These complementary strengths are why both services exist, and why many\nproduction systems can benefit from using them together. Later in this series,\nwe will examine the underlying technologies in more depth and provide detailed\nguidance for choosing the right approach for different scenarios.<\/p>\n<h2>Where Azure Content Understanding is going<\/h2>\n<p><span class=\"TextRun SCXW199748696 BCX8\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW199748696 BCX8\">With those complementary roles <\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">established<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">, <\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">we\u2019ll<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\"> turn to where <\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">we\u2019re<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\"> taking <\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">A<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">z<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">u<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">r<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">e<\/span> <span class=\"NormalTextRun SCXW199748696 BCX8\">Conte<\/span><span class=\"NormalTextRun SCXW199748696 BCX8\">nt Understanding next: expanding what developers can do with varied content while addressing the demands of production.\u00a0<\/span><\/span><span class=\"EOP Selected SCXW199748696 BCX8\" data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<p>We measure Azure Content Understanding against five practical criteria:\n<strong>quality<\/strong>, <strong>cost<\/strong>, <strong>latency<\/strong>, <strong>predictability<\/strong>, and <strong>enterprise readiness<\/strong>.<\/p>\n<p>Across most <strong>unstructured and semi-structured<\/strong> content &#8211; messy layouts, mixed\nmodalities, and reasoning-heavy fields &#8211; Azure Content Understanding is already\ndelivering better outcomes than traditional extraction pipelines. Customers\npoint to higher quality on unstructured content, simpler authoring, native\nmultimodal support, per-field grounding and confidence, and the ability to\nadopt advances in foundation models without reintegrating every release. Azure\nContent Understanding also extends beyond documents to images, audio, video,\nand agentic workflows.<\/p>\n<p><span data-contrast=\"none\">We are equally candid about where we need to invest. Our direction is to bring the strengths of purpose-built document extraction and generative AI closer together. Azure Document Intelligence continues to excel in highly structured document scenarios, drawing on task-specific models that understand two-dimensional layouts and spatial relationships. As we advance Azure Content Understanding, we are building toward higher-quality and richer document understanding, while bringing down model cost and simplifying model selection for enterprise production.<\/span><\/p>\n<p><span data-contrast=\"none\">That direction shapes our priorities: higher extraction quality, simpler integration, and more predictable production experiences. The next posts will explore these investments across five areas:<\/span><span data-ccp-props=\"{}\">\u00a0<\/span><\/p>\n<ul>\n<li><strong>Advanced Contextualization for Prebuilt Analyzers<\/strong> improves quality and\ncost efficiency for new industry-specific turnkey solutions.<\/li>\n<li><strong>Agentic mode<\/strong> supports tool-using workflows that reason over multi-step\nproblems in complex documents.<\/li>\n<li><strong>Synchronous Read and Layout APIs<\/strong> return results without requiring callers\nto manage asynchronous polling, enabling faster, more responsive experiences\nfor supported inputs.<\/li>\n<li><strong>Platform and framework integrations<\/strong> integrate Azure Content Understanding with\nFoundry IQ, Microsoft Agent Framework, LangChain, MarkItDown, and the Azure\nContent Understanding CLI, so its output flows into the pipelines teams\nalready run.<\/li>\n<li><strong>AI governance and enterprise readiness<\/strong> center privacy, security, and\nresponsible AI. Grounding, confidence scores, and dynamic human-in-the-loop\nworkflows help teams identify and review unreliable model output.<\/li>\n<\/ul>\n<p>For current availability and limitations, see <a href=\"https:\/\/learn.microsoft.com\/azure\/ai-services\/content-understanding\/whats-new\">What&#8217;s new in Azure Content\nUnderstanding<\/a>.<\/p>\n<p>Together, these investments improve more than the extraction pipeline. They\nexpand the range of content-heavy workflows teams can automate, including some\nthat might not look like content extraction at first.<\/p>\n<h2>Less obvious scenarios Azure Content Understanding unlocks<\/h2>\n<p>That broader potential becomes clearer in concrete examples. Here are eleven\nscenarios to watch:<\/p>\n<ul>\n<li><strong>Custom translation grounded in document layout and domain terminology.<\/strong>\nMove beyond &#8220;translate this paragraph&#8221; to &#8220;translate this contract while\npreserving numbered clauses, defined terms, and the confidence a legal\nreviewer needs to sign off.&#8221;<\/li>\n<li><strong>Multimodal search and knowledge mining.<\/strong> Turn images, video frames, and\naudio transcripts into the same queryable, grounded field structure as text,\nsupporting a single retrieval index across mixed-media archives.<\/li>\n<li><strong>Insights from charts, figures, and tables.<\/strong> Convert financial filings,\nscientific reports, and engineering drawings into structured data an agent\ncan reason over.<\/li>\n<li><strong>Audit-grade extraction with per-field grounding and confidence.<\/strong> Support\nregulated workflows such as insurance claims, tax filings, and healthcare\nintake, where the questions are not only &#8220;What does it say?&#8221; but also &#8220;How do\nwe know, and how sure are we?&#8221;<\/li>\n<li><strong>Security monitoring across short video clips and images.<\/strong> Detect and\nsummarize relevant events across visual evidence, producing structured\nfindings that security teams can review and act on.<\/li>\n<li><strong>Insurance claim triage.<\/strong> Combine forms, photos, estimates, and supporting\ndocuments to identify missing information, extract key facts, and route\nclaims for review.<\/li>\n<li><strong>Mortgage and loan underwriting.<\/strong> Cross-check income, employment, assets,\nand liabilities across application packages to surface inconsistencies and\nsupport verification.<\/li>\n<li><strong>Manufacturing incident analysis.<\/strong> Correlate incident reports, images,\nvideo, maintenance records, and sensor context to reconstruct events and\nhighlight contributing factors.<\/li>\n<li><strong>Vendor onboarding.<\/strong> Extract and validate information across applications,\ntax forms, certifications, and compliance documents to speed setup and\nexception handling.<\/li>\n<li><strong>Contract review.<\/strong> Identify key terms, obligations, deviations, and risks\nwhile grounding each finding in the source language.<\/li>\n<li><strong>Call centers.<\/strong> Turn conversation audio into structured insights that flow\ninto supervisor, compliance, and coaching workflows.<\/li>\n<\/ul>\n<p><span class=\"TextRun SCXW39828456 BCX8\" lang=\"EN-US\" xml:lang=\"EN-US\" data-contrast=\"none\"><span class=\"NormalTextRun SCXW39828456 BCX8\" data-ccp-parastyle=\"List Bullet\">Across these scenarios, the common requirement is not simply to read content. It is to give<\/span> <span class=\"NormalTextRun SCXW39828456 BCX8\" data-ccp-parastyle=\"List Bullet\">applications and agents information they can use &#8211; and people a way to verify it. That is why content extraction <\/span><span class=\"NormalTextRun SCXW39828456 BCX8\" data-ccp-parastyle=\"List Bullet\">remains<\/span><span class=\"NormalTextRun SCXW39828456 BCX8\" data-ccp-parastyle=\"List Bullet\"> foundational, even as models become more capable.<\/span><\/span><\/p>\n<h2>Where to go next<\/h2>\n<ul>\n<li>Explore the <a href=\"https:\/\/learn.microsoft.com\/azure\/ai-services\/content-understanding\/overview\">Azure Content Understanding overview<\/a>,\nthen try a prebuilt analyzer in the <a href=\"https:\/\/ai.azure.com\/\">Microsoft Foundry portal<\/a>\nor build a custom analyzer in <a href=\"https:\/\/contentunderstanding.ai.azure.com\/home\">Content Understanding Studio<\/a>.<\/li>\n<li>Review the <a href=\"https:\/\/learn.microsoft.com\/azure\/ai-services\/document-intelligence\/overview?view=doc-intel-4.0.0\">Azure Document Intelligence overview<\/a>\nand experiment in <a href=\"https:\/\/documentintelligence.ai.azure.com\/studio\">Document Intelligence Studio<\/a>.<\/li>\n<li>Follow the <a href=\"https:\/\/devblogs.microsoft.com\/foundry\/feed\/\">Microsoft Foundry Blog RSS feed<\/a>\nfor the rest of the series.<\/li>\n<\/ul>\n<p>We would love your feedback in the comments. Next week, we will take a closer\nlook at Azure Content Understanding and Azure Document Intelligence, including how they\ncompare with using LLMs directly.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Learn why stronger AI models make trustworthy content extraction more important, how Azure Content Understanding and Azure Document Intelligence serve different needs, and what comes next.<\/p>\n","protected":false},"author":218493,"featured_media":2941,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1],"tags":[6,144,162,4,185],"class_list":["post-2933","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-microsoft-foundry","tag-azure-ai","tag-content-understanding","tag-document-intelligence","tag-generative-ai","tag-multimodal-ai"],"acf":[],"blog_post_summary":"<p>Learn why stronger AI models make trustworthy content extraction more important, how Azure Content Understanding and Azure Document Intelligence serve different needs, and what comes next.<\/p>\n","_links":{"self":[{"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/posts\/2933","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/users\/218493"}],"replies":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/comments?post=2933"}],"version-history":[{"count":1,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/posts\/2933\/revisions"}],"predecessor-version":[{"id":2939,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/posts\/2933\/revisions\/2939"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/media\/2941"}],"wp:attachment":[{"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/media?parent=2933"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/categories?post=2933"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/devblogs.microsoft.com\/foundry\/wp-json\/wp\/v2\/tags?post=2933"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}