Foundry IQ: Inside the Managed Knowledge Layer That Turns RAG Into an Agent Tool Call
DEV Community

Foundry IQ: Inside the Managed Knowledge Layer That Turns RAG Into an Agent Tool Call

Foundry IQ: Inside the Managed Knowledge Layer That Turns RAG Into an Agent Tool Call Ask any team that shipped a "chat with your docs" bot in 2024 what happened six months later, and you'll hear a familiar story: the chunking strategy needed retuning, the reranker was hand-rolled and brittle, permissions leaked because the vector index didn't respect SharePoint ACLs, and every new agent needed its own copy-pasted retrieval pipeline. Retrieval-Augmented Generation (RAG) was never really the hard part - productionizing RAG as a shared, governed, multi-tenant capability was. Microsoft Foundry's answer to that problem is Foundry IQ, a managed knowledge layer released as the productized wrapper around Azure AI Search's agentic retrieval engine. It is one of the more quietly significant additions to the Foundry ecosystem this year, because it changes the unit of reuse in enterprise AI from "a RAG pipeline I built" to "a knowledge base I connect to N agents," with permission enforcement baked into the query path instead of bolted on after retrieval. This article is a deep, implementation-level walkthrough of Foundry IQ: what it actually is, how the agentic retrieval pipeline works internally, how it plugs into Foundry Agent Service over MCP, what the security model really enforces (and doesn't), and where it breaks down at scale. Table of Contents - Why This Exists - Core Concepts: Knowledge Sources, Knowledge Bases, Agentic Retrieval - Architecture: How a Query Actually Flows - The MCP Bridge Into Foundry Agent Service - Implementation Walkthrough - Permission Enforcement: What's Real and What's Marketing - Real-World Developer Scenario: HR Policy Assistant Across Three Data Silos - Production Considerations - Cost Considerations - Common Mistakes and Pitfalls - Alternatives and Trade-offs - Practical Recommendations - Conclusion - References 1. Why This Exists A Foundry model - even the largest ones you can deploy - has a knowledge cutoff and zero awareness of your tenant's SharePoint sites, blob containers, Fabric lakehouses, or internal wikis. The standard fix is RAG: chunk your documents, embed them, index them in a vector store, retrieve the top-k chunks at query time, and stuff them into the prompt. The problem isn't the concept - it's everything around it: - Query complexity. A single dense-vector similarity search handles "what is our parental leave policy" fine. It falls over on "compare our parental leave policy in the US and Germany and tell me which team's leads are most affected by the difference" - a query that actually needs decomposition into sub-questions. - Fragmentation across sources. Real enterprise knowledge is never in one place. It's in SharePoint, Blob Storage, a Fabric lakehouse, and sometimes it needs to come from the live web. Building one retrieval pipeline per source, per agent, doesn't scale organizationally. - Permission leakage. If your index doesn't carry ACL metadata and your query path doesn't check it, your RAG bot becomes a permission-escalation vector - the classic "I asked the HR bot and it told me the CEO's salary" failure mode. - Reuse. Ten different teams building ten different agents against the same underlying corporate knowledge shouldn't mean ten different embedding pipelines, ten different chunking strategies, and ten different bugs. Foundry IQ addresses this by promoting retrieval from "a pipeline you write" to "a first-class, shareable resource" - the knowledge base - that sits on top of Azure AI Search's agentic retrieval engine and is consumable by any number of Foundry agents (or Microsoft Agent Framework apps, or Copilot Studio agents) via a standard protocol. 2. Core Concepts: Knowledge Sources, Knowledge Bases, Agentic Retrieval Three objects matter here, and it's worth being precise about the layering because the docs use "Foundry IQ" and "agentic retrieval" almost interchangeably, which causes confusion. Knowledge source A knowledge source is a top-level Azure AI Search resource describing where content comes from and how it's queried. Knowledge sources are either: - Indexed - Azure AI Search ingests the content ahead of time via an indexer pipeline (chunking, embedding generation, metadata extraction). Supported indexed kinds: Search index (wraps an existing index), Azure Blob, Azure SQL (preview), File (preview), OneLake, and Indexed SharePoint (preview). - Remote - content is fetched live at query time, not pre-indexed. Supported remote kinds: Remote SharePoint (preview, uses the Copilot Retrieval API and enforces SharePoint's own permissions directly), Fabric Data Agent (preview), Fabric Ontology (preview), MCP server (preview - yes, a knowledge source can itself be a proxy to another MCP server), Work IQ (preview), and Web (via Bing). This indexed-vs-remote split matters architecturally: indexed sources trade freshness for query speed and semantic reranking quality; remote sources trade some query-time latency and reduced reranking control for zero duplication of source-of-truth data and native enforcement of the origin system's permission model. Knowledge base A knowledge base is the orchestration object. It references one or more knowledge sources and holds the parameters that control retrieval behavior - most importantly the retrieval reasoning effort (minimal , low , or medium ), which determines whether an LLM is used to plan/decompose the query before execution. Multiple agents can point at the same knowledge base. This is the reusable unit: build it once, govern it once, connect N agents to it. Agentic retrieval Agentic retrieval is the actual multi-query pipeline that a knowledge base executes when called. It is a genuinely distinct pattern from naive single-query vector search: - Query planning (skipped entirely at minimal effort): an LLM - an Azure OpenAI deployment you configure on the knowledge base - takes the user's query plus conversation history and decomposes it into a set of focused subqueries. This is where "compare parental leave in the US and Germany" becomes two or three separate, well-formed sub-questions instead of one blurry embedding. - Parallel query execution: every subquery runs concurrently against every configured knowledge source, using keyword, vector, or hybrid search as appropriate to each source. - Semantic reranking: each subquery's results are reranked with Azure AI Search's L2 semantic reranker to surface the truly relevant matches, not just the nearest-neighbor matches. - Result synthesis: everything is merged into a unified response. You always get merged extractive content; source references and an execution activity log are optional, and - in preview - full natural-language answer synthesis (an LLM writes the final grounded answer with citations, rather than the caller having to do that step itself). The important design decision here: agentic retrieval returns grounding data, not necessarily a final answer. Whether you consume it as raw extractive passages (GA path) or ask it to synthesize a natural-language answer (preview path) is your choice, made per knowledge base configuration. 3. Architecture: How a Query Actually Flows At a component level: | Component | Owning service | Role | |---|---|---| | Knowledge base | Azure AI Search | Orchestrates the pipeline; owns query parameters and reasoning effort | | Knowledge source(s) | Azure AI Search | Define what content is queried and how | | Search index | Azure AI Search | Backing store for indexed sources; holds text + vectors + semantic config | | Semantic ranker | Azure AI Search | L2 reranking of subquery results | | LLM | Azure OpenAI (via Foundry Models) | Powers query planning, web-result summarization, and answer synthesis | | Foundry Agent Service | Microsoft Foundry | Consumes the knowledge base as an MCP tool from a PromptAgentDefinition | Note what's not in this list: there is no separate "Foundry IQ service" runtime. Foundry IQ is the productized, governed front door - the naming and portal experience layer - over Azure AI Search's agentic retrieval, surfaced inside the Microsoft Foundry portal and consumable through Foundry Agent Service. This matters operationally: your quotas, region availability, and REST API versioning all live under Azure AI Search, not under a separate Foundry billing meter. Two API generations you must not mix As of the 2026-04-01 GA REST API, Azure AI Search supports agentic retrieval for GA knowledge source types with minimal reasoning effort only (i.e., no LLM-based query planning, extractive results only). The 2026-08-01-preview REST API version unlocks preview knowledge source types (SharePoint, Fabric, MCP-as-source, Web), non-minimal reasoning effort (LLM query planning), answer synthesis, and multi-turn message arrays. The Microsoft Foundry portal and Azure portal currently only expose the preview surface - meaning anything you wire up through the portal UI may need a deliberate migration pass before it's a supportable GA production configuration. If you're building for production today, decide explicitly which REST API version you're targeting rather than letting the portal default you into preview schemas you didn't intend to depend on. 4. The MCP Bridge Into Foundry Agent Service This is the part developers most need to internalize: Foundry Agent Service talks to a Foundry IQ knowledge base exclusively through the Model Context Protocol. The knowledge base itself exposes an MCP endpoint: {search_service_endpoint}/knowledgebases/{knowledge_base_name}/mcp?api-version=2026-08-01-preview That endpoint exposes exactly one MCP tool today: knowledge_base_retrieve . Your agent's PromptAgentDefinition gets an MCPTool pointed at that endpoint via a project connection - a RemoteTool connection category with ProjectManagedIdentity auth, which is specific to Foundry project connections and lets the project's system-assigned managed identity authenticate to Azure AI Search without you juggling API keys in agent config. This d

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.