Retrieval API plus Your Own LLM: The Middle Path to Ground a Custom Chatbot or App on SharePoint Content

Overall summary

The Retrieval API is the middle path. It sits between fully managed intelligence and a fully custom build. It gives you Microsoft’s permission-trimmed, freshness-aware grounding over your live SharePoint content, and you do not build or run a search index of your own. You bring the model, the orchestration, and the experience. Microsoft brings the retrieval, the security trimming, and the compliance boundary.

This makes it a good fit when you want targeted grounding, control over the model, and a predictable pipeline, without the cost and effort of running your own crawler and index. It is the wrong fit when you want Copilot-grade reasoning out of the box, when you want Microsoft to own the intelligence end to end, or when you need the full control of a custom Azure AI Search build.

Across the eight factors it rates Strong on five: Security, Governance, Data residency, Flexibility, and Data freshness. It rates Moderate on three: Cost, Retrieval quality, and Scalability.

Reach for the Retrieval API when:

  • You want permission-trimmed, live SharePoint grounding without building an index.
  • You want to choose your own model, including open-source or fine-tuned models.
  • You want targeted grounding over specific SharePoint sites or document sets.
  • You want retrieval behavior you can tune for your domain.
  • You want to avoid the work of crawling and indexing SharePoint yourself.

Look elsewhere when:

  • You want Copilot-grade reasoning and broad Microsoft 365 context. The Retrieval API does retrieval only.
  • You want Microsoft to own the intelligence. Here you must bring your own model.
  • You need the most control over ranking, chunking, and ingestion. That is the custom Azure AI Search path.
  • You need a fully delegated, user-level reasoning engine. That is Work IQ.

Bottom line: the Retrieval API is the balanced option. You keep control of the model and the experience. Microsoft keeps control of the security trimming and the freshness. You skip the work of building your own index, and you get targeted grounding that Work IQ does not give you. In return, you own the reasoning, and you pay to host and run your own model.

Introduction

In the baseline article, I said the three options differ mainly in one thing: how much of the pipeline you own. The Retrieval API plus your own LLM sits right in the middle. It is not fully managed intelligence, and it is not a fully custom build. It is a hybrid.

Microsoft handles the retrieval. You handle the reasoning.

The Retrieval API returns permission-trimmed, grounded content from the live Microsoft 365 index. Your model then turns that content into an answer. This split is the whole point of the middle path. You get secure, predictable grounding without building an index, and you keep full control over the model, the ranking logic, and the user experience.

In the rest of this article I put the Retrieval API through the same eight factors I used in the baseline and in the Work IQ article, so the comparison stays honest and consistent.

A few things to know first

  1. The Retrieval API is part of the Microsoft 365 Copilot API family and is available over REST.
  2. It returns grounded content, not answers. You must bring your own model.
  3. It trims content automatically based on the signed-in user’s Microsoft 365 identity.
  4. It does not need an index, embeddings, or chunking of your own.
  5. It supports targeted grounding over specific SharePoint sites or document sets, using a filter expression written in Keyword Query Language (KQL).
  6. It covers SharePoint, OneDrive for Business, and Microsoft 365 Copilot connectors as data sources.
  7. It is available in the commercial (Global) cloud only. It is not available in the US Government (GCC High, DoD) or China (21Vianet) clouds.

Security

What I am looking for: whether the solution honors SharePoint permissions, and how the user’s identity is carried through to the content.

Security is one of the Retrieval API’s strongest points. Like Work IQ, it is delegated only. Every call runs on behalf of the signed-in user and carries that user’s Microsoft Entra identity. The API returns only the content that user is already allowed to see. Sensitivity labels and access controls are honored automatically.

The difference from Work IQ is that the Retrieval API does not reason. It retrieves only. So your model never sees content the user is not entitled to, because the API trims it before your model ever receives it.

Authentication follows the same pattern as other Microsoft 365 delegated APIs. Your app signs the user in against Entra ID. The delegated scopes are Files.Read.All and Sites.Read.All for SharePoint and OneDrive, and ExternalItem.Read.All for Copilot connectors. Application permissions are not supported. Personal Microsoft accounts are not supported. So anonymous and unattended service-account scenarios are not possible.

One scope point is worth calling out. The API does not support Sites.Selected. So you cannot narrow the app’s consent down to a named set of SharePoint sites the way the Graph application model lets you. The app gets tenant-wide read scope. In practice the results are still trimmed to what the signed-in user can see, because the call is delegated. But the consent itself is broad. If you want to restrict the solution to specific sites, you do that at query time with a filter on path or site ID. Be clear on what that is: it is a query filter, not a permission boundary. It shapes what you ask for, not what the app is allowed to reach. If you need a true permission-level limit to specific sites, this is a point in favor of the custom build in Part 4.

There is a quiet trap in that filter. If your KQL has a syntax error, the call does not fail. It runs as if you passed no filter at all, across everything the user can see. This is not a permission breach, because results are still trimmed to the user. But it silently defeats the scoping you thought you had. For content where you meant to limit answers to an approved set of sites, that is a real risk. So validate and test your filter expressions, and do not assume a scoped call stayed scoped.

Bottom line: Strong. You get automatic permission trimming, sensitivity label enforcement, and delegated identity, and you write no access-control code. The limits are that anonymous and app-only scenarios are not supported, that site scoping is a query filter rather than a consent boundary, and that a broken filter fails open to the user’s full view.

Networking and security model

What I am looking for: whether your app needs private endpoints, an API gateway, or firewall changes to use this option.

In Part 2, I said the client app needs almost nothing on the network to reach Work IQ. The same is true for the Retrieval API, and for the same reason. It is a Microsoft Graph endpoint. You call https://graph.microsoft.com/v1.0/copilot/retrieval over HTTPS on port 443. The call is secured by a Microsoft Entra token and TLS. So your app needs outbound HTTPS to graph.microsoft.com, and nothing more. No private endpoint. No API gateway in front of it. No inbound firewall change. The endpoint is a public, multi-tenant Microsoft service, and there is no Private Link or VNet option to reach it privately. You rely on Entra and TLS, not on network isolation.

But here the middle path adds a second network surface that Work IQ does not have, and this is the honest difference. You host your own model. That model endpoint is yours. So its network posture is your decision and your responsibility. If you run your model on Azure OpenAI or your own inference service, you can and often should put it behind a private endpoint or a VNet, and lock down the egress from your orchestration layer so content flows only where you intend. With Work IQ, Microsoft owned the model, so this surface did not exist for you. Here it does.

Bottom line: reaching Microsoft is a simple outbound HTTPS call, with no private networking to set up. Securing your own model is on you, and it is where any real network and isolation work in this option lives.

Cost

What I am looking for: licensing, implementation, and operations.

Licensing and billing. There are two ways to pay. The first is per-user Microsoft 365 Copilot licenses. With those, the Retrieval API calls carry no extra per-call charge. The second is pay-as-you-go, which is in preview. With pay-as-you-go you can call the API for users who do not have a Copilot license, and you pay $0.10 per API call, billed through an Azure subscription. Two conditions apply to pay-as-you-go: the tenant still needs at least one Copilot license, and pay-as-you-go covers tenant-level sources (SharePoint and Copilot connectors) but not user-level OneDrive.

There is one more cost that is easy to miss. You host and run your own model. That model cost is separate from anything Microsoft charges, and it is usually the bigger number. So do not assume this path is cheaper than Work IQ. The per-call retrieval charge is small, but your model hosting and inference are on top. Whether the total is cheaper or pricier than Work IQ depends on your model choice and your volume.

Implementation. Moderate. You do not build an index, a crawler, or an ingestion pipeline. You do build: a model endpoint, a retrieval and orchestration layer, and the logic that turns the retrieved content into a grounded answer.

Operations. Moderate. You watch four things: your model hosting cost, your Copilot Credit use, Retrieval API throttling, and your model’s quality over time.

Bottom line: Moderate. You save the cost of building and running an index. You take on the cost of hosting and running your own model. The retrieval call itself is cheap. The model is where your bill grows.

Governance, including auditability and observability

What I am looking for: how well the approach fits enterprise controls, and whether you can see and prove what was retrieved.

Governance is strong for the retrieval half, because Retrieval API calls are Copilot API calls. They log in Microsoft Purview. Sensitivity labels, DLP, retention, and eDiscovery apply automatically. For the retrieval step, you get the same compliance posture as Work IQ.

The catch is that your model sits outside Microsoft’s governance boundary. So you have to govern it yourself. You must make sure your model does not store user content, that your hosting environment meets your compliance needs, and that your own logs do not capture sensitive data.

Bottom line: Strong for retrieval. You get Purview audit, eDiscovery, labels, and retention for the retrieved content. You govern your model separately.

Data residency

What I am looking for: where the content lives and travels, and whether it leaves the tenant boundary.

Here the middle path is actually cleaner than Work IQ, and for a simple reason. The Retrieval API does not run a model on Microsoft’s side. It returns grounded content only. So the retrieval and grounding stay inside Microsoft’s boundary, the semantic index honors the user’s identity, and there is no Microsoft-side model routing to worry about. Prompts, responses, and data accessed through Microsoft Graph are not used to train foundation models.

The residency question that remains is about your model, not Microsoft’s. You choose where your model runs. So the EU Data Boundary and any region rules for the reasoning step are your responsibility, because the reasoning happens on your side. Verify your model’s hosting region against your own requirements.

Bottom line: Strong. Retrieval stays inside Microsoft’s boundary and honors identity. Your model’s location is the one residency choice you own.

Retrieval quality and citations

What I am looking for: relevance, accuracy, control over ranking, and whether citations are trustworthy.

Retrieval quality is good, but not as strong as Work IQ. Work IQ uses Microsoft’s semantic index plus its own reasoning. The Retrieval API uses the same semantic index, but it returns raw grounded content without reasoning.

The strengths: permission-aware retrieval, live index freshness, real citations with source URLs, relevance scores on each extract, and targeted grounding over specific sites.

The limits: you own the ranking and how you use the results. You must merge multiple returned extracts into one coherent answer. And you must handle accuracy and hallucination control in your own model.

Two coverage limits matter for insurance content, because they decide what the index can actually read. First, the strong semantic retrieval only works for a short list of file types: .doc, .docx, .pptx, .pdf, .aspx, and .one. Every other file type falls back to lexical retrieval, which is keyword matching only. So content in other formats quietly loses the semantic quality that is the whole reason to use this index. Second, images, charts, and other non-text content are not read at all, and table text is read only from .doc, .docx, and .pptx files. Scanned documents, diagrams, and forms are common in Property and Casualty, so check how much of your knowledge actually lives in text your index can see.

One fair note so the comparison stays honest. These are semantic index limits, and Work IQ reads the same index. So they apply to Work IQ too. They set a floor both managed options share, and they do not change the ranking between them. The reason this option still rates below Work IQ on retrieval quality is the one above: you own the ranking and the answer synthesis.

Bottom line: Moderate. You get trustworthy citations and strong permission trimming. You build the ranking logic and the answer synthesis yourself, and you should confirm your content is in formats the index can read well.

Flexibility

What I am looking for: how much control you keep over the model, the orchestration, and the user experience.

Flexibility is where the Retrieval API shines. You keep the model choice, the model tuning, the ranking logic, the prompt design, and full control over the user experience.

You can use OpenAI models, Azure OpenAI models, open-source models, or a fine-tuned model of your own. You can build domain-specific ranking, custom multi-step reasoning, and your own retrieval-augmented pipeline on top.

You can also target the grounding. You point retrieval at specific SharePoint sites, libraries, or document sets using a KQL filter expression on properties like path and site ID. This is the exact control Work IQ does not give you at the raw API level, where a call runs across everything the user can see. Keep in mind this is a query-time filter, not a permission limit, as noted under Security.

One piece of the merge job sits with you. A single call reads from one data source only, set by the dataSource parameter, and the API does not interleave results across sources. So if an answer needs SharePoint and OneDrive and connectors together, you make separate calls and merge the results yourself. That is the middle path being honest about where the work lives: the flexibility is real, and so is the orchestration it hands you.

Bottom line: Strong. You keep control of the model, the ranking, the orchestration, and the experience. This is the Retrieval API’s defining strength, and you pay for it in the orchestration you own.

Scalability

What I am looking for: the ability to handle large volumes of content and many sites without hitting throttling limits.

Content scale is effectively unlimited, because Microsoft owns and runs the index. You never crawl. You never hit Graph crawl limits. You get Microsoft’s index, which already runs at cloud scale, on day one.

Query scale is the part you manage, and there is a concrete ceiling to design around: up to 200 requests per user per hour. For a busy user, that limit arrives sooner than you expect. So you should handle 429 responses with Retry-After, and 503 responses with backoff. The API also lets you batch up to 20 requests together, and return between 1 and 25 results per call, which helps you keep call volume under the limit.

Bottom line: Moderate. Content scale is unlimited. Query scale is capped per user per hour, so plan for it, and the rest depends on your model and your budget.

Data freshness and latency

What I am looking for: how quickly a change in SharePoint shows up in answers, and how fast answers return.

Freshness matches Work IQ, because both use the same semantic index. Edits to existing documents show up within minutes. Brand-new documents can take up to about a day to reach the semantic layer.

Latency has two parts. There is the Retrieval API response time, which Microsoft controls. And there is your model’s inference time plus your own orchestration, which you control.

Bottom line: Strong. You get Microsoft’s freshness for free. You control your model’s latency.

Other limits worth knowing

A few smaller limits round out the picture. The query string is capped at 1,500 characters. A single call returns at most 25 results. Files larger than 512 MB (for .docx, .pptx, and .pdf) or 150 MB (for anything else) are skipped. And the whole API carries the same limits as the Microsoft 365 Copilot semantic index, since that is the index it reads. None of these change the verdict, but they are the kind of thing you want to know before you build, not after.

Scorecard

FactorRetrieval API plus your own LLMIn one line
SecurityStrongPermission trimming and labels enforced automatically.
CostModerateNo index to build, but you host and run the model.
GovernanceStrongRetrieval governed by Purview; the model is yours to govern.
Data residencyStrongRetrieval stays in Microsoft’s boundary; your model’s region is your call.
Retrieval qualityModerateStrong citations, but ranking and synthesis are yours; drops toward Limited for scanned, image, or unsupported-format content.
FlexibilityStrongFull control over model, ranking, and experience.
ScalabilityModerateUnlimited content; query scale depends on your model.
Data freshnessStrongEdits show in minutes; new documents within about a day.

Bottom line

The Retrieval API is the middle path for a reason. Microsoft keeps the retrieval, the security trimming, and the freshness. You keep the model, the ranking, and the experience. You skip the cost of building an index, and you take on the cost of running a model.

Choose it when you want Microsoft’s secure, live grounding but your own model and your own control. Move past it when you want Microsoft to own the intelligence, or when you need the full control of a custom build.

In the next and final article, I take the most custom option: a custom RAG solution built on Azure AI Search, where you own the whole pipeline from ingestion to answer.


Written by Venugopal Reddy P. I write about governance, compliance, AI, and automation at AK Tech Hub. Views are my own, and product details in this fast-moving space change often, so verify current capabilities and pricing before making decisions.

2 responses to “Retrieval API plus Your Own LLM: The Middle Path to Ground a Custom Chatbot or App on SharePoint Content”

  1. […] Copilot Retrieval API plus your own LLM. The middle path. Microsoft returns permission-trimmed, grounded content from the live index, and you bring your own model and orchestration to turn it into an answer. You keep control of the model and the experience without ever having to build or maintain an index of your own. […]

    Like

  2. […] the next article I take the second option, the Copilot Retrieval API paired with your own LLM. It trades some of that managed convenience back for control. Microsoft still returns […]

    Like

Leave a comment