Generative AI integration is the architecture connecting a model to the systems an organization already runs. We build these connections at Xavor Corporation in Irvine, California through our generative AI development services. Three mechanisms carry enterprise context to a model, and each answers a different kind of question. Each decision below carries a cost on both sides. Each enterprise system already answers who may see what, and the integration either asks it or overrides it.
A pilot works because one engineer has access to everything, and production fails because permissions suddenly apply.
RAG, tool calls, or MCP: how context reaches the model
Three mechanisms carry enterprise context to a model, and each answers a different kind of question. Retrieval searches documents by meaning. Tool calls query systems by parameter. The Model Context Protocol standardizes how tools are exposed so one interface reaches many of them.
| Retrieval | Tool calls | MCP | |
| Fetches | Unstructured documents | Structured records | Either, through a standard |
| Freshness | As fresh as the index | Live at query time | Live at query time |
| Can write | No | Yes | Yes |
| Matching | Semantic similarity | Exact parameters | Exact parameters |
| Suits | Knowledge questions | Transactional systems | Many tools, one interface |
Retrieval answers what a document says, and a tool call answers what a record holds right now.
The vocabulary differs by provider, which trips up teams reading two sets of documentation. OpenAI uses function calling and tool calling interchangeably. Anthropic says tool use and separates client tools, which run in your application, from server tools, which run on the provider’s infrastructure.
The distinction that survives both is where execution happens. Retrieval runs before the model generates. A tool call runs during, and your code answers it.
Production systems increasingly run these together in a loop, with the model deciding when it has enough context. That pattern is where the field has moved rather than a settled standard, and the terminology around it is still shifting.
Choosing retrieval over fine-tuning is a separate question, and we cover when retrieval beats fine-tuning for adapting a model there.
The seven decisions that decide what ships
Each decision below carries a cost on both sides, and each one has broken a deployment somewhere. None has a universally correct answer. What separates a shipped system from a stalled pilot is making all seven deliberately rather than inheriting them.

Published guidance names these as best practices, which hides the fact that each one costs you something.
Group one: how context arrives
Two decisions govern what the model sees before it generates anything at all.
1. Retrieval, tool calls, or a loop. Retrieval is cheap and predictable, and stale between index refreshes. Tool calls are current, slower, and harder to constrain. A loop costs the most per query.
Breaks when: retrieval is chosen for a question needing live data, so the answer was correct last Tuesday.
2. Raw systems or a governed layer. Reading straight from source tables is available today. Reading from agreed definitions takes longer and is the only version where two questions about the same measure agree.
Breaks when: the model derives its own definitions per query and revenue means three things. The definitions themselves are a separate build, covered in what a governed semantic layer holds.
Group two: what the model is allowed to do
Three decisions determine what the model can see and change once it runs. These are the ones that stop pilots at security review.
3. Service account or user identity. One service account ships in a week. Propagating the end user’s identity through every call is harder to build and is the only version that survives scrutiny.
Breaks when: an index built under a service account returns documents the asking user was never cleared to see. This is the most common reason pilots do not reach production.
4. Read-only or write-capable. Answering questions is safe and caps the value. Changing records delivers the automation case and creates a blast radius.
Breaks when: write access is granted broadly rather than per operation, so one malformed call updates a thousand records.
5. Where permissions are enforced. Filtering at retrieval is fastest and needs permission metadata in the index. Enforcing at the tool reuses the system’s own rules. Checking after generation is weakest.
Microsoft’s enterprise RAG accelerator applies permission trimming at retrieval across blended sources, on a Zero-Trust architecture. That is a reference architecture enforcing at the earliest point available.
Breaks when: the check runs last, so sensitive content reached the model even if it never reached the user.
Group three: how it runs and how it fails
Two decisions cover timing and failure, and teams frequently make neither one deliberately.
6. Synchronous or event-driven. Request and response is simple and bounded. Reacting to a record change covers the workflows where value concentrates, and brings queueing, retries, and ordering.
Breaks when: a synchronous design is retrofitted to a workflow that needed to fire on a state change.
7. Review, thresholds, or autonomy. Human review on every action is safe and removes the throughput case. Confidence thresholds need a signal worth trusting. Autonomy needs a rollback path.
Breaks when: no failure behavior is defined, so the first wrong answer becomes an incident rather than an exception. Designing that behavior is part of how we approach agentic AI development.
What changes when the system is a CRM, an ERP, or a PLM
Each enterprise system already answers who may see what, and the integration either asks it or overrides it. The seven decisions do not resolve the same way across four systems, which is why answering them once globally produces a design that fails in three places.
| System | The permission model | What it does to the integration |
| CRM | Record-level sharing rules | Two reps see different pipelines |
| ITSM | Approval chains and access governance | A writing agent must respect them |
| PLM | Item and change-level control | Unreleased parts are a compliance question |
| Knowledge stores | Folder-inherited document permissions | Nobody has audited them in years |
In Salesforce, sharing rules mean the same query returns different records depending on who asks. An agent carrying a service account flattens that, and the flattening is invisible until someone notices.
ServiceNow already encodes approval chains. An agent with write access either routes through them or creates a second, unaudited path to the same outcome.
PLM is stricter again. In Aras Innovator and Oracle Agile PLM, visibility of an unreleased item is a controlled question, not a convenience setting.
Every one of these systems spent years encoding who may see what, and an integration can undo that in an afternoon.
Knowledge stores are the quietest risk, because folder permissions accumulate over years and rarely get reviewed. We built a private ChatGPT on governed supply chain knowledge where that inventory came first and the model came second.
From pilot to production: what generative AI integration takes
A pilot needs one API key, and production needs a list of systems and their permissions. The gap between them is rarely about model quality. It is about which systems are in scope and what each already enforces.

Three steps close it:
- Inventory the systems and what each already enforces: list every system the model will reach and record its access model before designing anything.
- Answer decisions 3, 4, and 5 per system: identity, write capability, and enforcement point frequently resolve differently across a CRM and a PLM.
- Define the failure behavior before the first write: decide what happens on a wrong answer while it is still cheap to decide.
The work between a working demo and a shipped system is almost entirely about permissions and failure behavior.
Most of that work is connective rather than generative, and it sits closer to the integration work that connects enterprise systems than to model selection.
Decide before you build
Seven decisions made deliberately beat seven decisions inherited from whatever the pilot happened to do.
The pilot answered all seven by default. One service account, read-only, whatever data the engineer could reach. Those defaults are what production has to undo, and undoing them later costs more than choosing them now.
A pilot needs one API key. Production needs a list of systems, the permissions attached to each, and a decision on all seven of the above. If you want that list built against your own estate, [email protected] reaches our integration team.
FAQs
Generative AI integration connects a model to the systems an organization already runs, covering how context reaches it, what it can change, and whose permissions it operates under. It differs from a standalone AI application by operating inside existing workflows and access controls.
Decide how context arrives, what the model reads from, whose identity it carries, what it can change, where permissions are enforced, when it runs, and what happens when it is wrong. Answer those per system rather than once globally.
Retrieval searches unstructured documents by semantic similarity and returns text for grounding. API integration queries systems by exact parameter and can write as well as read. Retrieval suits knowledge questions; API calls suit transactional systems.