π Deployment Guide
Use this page as the canonical installation guide. Start with Basic Deployment for a simple environment, or Zero Trust Deployment when network isolation is required.
Note: You can change parameter values in
main.parameters.jsonor set them withazd env setbefore runningazd provision. This applies only to parameters that support environment variable substitution.Underlying infrastructure: GPT-RAG provisions the Azure AI Landing Zone (AILZ) Bicep module as its infrastructure foundation. For the full list of parameters, opt-in features (IP allow-lists, BYO Private DNS / Log Analytics, hub-and-spoke integration, etc.), and the v2 migration path, see the AILZ parameterization reference and the v2-migration guide.
Prerequisites
Required Permissions:
- Azure subscription with Contributor and User Access Admin roles
- Agreement to Responsible AI terms for Azure AI Services
Required Tools:
- Azure Developer CLI
- PowerShell 7+ (Windows only)
- Git
- Python 3.12
Basic Deployment
Quick setup for demos without network isolation. In this mode, the workstation can run the full flow: provision, post-provision configuration, and service deployment.
azd init -t azure/gpt-rag
az login
azd auth login
azd env set NETWORK_ISOLATION false
azd provision
azd deploy
Add
--tenantforazor--tenant-idforazdif you want a specific tenant.Resource naming: Starting with GPT-RAG v3.1.0 and AI Landing Zone v2.2.0, fresh deployments name resources using the Cloud Adoption Framework pattern (for example
cosmos-<hash>-<env>-<region>-001). No extra variables are required. If you need to keep the pre-v3.1.0 names, setRESOURCE_NAMING_MODE=legacybeforeazd provision. See the resource naming guide for details, override options, and a before/after table.
azd provision runs GPT-RAG preflight checks before Azure Resource Manager deployment starts. These checks validate the selected region, jumpbox VM SKU restrictions, provider/location support for AI Search, Cosmos DB, Container Apps, and AI Foundry/Cognitive Services, and Azure OpenAI model quota for the configured deployments. If model quota is insufficient, the hook fails early and suggests candidate regions when possible.
Some transient Azure capacity failures are not exposed by reliable pre-create APIs. For example, Cosmos DB can still fail later with regional high-demand ServiceUnavailable; the preflight reports this limitation explicitly. Use GPT_RAG_REGIONAL_PREFLIGHT_SKIP=true only to bypass GPT-RAG regional checks, or PREFLIGHT_SKIP=true to bypass all preflight hooks.
For current published GPT-RAG umbrella releases, the postProvision hook runs
locally after azd provision, and azd deploy publishes the component services
pinned by manifest.json. Genuinely fresh deployments default to the
hosted/no-panel topology, in which the orchestrator runs as a Foundry hosted
agent instead of a Container App. The classic Container Apps topology, which
also deploys the orchestrator Container App, remains supported and is selected
explicitly through CHAT_BACKEND=orchestrator.
Chat runtime modes
Shipped in GPT-RAG v3.8.3
GPT-RAG v3.8.3
pins UI v2.6.2, orchestrator v4.1.1, ingestion v2.7.3, and AILZ
v2.5.1 at their exact release commits, and validates the hosted chat path
end to end. Classic, hosted/no-panel, and explicitly selected hosted-panel
are supported topologies. Continuity, user-history, owner-binding
validation, and operator-surface evidence gates remain deployment-published
false. See the
exact integration matrix.
The platform implementation resolves one canonical topology before provisioning and materializes the corresponding legacy flags and App Configuration values. Topology never changes automatically during a chat request.
| Environment or operator choice | Resolved settings | Resulting topology |
|---|---|---|
| Genuinely fresh environment | DEPLOYMENT_TOPOLOGY=hosted-no-panel, DEPLOY_HOSTED_AGENT_ORCHESTRATION=true, DEPLOY_ADMINISTRATIVE_PANEL=false, CHAT_BACKEND=hosted_agent |
Web UI and ingestion remain in Container Apps. Chat runs in a Microsoft Foundry hosted agent. No orchestrator Container App or panel-only Cosmos DB is provisioned. |
| Existing environment with persisted topology | Existing topology and CHAT_BACKEND stay sticky. An unmarked pre-cutover environment resolves to classic. |
Upgrade does not implicitly migrate identity, conversation, authorization, or cost semantics. |
| Explicit Container Apps fallback | DEPLOYMENT_TOPOLOGY=classic, materialized hosted and panel flags false, CHAT_BACKEND=orchestrator |
UI routes chat to the orchestrator Container App. Classic history and panel data remain available. |
| Explicit migration to hosted/no-panel | DEPLOYMENT_TOPOLOGY=hosted-no-panel, delegated hosted scope configured, panel false |
Runs the two-phase hosted lifecycle below, then validates the hosted request path before the classic chat path is removed or deactivated. |
| Explicit hosted panel | DEPLOYMENT_TOPOLOGY=hosted-panel, hosted and administrative-panel flags true, CHAT_BACKEND=hosted_agent |
Deploys UI and ingestion, omits the orchestrator Container App, and provisions only the owner-index and feedback metadata containers. User-history and operator routes remain off/503 because their evidence gates stay false. |
Hosted-panel is never selected implicitly. Operators must set
DEPLOYMENT_TOPOLOGY=hosted-panel or explicitly set both legacy hosted and
administrative-panel flags to true. A stray panel flag while hosted
orchestration is false remains classic.
The deployment hooks publish the shared runtime contract under the App
Configuration label gpt-rag:
| Setting | Operator contract |
|---|---|
DEPLOYMENT_TOPOLOGY |
Canonical deployment choice: hosted-no-panel, hosted-panel, or classic. Hosted-panel requires explicit selection. |
CHAT_BACKEND |
UI v2.6.2 treats missing or blank as hosted_agent; an umbrella deployment must publish the resolved sticky value. orchestrator is the explicit fallback. Unknown values fail startup. Environment configuration takes precedence over App Configuration. |
ORCHESTRATOR_BASE_URL |
Classic service root, used only when CHAT_BACKEND=orchestrator. The UI calls the /orchestrator route on this endpoint. |
HOSTED_AGENT_BASE_URL |
Required HTTPS hosted service root. Orchestrator v4.1.1 defines stateless POST /responses; UI v2.6.2 currently sends complete ordered messages through the distinct POST /invocations compatibility route. Continuity remains off until the live call route satisfies the protocol evidence gate. |
HOSTED_AGENT_RESOURCE_SCOPE |
Required explicit non-ARM hosted data-plane Entra scope ending in /.default, for example api://<application-id>/.default. |
HOSTED_AGENT_AUTH_MODE |
user_delegated is the default and required continuity path. Under OQ-OWN, it means the trusted UI BFF derives x-ms-user-identity; it does not mean an OBO token is sent to the agent. OBO remains a separate retrieval flow. service_identity is an explicit reviewed exception that is incompatible with owner-bound continuity, so continuity stays off/503 in that mode. |
HOSTED_AGENT_SSE_IDLE_TIMEOUT_SECONDS |
Finite positive wait for the next SSE event. The UI default is 60; an infinite timeout is rejected. |
HOSTED_AGENT_IMAGE_VERSION |
Canonical lowercase immutable digest in sha256:<64-hex-characters> form. Mutable tags are rejected. |
SEARCH_SERVICE_UAI_RESOURCE_ID |
Required identity boundary for private Search. Post-provisioning preserves an explicit value or resolves the single Search user-assigned identity from the Search resource; it must not publish an empty replacement. |
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT |
Generative-AI prompt and completion telemetry capture. Defaults to false. Set to true only when the deployment's data-handling policy explicitly permits sensitive content telemetry. |
Hosted configuration, authentication, connection, timeout, protocol, and
runtime failures are terminal for startup or the affected request. The UI does
not silently switch to the orchestrator or to managed identity. Foundry passes
opaque x-agent-foundry-call-id context to Toolbox; user and delegated bearer
tokens are not copied into tool payloads or client-defined identity headers.
Hosted conversation continuity platform gate
Delegated owner binding remains evidence-gated
HOSTED_CONTINUITY_ENABLED defaults to false and must stay false. The
exact component releases are pinned, but deployment must still prove the
live 100%-routed Responses protocol 2.0.0, trusted
UI BFF identity derivation, and the two exact direct agent-scoped roles
before recording
HOSTED_CONVERSATION_OWNER_BINDING_VALIDATED=true. Otherwise compatible
history endpoints return HTTP 503.
The trusted UI BFF derives x-ms-user-identity from the authenticated
server-side principal and sends it on the hosted Responses request. This owner
header is not an OBO token: OBO remains a separate downstream retrieval flow
with its own audience and bearer token.
Activation assigns only the UI BFF:
- built-in Foundry Agent Consumer
(
eed3b665-ab3a-47b6-8f48-c9382fb1dad6), whose only DataAction isMicrosoft.CognitiveServices/accounts/AIServices/endpoints/interact/action; and - custom GPT-RAG Hosted Agent User Identity Impersonation
(
bef66abe-a495-530a-be1d-5d882fecff03), containing only the reviewedMicrosoft.CognitiveServices/accounts/AIServices/agents/endpoints/UserIdentityImpersonation/actionDataAction.
Both assignments must be direct ServicePrincipal assignments scoped to the
individual hosted agent. Broader, inherited, group-derived, wildcard,
custom-equivalent, or extra-DataAction access fails validation. Foundry User
and Project Runtime User are prohibited substitutes. The hosted runtime is
not an identity-header source and receives no Conversation-capability key,
Conversation or impersonation RBAC, or Cosmos DB in hosted/no-panel.
| Setting or gate | Required posture |
|---|---|
HOSTED_CONTINUITY_ENABLED |
false until delegated owner binding validates; false/missing validation means history HTTP 503. |
HOSTED_CONVERSATION_OWNER_BINDING |
Defaults to delegated; capability is the only accepted explicit fallback. |
HOSTED_CONVERSATION_OWNER_BINDING_VALIDATED |
Becomes true only after the live protocol, identity-source, role-definition, assignment, and scope checks pass. |
HOSTED_CONVERSATION_DELEGATED_IDENTITY_HEADER |
Must be exactly x-ms-user-identity. |
HOSTED_CONVERSATION_DELEGATED_IDENTITY_SOURCE |
Must be exactly authenticated_ui_bff_principal. |
HOSTED_CONVERSATIONS_TOKEN_AUDIENCE |
Exact Foundry audience https://ai.azure.com; distinct from x-ms-user-identity and downstream OBO audiences. |
HOSTED_AGENT_RESPONSES_PROTOCOL_VERSION |
Must be exactly 2.0.0. |
HOSTED_CONTINUITY_UNAVAILABLE_STATUS_CODE |
Must be 503. |
HOSTED_HISTORY_MAX_ITEMS |
Default 100; accepted range 1-1,000. |
HOSTED_HISTORY_MAX_TOKENS |
Default 32000; accepted range 1-1,000,000. |
HOSTED_HISTORY_TRUNCATION |
Must be drop_oldest. |
Capability/HMAC is a disabled fallback only. The delegated primary path does
not create a capability key, require
HOSTED_CONTINUITY_KEY_VAULT_URI/HOSTED_CONTINUITY_KEY_VAULT_NAME, publish
HOSTED_CONVERSATION_CAPABILITY_KEY, or grant a capability-secret role.
Fallback key ID, TTL, vault, reference, and retained key-history behavior apply
only if a future release explicitly selects and validates capability mode. See
the
hosted conversation continuity platform contract
for the complete trust and rollout boundary.
POST /responses and POST /invocations are distinct protocols, not aliases,
and their request bodies are not interchangeable. Orchestrator v4.1.1
implements a stateless hosted Responses contract: callers send complete ordered
text history in input on every request. A non-empty string is valid for one
turn. An array must be non-empty, text-only, ordered oldest to newest, and end
with a non-empty user message. Top-level conversation and
previous_response_id are rejected with HTTP 422.
{
"input": [
{
"role": "user",
"content": "What is the document retention policy?"
},
{
"role": "assistant",
"content": "The policy states 30 days."
},
{
"role": "user",
"content": "Who approves an exception?"
}
],
"stream": true,
"store": false
}
The compatibility POST /invocations route retains the legacy messages-based schema:
{
"messages": [
{
"role": "user",
"content": "What is the document retention policy?"
}
],
"conversation_id": "<conversation-id>",
"metadata": {}
}
The hosted runtime constructs no managed-Conversations client and performs zero
managed state operations. The compatibility conversation_id is an opaque label
for tagging and local retrieval scoping only; it is not an ownership credential.
Managed history, user list/read/feedback/delete, and opaque handle validation are
owned by the trusted UI BFF. See the
hosted-agent release matrix.
Two-phase hosted deployment
For a fresh hosted/no-panel deployment, configure the delegated data-plane scope before the first provision:
azd env set HOSTED_AGENT_RESOURCE_SCOPE "api://<application-id>/.default"
azd env set HOSTED_AGENT_SSE_IDLE_TIMEOUT_SECONDS 60
# Keep continuity disabled until live protocol, identity, and role evidence validates ownership.
azd env set HOSTED_CONTINUITY_ENABLED false
azd provision
pwsh scripts/prepareHostedDeployment.ps1
azd provision
azd deploy
On POSIX systems, use scripts/prepareHostedDeployment.sh for the preparation
step. For an explicit migration, set
DEPLOYMENT_TOPOLOGY=hosted-no-panel before the first azd provision.
For a network-isolated VPN host, the
private-host workflow
adds explicit post-provision deferral and route/connectivity checks around
both provisions. Do not use the compact sequence above to skip those
checks or to bypass the published v3.8.4 host restriction.
For the supported hosted-panel topology, use the same two-phase flow with:
azd env set DEPLOYMENT_TOPOLOGY hosted-panel
azd env set HOSTED_CONTINUITY_ENABLED false
This composes UI and ingestion plus only the two panel metadata containers.
PANEL_HISTORY_ENABLED, PANEL_HISTORY_OWNER_BINDING_VALIDATED, and
PANEL_OPERATOR_SURFACES_ENABLED remain deployment-published false; do not
override them before their separate evidence and authorization procedures
complete. The corresponding routes return HTTP 503 while disabled.
Panel post-provisioning resolves exactly one managed-identity principal from the
frontend Container App and exactly one from ingestion. It then creates only
container-scoped Cosmos SQL grants on panel-conversation-owner-index and
panel-feedback: Cosmos DB Built-in Data Contributor for frontend and
Cosmos DB Built-in Data Reader for ingestion. Missing or ambiguous
Container App identities fail setup. Do not substitute account-scope grants,
grant ingestion write access, or grant the hosted agent any panel Cosmos role.
The first provision creates hosted prerequisites with image preparation
enabled but hosted deployment disabled. The preparation command clones and
verifies the manifest-pinned orchestrator source, builds the standard image and
the hosted-entrypoint derivative, resolves the pushed manifest to an immutable
digest, and persists that digest and source provenance. The second provision
materializes the digest-backed hosted handoff; azd deploy then deploys the
hosted agent.
Public deployments use shared ACR Tasks. Network-isolated deployments use the
dedicated VNet-connected ACR Tasks agent pool; shared ACR Tasks cannot reach a
private endpoint. Operators may pass an already-built immutable
sha256:<64-hex-characters> digest to the preparation command to skip builds.
No lifecycle hook recursively invokes azd provision.
The child hosted-agent/azure.yaml service definition is part of the prebuilt
handoff contract. It must declare language: docker and
docker.remoteBuild: true. The parent pre-deploy hook sets
AZD_AGENT_SKIP_ACR=true in the child azd environment before
azd deploy orchestrator-agent, so the already-prepared immutable image is used
instead of triggering another ACR build. Do not remove any of these three
settings from a prebuilt hosted deployment.
Capturing conversation content in Foundry telemetry
The Foundry portal's per-agent Conversations view reports
"No conversation turns found" unless the agent's GenAI spans carry the prompt
and completion text. Those spans (invoke_agent) are always emitted, but the
gen_ai.input.messages, gen_ai.output.messages, and
gen_ai.system_instructions attributes are only populated when message-content
capture is enabled. GPT-RAG leaves it disabled by default, because capturing
message content sends user prompts and model answers to Application Insights,
which is a data-privacy decision each deployment has to make for itself.
Opt in per environment before deploying the hosted agent:
azd env set HOSTED_AGENT_CAPTURE_MESSAGE_CONTENT true
azd deploy
The pre-deploy hook copies the root .azure environment into the
hosted-agent azd project, so a single azd env set at the root is enough. The
value is published to the hosted agent as
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT, declared in the env:
block of hosted-agent/azure.yaml. It cannot be supplied as an App
Configuration key: the orchestrator reads it from the process environment before
any App Configuration resolution happens.
Hosted-agent environment variables are baked into an agent version and are
immutable once that version exists, so changing this value always creates a new
version. The image digest is unaffected, so no rebuild occurs. Set only this
variable, never the legacy AZURE_TRACING_GEN_AI_CONTENT_RECORDING_ENABLED:
when both are present and disagree, the Azure AI instrumentor treats it as a
configuration error.
Runtime readiness passes; the remaining evidence gates stay closed
Readiness has been re-run under NETWORK_ISOLATION=true against the pins
shipped in GPT-RAG v3.8.0:
the hosted agent reaches active state, session readiness succeeds, and a
grounded answer returns with its citation. That covers item 1 of the
evidence gate only. Keep continuity and panel evidence gates false/off/503,
keep the classic rollback available, and do not describe items 2 through 8
as validated until each one succeeds. See the
hosted-agent component release matrix.
Hosted runtime bootstrap permissions
Hosted bootstrap patch; component pins unchanged
GPT-RAG v3.8.4 adds this scoped deployment bootstrap without repinning
UI v2.6.2, orchestrator v4.1.1, ingestion v2.7.3, or AI Landing Zone
v2.5.1 from the v3.8.3
integration matrix.
Topology defaults and continuity and panel evidence gates are unchanged.
An active hosted version and a successful readiness probe do not establish
access to configuration, secrets, models, or documents. The bootstrap targets
the actual deployed Foundry agent's instance_identity.principal_id, discovered
after agent creation. The Foundry project identity, deployment operator, and
Container App identities are not substitutes; bootstrap creates no new identity.
The deployed agent must have one route receiving 100% of traffic. That route
may select an explicit numeric version or @latest; bootstrap resolves
@latest through the returned latest-version metadata and verifies the concrete
hosted version before planning grants for the agent instance identity. That
identity comes from the agent root's instance_identity.principal_id, not
versions.latest or the Foundry project. Malformed or ambiguous bindings fail
before any role-assignment writes.
The shared config/deployment/hosted_access.py contract separates a read-only
plan from an explicit apply. Planning discovers the instance identity,
dependency scopes, and existing assignments without changing resources. Apply
creates only missing, exact-scoped assignments from the following allowlist:
| Runtime dependency | Bootstrap role | Assignment scope |
|---|---|---|
| Solution App Configuration | App Configuration Data Reader | Exact configuration store |
| Solution Azure OpenAI models | Cognitive Services OpenAI User | Exact model account |
Configured AUDIT_HMAC_KEY Key Vault reference, when present |
Key Vault Secrets User | Exact referenced secret, not the whole vault |
The audit reference is inspected as metadata only: neither plan nor apply retrieves or logs the secret value. An absent audit reference means no secret grant and no secret creation. Missing or malformed identity, resource mapping, or reference metadata must fail without choosing another identity or broadening the scope. Audit signing is distinct from disabled Conversation-capability/HMAC fallback; bootstrap does not enable that fallback.
The deployment operator needs discovery read access and
Microsoft.Authorization/roleAssignments/write at the planned scopes to create
missing grants. Alternatively, an authorized access administrator can pre-create
the exact assignments. Matching unconditional assignments are reused on repeat
runs. A conditional assignment for an otherwise exact principal/role/scope
tuple blocks all bootstrap writes; bootstrap does not add another unconditional
assignment to bypass that condition. Have an authorized access administrator
review the conflict separately. If a later grant fails during apply, valid
earlier grants remain in place for idempotent recovery; do not roll them back
or add broader permissions to bypass the failure.
The child hosted-agent/azure.yaml registers
services.orchestrator-agent.hooks.postdeploy with
../scripts/bootstrapHostedAccess.ps1 on Windows and
../scripts/bootstrapHostedAccess.sh on POSIX. These hooks explicitly apply the
allowlisted grants after agent creation, covering direct child deployment as
well as deployment from the root. The direct child hook is RBAC-only: it does
not invoke the model or incur a smoke-request inference charge.
Root deployment waits for the child hook to succeed, then sends Hello! through
POST /invocations in a new hosted session. The shared smoke validator requires
a completed response containing non-empty assistant text and rejects error,
failed, incomplete, or cancelled outcomes. It does not require an exact reply
marker. Bootstrap failure or a failed greeting stops UI cutover.
Document access is not included. Bootstrap grants no Search or Storage Blob roles, elevated read, managed Conversation or impersonation access, Owner, or Contributor. It makes no network changes and cannot repair private connectivity. Keep retrieval permissions and delegated document authorization separate. An explicitly approved synthetic service-identity container grant also covers future files in that container; it must never become a product default.
Plan and recover hosted access
Run these commands from the repository root of the implementation containing
the bootstrap fix, after the agent has been deployed. --azd-env reads context
from the hosted child project's azd environment. To select an environment
explicitly, first set process-local AZURE_ENV_NAME with
$env:AZURE_ENV_NAME = "<environment-name>" in PowerShell or
export AZURE_ENV_NAME="<environment-name>" in POSIX. Otherwise the saved child
azd environment is used; no persistent environment configuration is needed.
PowerShell restores both the working directory and the previous process-local
PYTHONPATH:
$previousPythonPath = $env:PYTHONPATH
try {
$env:PYTHONPATH = (Get-Location).Path
Push-Location hosted-agent -ErrorAction Stop
try {
python -m config.deployment.hosted_access --azd-env --plan
}
finally {
Pop-Location
}
}
finally {
$env:PYTHONPATH = $previousPythonPath
}
POSIX keeps the directory and import-path changes inside a subshell:
( cd hosted-agent && PYTHONPATH=.. python3 -m config.deployment.hosted_access --azd-env --plan )
Planning is read-only and is the default when neither --plan nor --apply is
specified. A valid plan can exit 0 while required grants are missing. Inspect
grants[].exact_unconditional_assignment to see whether each exact grant
already exists. data_plane_readiness is not-tested: neither a successful
plan nor an assignment visible through Azure Resource Manager proves that
authorization has propagated to the data plane.
After reviewing the plan and obtaining authorization for missing grants, repeat
the PowerShell block with --plan replaced by --apply; its inner command is
python -m config.deployment.hosted_access --azd-env --apply. On POSIX:
( cd hosted-agent && PYTHONPATH=.. python3 -m config.deployment.hosted_access --azd-env --apply )
If a greeting fails while exact assignments are present, wait for propagation and rerun the idempotent bootstrap, then perform a separately approved, bounded smoke test in a new hosted session. That smoke request invokes the model and may incur inference charges; the recovery CLI does not run it automatically. Keep UI cutover blocked until the smoke validator succeeds. Do not rebuild the image to repair permissions or reuse a process that cached unavailable configuration as evidence that a grant failed.
A greeting exercises model and configuration access only. Successful synthetic
retrieval under service identity does not establish end-user document-level
authorization. Serializer compatibility with the released orchestrator
v4.1.1 was tested offline. A real read-only plan resolved @latest to concrete
version 1, exited 0, and found all three minimal grants already present;
data_plane_readiness remained not-tested. That plan performed no apply or
inference. A fresh automated live deployment/bootstrap-apply/smoke flow has
not been performed for this bootstrap change; these checks do not establish
data-plane readiness. See
runtime-access troubleshooting
and the application/document identity boundary.
Explicit classic fallback
Fallback is a deployment operation, not a request-time retry:
azd env set DEPLOYMENT_TOPOLOGY classic
azd provision
azd deploy
This restores the orchestrator Container App and publishes
CHAT_BACKEND=orchestrator without deleting hosted Conversations or existing
classic panel data.
Current release
GPT-RAG v3.8.5 is the
latest published umbrella release. It replaces the jumpbox deployment flag
gate with actual private DNS/TCP/TLS checks while retaining the exact
component pins, topology defaults, and evidence gates from v3.8.4. See the
network-isolation workflow. Host checks do not
establish fresh installation acceptance; a fresh automated live
installation/bootstrap/smoke run has not been performed for this release.
The preceding v3.8.4 adds the
hosted runtime bootstrap without
changing the component pins, topology defaults, or evidence gates from
v3.8.3. A fresh automated live deployment/bootstrap-apply/smoke flow has not
been performed for this patch.
The preceding GPT-RAG v3.8.3
pins UI v2.6.2, orchestrator v4.1.1,
ingestion v2.7.3, and AI Landing Zone v2.5.1, and it makes hosted/no-panel
the default topology for genuinely fresh deployments. It repins ingestion only:
v2.7.3 stops the data-ingestion administrative surface from being mounted in
deployments that did not ask for it. UI, orchestrator, and AI Landing Zone are
unchanged from v3.8.2. The previous classic-only
umbrella release was v3.7.0, pinning UI v2.3.13, orchestrator v3.8.0,
ingestion v2.5.0, and AI Landing Zone v2.3.0; do not mix the v3.8.3
topology and configuration contract into the older v3.7.0 hooks or manifest.
Do not deploy v3.8.0 or v3.8.1
Both releases are superseded and neither reaches a running deployment.
v3.8.0 pins orchestrator v4.1.0, whose frontend/ SPA build fails in the
first Dockerfile stage, so the orchestrator image cannot be built.
v3.8.1 repairs that build but still pins ingestion v2.7.1, which builds
successfully and then crashes on boot on an OpenTelemetry exporter import.
v3.8.2 is the first release in the v3.8.x line that deploys end to end.
Upgrade if you deployed either.
Retrieval backend
GPT-RAG can retrieve grounding content directly from Azure AI Search or through a
Foundry IQ knowledge base. Starting with GPT-RAG v3.0.2 and AI Landing Zone
v2.1.2, new deployments use Foundry IQ by default through a native Azure Blob
Knowledge Source. Existing deployments can stay on RETRIEVAL_BACKEND=ai_search
until you explicitly migrate.
Use the grounding sources overview to
understand the default Foundry IQ path, when to keep using Azure AI Search, and
when to use the searchIndex pattern for custom GPT-RAG ingestion pipelines.
The most important settings are:
| Setting | Typical value | Purpose |
|---|---|---|
RETRIEVAL_BACKEND |
foundry_iq for new deployments, ai_search for existing compatibility or rollback |
Selects the retrieval path. |
FOUNDRY_IQ_PATTERN |
azureBlob by default, or searchIndex for custom GPT-RAG ingestion |
Selects the Foundry IQ setup choice. |
KNOWLEDGE_BASE_NAME |
<env>-knowledge-base |
Foundry IQ knowledge base name. |
KNOWLEDGE_BASE_CONNECTION_ID |
Generated by AILZ | Dedicated Foundry connection for knowledge-base use. |
FOUNDRY_IQ_API_VERSION |
2026-05-01-preview |
Required for per-user permissions and custom ingestion path filterAddOn. |
FOUNDRY_IQ_KNOWLEDGE_RETRIEVAL_BILLING_PLAN |
free or standard |
Controls Azure AI Search agentic retrieval billing. |
With the default Blob path, Foundry IQ processes files directly from the
documents container. GPT-RAG ingestion is not used in that path. Use
FOUNDRY_IQ_PATTERN=searchIndex only when you intentionally keep a custom
GPT-RAG ingestion pipeline that writes chunks to Azure AI Search.
Two optional Foundry IQ Knowledge Sources can run alongside the documents source on the same Knowledge Base. Both are off by default and require signed-in users:
- Foundry IQ: Work IQ (Microsoft 365) blends in mail, meetings, files, chats, and people from the signed-in user's M365 world. Gated public preview.
- Foundry IQ: Fabric ontology (Microsoft Fabric)
- Foundry IQ: Fabric Data Agent (Microsoft Fabric) blends in analytical data from a Fabric ontology (semantic model, lakehouse, warehouse, KQL). Preview. Review data-egress caveats before enabling.
Demo video:
Zero Trust Deployment
For deployments that require network isolation.
Provision infrastructure separately from private data-plane configuration and service deployment:
| Phase | Where to run | Command |
|---|---|---|
| Provision infrastructure | Workstation | azd provision |
| Configure data-plane resources | Host with private VNet/VPN connectivity, subject to the release rules below | scripts/postProvision.ps1 |
| Deploy services | Host with private VNet/VPN connectivity, subject to the release rules below | azd deploy |
Historical v3.8.4: the pre-deploy hook requires
RUN_FROM_JUMPBOX=true when NETWORK_ISOLATION=true. Do not set that flag
on a local machine to work around the release's host restriction.
If you cannot sign in to the administration VM, follow the P2S VPN how-to to prepare private access from a managed local Windows machine. Its simple topology uses one new dedicated resource group for VPN/DNS first and GPT-RAG later, not an existing corporate network group. Read the release boundary and host checks before creating resources.
Private host checks (v3.8.5)
Starting with v3.8.5, a connected VPN/VNet host does not need a jumpbox
declaration to deploy. Use a complete published release containing these
checks, not a development commit or a mixture of release files.
With NETWORK_ISOLATION=true, the deployment hook checks the operating system's DNS
results for RFC1918 private IPv4 destinations, TCP 443, and normal TLS
certificate/hostname validation with SNI. It checks APP_CONFIG_ENDPOINT
before post-provision configuration; APP_CONFIG_ENDPOINT plus
AZURE_AI_PROJECT_ENDPOINT before hosted deployment; and the actual build
registry's AZURE_CONTAINER_REGISTRY_ENDPOINT before a hosted image build.
Public mode does not probe. A prebuilt or reused image
digest avoids an unnecessary hosted-build ACR probe.
An invalid explicit NETWORK_ISOLATION value is an error, not a switch to
public mode.
The checks send no Azure credentials. Invalid/missing endpoints, public or mixed DNS results, certificate errors, and timeouts fail the guarded phase before its writes. No flag bypass, disabled TLS validation, proxy bypass, or public endpoint fallback is part of the workflow.
| Setting | Post-provision behavior |
|---|---|
RUN_FROM_JUMPBOX unset, warning flag false/unset |
Check connectivity and configure; no VPN confirmation prompt or noninteractive automatic skip |
RUN_FROM_JUMPBOX=false, 0, no, or skip |
Explicitly defer configuration; this is not a local-machine selector |
AZURE_SKIP_NETWORK_ISOLATION_WARNING=true, jumpbox flag unset |
Explicitly defer configuration until the operator verifies routes and private access |
Truthy RUN_FROM_JUMPBOX |
Takes precedence over the warning flag, but still requires the same host checks; unnecessary for a connected VPN host |
For the VPN sequence, leave the jumpbox key unset and use only the explicit
warning flag to defer both infrastructure phases. See the
bounded migration/resume note
if a previous guide set the key; do not delete an azd environment to unset it.
When resuming configuration, set both the saved azd warning flag and the
current process AZURE_SKIP_NETWORK_ISOLATION_WARNING to false; an inherited
process value can still defer when the selected environment has no saved key.
The VPN guide shows both assignments around each provision/configuration phase.
These flags affect only post-provision deferral, not pre-deploy or hosted-build
checks. A deferred post-provision hook warns that configuration is incomplete
and can exit successfully; that is not a successful configuration phase.
Failed, empty, or malformed selected-environment reads fail closed rather
than using a stale .azure directory.
Host-check success does not prove RBAC, API health, effective Azure VNet ownership, access to every service, remote build-pool egress, or success of a fresh automated installation. Existing runtime/document permission boundaries and continuity/panel evidence gates are unchanged.
Network Isolation runbook
v3.8.5 sequence: for a fresh VPN-host setup,
use the detailed P2S VPN guide. In outline:
- Select the released source and local
azdenvironment. Enable network isolation; leaveRUN_FROM_JUMPBOXunset. - Set
AZURE_SKIP_NETWORK_ISOLATION_WARNING=trueand run the firstazd provisionin the selected group, explicitly deferring data-plane configuration. - Connect through approved VNet/VPN access and verify private DNS, the P2S return route, TLS, and authorized service reads.
- Sign in with your own account on the local machine, or the VM managed identity when actually using the jumpbox. Clear the deferral with
AZURE_SKIP_NETWORK_ISOLATION_WARNING=false, then runscripts/postProvision.ps1. - For hosted deployment, prepare the image using the provisioned VNet-connected ACR Tasks pool. Restore deferral to
truebefore the secondazd provision; afterward, recheck any manually added return route and private connectivity. - Clear deferral to
false, rerun post-provision configuration successfully, then runazd deploy. A host-check failure is a stop condition, not a reason to set the jumpbox flag or expose a private service publicly.
BUILD_MODE is normally not required when deploying the UI, orchestrator, or ingestion services. The hosted-agent derivative image can use the same dedicated pool. Shared ACR Tasks cannot reach a private endpoint.
Regional preflight
Run preflight before every Zero Trust deployment. It is much faster to fail in the first few minutes than to wait for a long network-isolated deployment and then discover that a regional dependency cannot be created.
azd provision runs the scripts/preProvision hook. The hook invokes
scripts/Invoke-RegionalPreflight.ps1 before the Azure Resource Manager
deployment starts.
Preflight checks include:
- the selected Azure region and provider support,
- common regional readiness checks for Azure AI Search, Cosmos DB, Container Apps, AI Foundry, and Cognitive Services,
- jumpbox VM SKU availability and restrictions,
- Azure OpenAI model quota for the configured deployments.
Preflight is an early warning, not a live capacity reservation. Azure capacity can still change after the check passes, and some regional capacity errors are only returned when Azure creates the resource. Recent examples include Azure AI Search Standard capacity in Sweden Central and Cosmos DB zonal capacity in West Europe.
Use the result this way:
| Result | Operator action |
|---|---|
FAIL |
Stop. Fix the subscription, quota, region, or parameter issue before provisioning. |
WARN |
Review the warning before continuing. If it mentions capacity or regional risk, consider changing region first. |
| Pass | Continue, but keep the deployment logs open because live capacity can still change. |
If a region fails or warns on a critical dependency, try another fully supported
region instead of waiting 30 minutes or more for a deployment that is likely to
fail. Use GPT_RAG_REGIONAL_PREFLIGHT_SKIP=true only when you intentionally
bypass regional checks, or PREFLIGHT_SKIP=true to bypass all preflight hooks.
Before Provisioning
Enable network isolation in your environment:
azd env set NETWORK_ISOLATION true
Optional v2 parameters can be set before provisioning:
azd env set DEPLOYMENT_MODE standalone
azd env set VM_SIZE Standard_D2s_v3
azd env set ENABLE_COSMOS_ANALYTICAL_STORAGE false
ALLOWED_IP_RANGES is also available for CIDR allow-listing, but because it is an array parameter, prefer editing main.parameters.json or using a parameter overlay rather than storing a complex array in the azd environment.
Make sure youβre signed in with your Azure user account:
az login
azd auth login
Add
--tenantforazor--tenant-idforazdif you want a specific tenant.
Provision Infrastructure
azd env set AZURE_SKIP_NETWORK_ISOLATION_WARNING true # explicitly defers data-plane configuration
azd provision
Post-Provision Configuration
With NETWORK_ISOLATION=true, data-plane configuration needs private VNet/VPN
access. For the VPN workflow, keep the jumpbox key unset and explicitly
defer configuration during infrastructure-only phases. Once connectivity is
verified, set AZURE_SKIP_NETWORK_ISOLATION_WARNING=false and run
scripts/postProvision.ps1; actual checks replace a declaration or prompt.
Without explicit deferral, the v3.8.5 hook fails when connectivity is unavailable
rather than silently skipping a noninteractive run. The older v3.8.4
still uses its historical confirmation/defer behavior.
Using the Jumpbox VM
1) Reset the VM password in the Azure Portal (required on first access if not set in deployment parameters):
- Go to your VM resource β Support + troubleshooting β Reset password β Set new credentials
- Default username is
testvmuser
2) Connect via Azure Bastion
3) Authenticate with the VM's Managed Identity:
az login --identity
azd auth login --managed-identity
Add
--tenantforazor--tenant-idforazdif you want a specific tenant.
4) Run the post-provision script:
The RUN_FROM_JUMPBOX=true example below is for an actual jumpbox and is
compatible with the v3.8.5 workflow, but never bypasses its checks.
It takes precedence over the warning flag, so use the separate
VPN deferral sequence
for a local machine.
PowerShell:
cd C:\github\GPT-RAG
azd env set RUN_FROM_JUMPBOX true
.\scripts\postProvision.ps1
Bash:
cd /mnt/c/github/gpt-rag
./scripts/postProvision.sh
Note: If you have re-initialized or cloned the gpt-rag repo again, refresh your
azdenvironment before running the postProvision script so it points to the existing deployment:azd init -t azure/gpt-ragthenazd env refresh. When prompted, select the same Subscription, Resource Group, and Location as the original provisioning soazdcorrectly links to your environment.
Existing Platform / AI Landing Zone Integrated
Use these settings when GPT-RAG must deploy into an existing enterprise platform, such as a hub-spoke network with centrally managed Private DNS Zones, Log Analytics, Application Insights, Bastion, NAT Gateway, or Azure Firewall.
Core mode: set DEPLOYMENT_MODE to ailz-integrated, then pass the existing resource IDs that your platform team owns. The default remains standalone, so basic deployments do not require these settings.
azd env set DEPLOYMENT_MODE ailz-integrated
azd env set USE_EXISTING_VNET true
azd env set EXISTING_VNET_RESOURCE_ID "/subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.Network/virtualNetworks/<vnet>"
Existing Private DNS Zones: set the zone resource IDs for services already managed by the platform. Common values include EXISTING_PRIVATE_DNS_ZONE_OPENAI_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_AISERVICES_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_SEARCH_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_COSMOS_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_BLOB_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_KEYVAULT_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_APPCONFIG_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_CONTAINERAPPS_RESOURCE_ID, EXISTING_PRIVATE_DNS_ZONE_ACR_RESOURCE_ID, and Azure Monitor / App Insights zone IDs.
azd env set EXISTING_PRIVATE_DNS_ZONE_SEARCH_RESOURCE_ID "/subscriptions/<sub>/resourceGroups/<dns-rg>/providers/Microsoft.Network/privateDnsZones/privatelink.search.windows.net"
azd env set EXISTING_PRIVATE_DNS_ZONE_OPENAI_RESOURCE_ID "/subscriptions/<sub>/resourceGroups/<dns-rg>/providers/Microsoft.Network/privateDnsZones/privatelink.openai.azure.com"
azd env set DNS_ZONE_LINK_SUFFIX "<unique-spoke-name>"
Shared platform resources: set EXISTING_LOG_ANALYTICS_WORKSPACE_RESOURCE_ID, EXISTING_APPLICATION_INSIGHTS_RESOURCE_ID, EXISTING_APPLICATION_INSIGHTS_CONNECTION_STRING, HUB_INTEGRATION_HUB_VNET_RESOURCE_ID, HUB_INTEGRATION_EGRESS_NEXT_HOP_IP, or HUB_INTEGRATION_EXISTING_ROUTE_TABLE_RESOURCE_ID when those resources are centrally managed.
Spoke resource switches: use DEPLOY_JUMPBOX, DEPLOY_BASTION, DEPLOY_NAT_GATEWAY, EXISTING_JUMPBOX_RESOURCE_ID, EXISTING_BASTION_RESOURCE_ID, EXISTING_NAT_GATEWAY_RESOURCE_ID, DEPLOY_AZURE_FIREWALL, and DEPLOY_ACR_TASK_AGENT_POOL to align GPT-RAG with the platform topology. These are optional and preserve the default standalone behavior when unset.
Deploy GPT-RAG Services
Note: For Zero Trust deployments with network isolation, deploy services from the jumpbox or another host with VNet connectivity. If using the jumpbox VM, the repositories are located in the
C:\githubdirectory.
Once the GPT-RAG infrastructure is provisioned, you can deploy the services.
To deploy all services at once from the actual jumpbox, navigate to the
gpt-rag directory (with azd environment configured) and run the example below.
For a local VPN host, use the v3.8.5
two-phase sequence
instead of copying its jumpbox path and flag.
cd C:\github\GPT-RAG
azd env set RUN_FROM_JUMPBOX true
azd env set NETWORK_ISOLATION true
azd env set ACR_TASK_AGENT_POOL build-pool
azd deploy
This command deploys the services selected by the active mode. The UI, orchestrator, ingestion, and hosted-agent derivative images can use Azure Container Registry remote builds (az acr build) against the dedicated VNet-connected ACR Tasks agent pool provided by AILZ v2.4.1. Shared ACR Tasks cannot reach a private endpoint.
The deploy hook uses NETWORK_ISOLATION as the source of truth, not the older
AZURE_ZERO_TRUST variable. The older v3.8.4 requires
RUN_FROM_JUMPBOX=true for isolated deployment. Starting with v3.8.5,
that gate is replaced by private host checks:
a connected VPN/VNet host does not need the flag, and setting it cannot make
a failed DNS/TCP/TLS check pass.
If you prefer to deploy a single service, for example, when updating only that service, you can deploy it individually. Below is an example using the orchestrator service. The same approach applies to other services (frontend, dataingest, mcp).
Deploy Individual Services
Make sure you're logged in to Azure:
az login
Example: Deploying the Orchestrator
Using azd (recommended):
Initialize the template:
azd init -t azure/gpt-rag-orchestrator
Important: Use the same environment name with
azd initas in the infrastructure deployment to keep components consistent.
Update environment variables then deploy:
azd env refresh
azd deploy
Important: Run
azd env refreshwith the same subscription and resource group used in the infrastructure deployment.
Using a shell script:
Clone the repository, set the App Configuration endpoint, and run the deployment script.
PowerShell (Windows):
git clone https://github.com/Azure/gpt-rag-orchestrator.git
$env:APP_CONFIG_ENDPOINT = "https://<your-app-config-name>.azconfig.io"
cd gpt-rag-orchestrator
.\scripts\deploy.ps1
Bash (Linux/macOS):
git clone https://github.com/Azure/gpt-rag-orchestrator.git
export APP_CONFIG_ENDPOINT="https://<your-app-config-name>.azconfig.io"
cd gpt-rag-orchestrator
./scripts/deploy.sh
Permissions
The role tables below describe the classic Container Apps topology. The pinned hosted matrix omits the orchestrator Container App assignments and adds Foundry data-plane, delegated-user, Toolbox, and immutable-image pull assignments. Hosted-panel adds only the narrow panel metadata-container assignments; its user and operator surfaces remain off/503 behind independent evidence gates.
Microsoft Foundry Role and AI Search Assignments
| Resource | Role | Assignee | Description |
|---|---|---|---|
| GenAI App Search Service | Search Index Data Reader | Microsoft Foundry Project | Read index data |
| GenAI App Search Service | Search Service Contributor | Microsoft Foundry Project | Create AI Search connection |
| GenAI App Storage Account | Storage Blob Data Reader | Microsoft Foundry Project | Read blob data |
| Microsoft Foundry Account | Cognitive Services User | Search Service | Allow Search Service to access vectorizers |
Container App Role Assignments
| Resource | Role | Assignee | Description |
|---|---|---|---|
| GenAI App Configuration Store | App Configuration Data Reader | ContainerApp: orchestrator | Read configuration data |
| GenAI App Configuration Store | App Configuration Data Reader | ContainerApp: frontend | Read configuration data |
| GenAI App Configuration Store | App Configuration Data Reader | ContainerApp: dataingest | Read configuration data |
| GenAI App Configuration Store | App Configuration Data Reader | ContainerApp: mcp | Read configuration data |
| GenAI App Container Registry | AcrPull | ContainerApp: orchestrator | Pull container images |
| GenAI App Container Registry | AcrPull | ContainerApp: frontend | Pull container images |
| GenAI App Container Registry | AcrPull | ContainerApp: dataingest | Pull container images |
| GenAI App Container Registry | AcrPull | ContainerApp: mcp | Pull container images |
| GenAI App Key Vault | Key Vault Secrets User | ContainerApp: orchestrator | Read secrets |
| GenAI App Key Vault | Key Vault Secrets User | ContainerApp: frontend | Read secrets |
| GenAI App Key Vault | Key Vault Secrets User | ContainerApp: dataingest | Read secrets |
| GenAI App Key Vault | Key Vault Secrets User | ContainerApp: mcp | Read secrets |
| GenAI App Search Service | Search Index Data Reader | ContainerApp: orchestrator | Read index data |
| GenAI App Search Service | Search Index Data Contributor | ContainerApp: dataingest | Read/write index data |
| GenAI App Search Service | Search Index Data Contributor | ContainerApp: mcp | Read/write index data |
| GenAI App Storage Account | Storage Blob Data Reader | ContainerApp: orchestrator | Read blob data |
| GenAI App Storage Account | Storage Blob Data Reader | ContainerApp: frontend | Read blob data |
| GenAI App Storage Account | Storage Blob Data Contributor | ContainerApp: dataingest | Read/write blob data |
| GenAI App Storage Account | Storage Blob Data Contributor | ContainerApp: mcp | Read/write blob data |
| GenAI App Cosmos DB | Cosmos DB Built-in Data Contributor | ContainerApp: orchestrator | Read/write Cosmos DB data |
| Microsoft Foundry Account | Cognitive Services User | ContainerApp: orchestrator | Access Cognitive Services |
| Microsoft Foundry Account | Cognitive Services User | ContainerApp: dataingest | Access Cognitive Services |
| Microsoft Foundry Account | Cognitive Services User | ContainerApp: mcp | Access Cognitive Services |
| Microsoft Foundry Account | Cognitive Services OpenAI User | ContainerApp: orchestrator | Use OpenAI APIs |
| Microsoft Foundry Account | Cognitive Services OpenAI User | ContainerApp: dataingest | Use OpenAI APIs |
| Microsoft Foundry Account | Cognitive Services OpenAI User | ContainerApp: mcp | Use OpenAI APIs |
Executor Role Assignments
| Resource | Role | Assignee | Description |
|---|---|---|---|
| GenAI App Configuration Store | App Configuration Data Owner | Executor | Full control over configuration settings |
| GenAI App Container Registry | AcrPush | Executor | Push container images |
| GenAI App Container Registry | AcrPull | Executor | Pull container images |
| GenAI App Key Vault | Key Vault Contributor | Executor | Manage Key Vault settings |
| GenAI App Key Vault | Key Vault Secrets Officer | Executor | Create Key Vault secrets |
| GenAI App Search Service | Search Service Contributor | Executor | Create/update search service elements |
| GenAI App Search Service | Search Index Data Contributor | Executor | Read/write search index data |
| GenAI App Search Service | Search Index Data Reader | Executor | Read index data |
| GenAI App Storage Account | Storage Blob Data Contributor | Executor | Read/write blob data |
| GenAI App Cosmos DB | Cosmos DB Built-in Data Contributor | Executor | Read/write Cosmos DB data |
| Microsoft Foundry Account | Cognitive Services OpenAI User | Executor | Use OpenAI APIs |
Jumpbox VM Role Assignments
| Resource | Role | Assignee | Description |
|---|---|---|---|
| GenAI App Container Apps | Container Apps Contributor | Jumpbox VM | Full control over Container Apps |
| Azure Managed Identity | Managed Identity Operator | Jumpbox VM | Assign and manage user-assigned identities |
| GenAI App Container Registry | Container Registry Repository Writer | Jumpbox VM | Write to ACR repositories |
| GenAI App Container Registry | Container Registry Tasks Contributor | Jumpbox VM | Manage ACR tasks |
| GenAI App Container Registry | Container Registry Data Access Configuration Administrator | Jumpbox VM | Manage ACR data access configuration |
| GenAI App Container Registry | AcrPush | Jumpbox VM | Push container images |
| GenAI App Configuration Store | App Configuration Data Owner | Jumpbox VM | Full control over configuration settings |
| GenAI App Key Vault | Key Vault Contributor | Jumpbox VM | Manage Key Vault settings |
| GenAI App Key Vault | Key Vault Secrets Officer | Jumpbox VM | Create Key Vault secrets |
| GenAI App Search Service | Search Service Contributor | Jumpbox VM | Create/update search service elements |
| GenAI App Search Service | Search Index Data Contributor | Jumpbox VM | Read/write search index data |
| GenAI App Storage Account | Storage Blob Data Contributor | Jumpbox VM | Read/write blob data |
| GenAI App Cosmos DB | Cosmos DB Built-in Data Contributor | Jumpbox VM | Read/write Cosmos DB data |
| Microsoft Foundry Account | Cognitive Services Contributor | Jumpbox VM | Manage Cognitive Services resources |
| Microsoft Foundry Account | Cognitive Services OpenAI User | Jumpbox VM | Use OpenAI APIs |