Provision a GPU-enabled TauGrid AKS environment

Build AKS, GPU capacity, TauGrid, and Portal with the repository Terraform root.

Use the repository’s terraform/aks root to create a GPU-enabled AKS environment. It provisions a system pool and a GPU pool, enables OIDC and Azure Workload Identity, and invokes the supported tau cluster install command.

tau cluster install installs the versioned TauGrid distribution that owns Kueue, KubeRay, the Tau controller, GPU monitoring, the baseline Kueue queue, and Portal; skip installing those components separately with Helm.

Prerequisites

  • an Azure subscription that passes the AKS cluster prerequisites, has GPU quota for the selected region and SKU, and an approved Terraform identity;
  • Azure CLI, Terraform 1.9 or later, kubectl, Helm, and local Python dependencies;
  • Azure credentials accepted by the AzureRM Terraform provider and permission to provision AKS, networking, storage, and identities; and
  • tau and PowerShell 7 on PATH. Linux and macOS users can configure the Terraform command interpreter to use Bash.

The default deployment creates one Standard_NC24ads_A100_v4 node. This node has one A100 80 GB GPU and is billable. Before applying, verify that the target region has capacity for the corresponding VM family. Change gpu_vm_size, gpu_count_per_node, and gpu_monitoring_sku_name together when selecting another GPU SKU.

Deploy

From a checkout of this repository:

cd terraform/aks
terraform init
terraform apply -var="subscription_id=<your-subscription-id>"

By default, Terraform uses gpu_stack_mode = "self_managed": it installs a standalone NVIDIA device plugin plus the upstream NVIDIA DCGM exporter. NVIDIA GPU Operator remains a separate existing-cluster model, while the cluster platform retains ownership of the underlying NVIDIA driver. Terraform also creates a local ignored admin kubeconfig and values file under terraform/aks/generated/. For the default A100 pool, it then normalizes MIG mode, restarts the GPU VM scale set, and waits for allocatable GPUs before running:

tau cluster install --values generated/taugrid-values.yaml --version 0.4.0

In this mode, the generated values configure TauGrid GPU monitoring with dcgmHealth.source: exporter and exporterUrl: http://dcgm-exporter.dcgm-exporter.svc:9400/metrics, pointing at the standalone DCGM exporter’s own Service rather than a host-local endpoint.

Set gpu_stack_mode = "aks_managed_preview" to use AKS Managed GPU Experience instead. This preview mode uses EnableManagedGPUExperience=true at GPU pool creation; AKS then owns the NVIDIA driver, device plugin, and a node-local DCGM exporter host service at port 19400. TauGrid GPU monitoring uses dcgmHealth.source: host-dcgmi with the default http://localhost:19400/metrics in this mode. Before applying, register the feature, wait for Registered, and refresh the AKS resource provider:

az feature register --namespace Microsoft.ContainerService --name ManagedGPUExperiencePreview
az feature show --namespace Microsoft.ContainerService --name ManagedGPUExperiencePreview --query properties.state --output tsv
az provider register --namespace Microsoft.ContainerService
az provider show --namespace Microsoft.ContainerService --query registrationState --output tsv

GPU cluster autoscaling is not supported in the managed preview. The current AzureRM provider does not expose gpuProfile.nvidia.managementMode, so the tag is a temporary workaround. When AzureRM exposes that field, replace the tag with the provider setting; retain the feature registration until AKS makes the feature generally available.

If the local installation step fails after AKS has been created, correct the local prerequisite and rerun terraform apply. Terraform replaces that step when its cluster, generated values, or TauGrid version changes.

An existing cluster with an externally managed GPU Operator is a third, distinct operational model alongside Terraform’s standalone and AKS-managed modes. GPU Operator owns the GPU software stack according to its own ClusterPolicy, and TauGrid consumes the Operator’s DCGM exporter endpoint with dcgmHealth.source: exporter and an explicit non-loopback Service URL (commonly port 9400). See the GPU software stack models comparison and the GPU monitoring chart’s DCGM health sources documentation. Regardless of which of these three models owns the stack, workload configs always request standard Kubernetes nvidia.com/gpu resources unchanged.

Verify and provision a workspace

Fetch an operator kubeconfig:

terraform output -raw get_credentials_command

Run the command printed above to fetch the operator kubeconfig.

Verify both TauGrid and GPU capacity:

tau cluster validate installation --timeout 10m
kubectl get nodes -l accelerator=nvidia
kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}'

The default path is an operator sandbox using local administrator credentials. It creates platform infrastructure only, so a platform owner provisions the researcher workspace with an explicit Entra group object ID before submitting workloads:

tau workspace create taugrid-default \
  --namespace taugrid-default \
  --principal-name <entra-group-object-id> \
  --apply
tau workspace check taugrid-default

To provision the Entra-backed workspace in the same apply, set the following before the first terraform apply. Terraform applies a native TauWorkspace after tau cluster install; the controller reconciles the workload Namespace, jobqueue LocalQueue, and namespace-scoped researcher RBAC.

workspace_namespace = "taugrid-default"

bootstrap_workspace = {
  name                  = "taugrid-default"
  entra_group_object_id = "<entra-group-object-id>"
}

When bootstrap is enabled, Terraform configures Portal’s computed Jobs board with an explicit operator scope limited to this workspace namespace and jobqueue. This is still an operator-only ClusterIP diagnostic path, not a researcher-facing authenticated Portal endpoint.

On a retained cluster, removing bootstrap_workspace stops Terraform from applying the CR but intentionally does not delete an existing workspace or its workloads. Remove a workspace through the workspace administration workflow after reviewing the impact.

Optional ADX and lifecycle history

ADX observability is opt-in because it creates additional billable resources. Both ADX and lifecycle history can be enabled in the initial plan and apply:

enable_adx                = true
adx_cluster_name          = ""
enable_lifecycle_recorder = true
workspace_namespace       = "taugrid-default"

An empty adx_cluster_name generates a stable 20-character taugrid<13-hex-characters> candidate from the subscription, resource group, and AKS cluster names. Set an explicit globally unique name only if Azure reports that the candidate is in use. Terraform preserves the deployed name on later applies, including a legacy 15-character automatic name, so upgrading does not replace the cluster. After apply, inspect the deployed name with terraform output -raw adx_cluster_name. An explicit name is used unchanged for a new cluster; changing the name of an existing cluster requires an explicit data-preserving migration.

Terraform installs adx-mon alongside TauGrid. In self-managed mode it discovers the upstream DCGM exporter Pod. TauGrid GPU monitoring uses that exporter’s node-local Service in self-managed mode; in AKS managed preview mode it collects from the GPU node host service.

When lifecycle history is enabled, Terraform creates workspace_namespace before installing TauGrid so the chart can render the recorder’s namespace scoped RBAC in the same apply. If bootstrap_workspace is configured, Terraform applies the TauWorkspace after installing TauGrid. Otherwise, the tau workspace create command above adopts the bootstrapped namespace. Both paths let the controller reconcile its labels, LocalQueue, RBAC, and other workspace settings.

For a real GPU workload, use the A100 GPU quickstart after the GPU allocatable-resource check succeeds. It verifies CUDA execution, not only scheduling, and can incur additional GPU cost.

Portal

Terraform uses the TauGrid distribution default and installs Portal as tau-portal in the tau-system namespace. Its Service is ClusterIP-only, keeping it reachable only from inside the cluster network. An operator can inspect it with:

kubectl port-forward service/tau-portal 18080:80 --namespace=tau-system

This is an operator diagnostic path; researcher access requires a platform-owned authenticated HTTPS proxy in front of Portal. The default Portal serves Kubernetes-backed boards. Experiment, cluster-health, and cost boards require separately configured ADX and Azure Workload Identity.

Destroy

Destroy the environment when it is no longer needed to stop GPU billing:

terraform destroy -var="subscription_id=<your-subscription-id>"
Last modified August 31, 2026: Prepare TauGrid 0.4.0 Release (#217) (c9245db)