Platform admin guide

Install TauGrid, provision workspaces, and operate clusters

This guide is for platform owners and on-call operators who install TauGrid, provision governed workspaces, and keep clusters healthy.

Installation guides

Setup guides

Operate

Troubleshooting

A run stuck or failed? Start with Troubleshoot a run, the same canonical, layer-by-layer decision path used across these pages to identify which layer owns the problem.


Installation guides

Install TauGrid on Kubernetes or prepare an AKS cluster

Setup guides

Prepare workspaces and optional platform services

Identity and security boundaries

Separate human Kubernetes authorization from workload cloud identity

Queue, quota, topology, and GPU placement

How policy intent becomes admitted and scheduled pods

Observability and evidence

Immediate lifecycle, durable experiment state, and fleet telemetry

Troubleshooting

Diagnose failed runs and recover safely

Workload profile migration

Move placement catalogs into TauCluster and operate profile revisions safely

Multi-cluster execution

Deterministic dispatch to preconfigured MultiKueue workers