Cloud-native AI infrastructure for GPU workloads on Kubernetes

TauGrid brings the tau CLI, Kueue queueing, KubeRay orchestration, GPU-node health monitoring, and observability into one stack, spanning data preparation, distributed training, fine-tuning, and inference. Platform teams get an integrated foundation; researchers focus on their code, not Kubernetes plumbing.

From submission to healthy GPUs

A workload submitted with tau run enters a shared Kueue ClusterQueue, receives quota, and is placed by Kubernetes on healthy GPUs. Queue state, placement, and live training evidence stay visible along the way.

An animated TauGrid control room showing jobs from three workspaces sharing a Kueue ClusterQueue while the front job receives GPU quota and Kubernetes places it on two available healthy GPUs.
One shared queue across workspaces Watch jobs enter the ClusterQueue, receive best-effort FIFO admission, and start on healthy GPU capacity.

Across the GPU workload lifecycle

The same stack carries a workload from raw data through to a running endpoint.

An animated TauGrid pipeline carrying a model workload through data preparation, distributed training, fine-tuning, and a ready inference endpoint.
One workload, one continuous record Move from prepared data to a serving endpoint while TauGrid preserves workload evidence.

Built in the open

TauGrid is MIT-licensed and developed in the open on GitHub.