Skip to main content
Poolside supports model inference deployments in GPU-backed Kubernetes environments. Use this page to review supported environments and size the infrastructure that serves Poolside models.
Poolside support covers the environments and minimum requirements documented on this page. For other Kubernetes environments, GPU configurations, storage backends, or security requirements, contact your Poolside account team to discuss support options.

Supported environments

Inference requirements

Poolside models have different minimum requirements. Use this table to size your inference nodes. For concurrent-agent capacity and developer-seat estimates, contact your Poolside account team. For local development on your own hardware, see How to run a Poolside model locally. For cloud deployments, storage requirements depend on whether S3-compatible storage is colocated on the inference node. For on-premises deployments, see Storage requirements for guidance on sizing and configuring storage.

GPU memory reference

Poolside model inference targets the following NVIDIA GPU families: RTX 6000 Blackwell, H100, and H200. The memory per GPU is listed here for reference against the per-model minimum GPU memory table above: Which GPUs suit a given model is model-dependent. The per-model minimum GPU memory table is a floor, not a GPU selector: context length, batch size, and the number of GPUs all affect what actually serves a model. Confirm the right combination of model, GPU type, and GPU count for your workload with your Poolside account team.

Support and compatibility

For questions about integration with specific enterprise tooling or deployment workflows, contact Poolside.