Supported environments
Plan your deployment
Choose infrastructure based on your model, expected concurrency, context length, and latency requirements. GPU memory alone does not determine a supported configuration. When planning a deployment, consider:- Model: Larger models require more GPU memory and compute capacity.
- Concurrency and context length: More simultaneous requests and longer contexts increase memory requirements.
- GPU topology: The required GPU count depends on the GPU family and available memory per GPU.
- Host resources: CPU and host memory requirements depend on your deployment topology and workload.
Planning estimates
Use the estimates below as a starting point for infrastructure planning, not as minimum requirements, a supported configuration, or a guarantee of performance for a particular workload. Context length, concurrency, batch size, latency requirements, and GPU topology all affect the resources a deployment actually needs. For concurrent-agent capacity, developer-seat estimates, and models not listed here, contact your Poolside account team for configuration guidance.
For cloud deployments, storage requirements depend on whether S3-compatible storage is colocated on the inference node.
For on-premises deployments, see Storage requirements for guidance on sizing and configuring storage.
CPU and host memory
CPU and host memory needs scale with GPU count rather than with a specific model. As an approximate planning estimate, independent of which model a node serves:- At least 8 CPU cores per GPU
- Host memory at least equal to total GPU memory
GPU types
Poolside model inference targets the following NVIDIA GPU families:
The per-model table provides planning guidance, not a strict floor or a GPU selector. Context length, batch size, and the number of GPUs affect the configuration that serves a model.