When consulting about GPU clusters or reserved resources, don't just provide a GPU model and quantity. First, write a checklist covering what this batch of GPUs will run, whether high-speed communication is needed between GPUs, how data will be read and written, how long it will be used, and what average utilization to expect, then submit it through the enterprise contact entry. This way, the other party can directly determine whether suitable nodes are available, how long delivery will take, and how billing works, without repeated follow-up questions.
Below, we explain in practical order: first determine whether consultation is needed, then organize the checklist, and finally clarify what to ask during communication.
First Determine: Does Your Requirement Need Manual Coordination?
Not all multi-GPU requirements require contacting sales. On NexGPU, single-GPU and standard multi-GPU instances can be ordered self-service in the console. They are billed by the hour, metered by the second, have no minimum spend, and require no contract. The unit price is locked at order time and remains unchanged until destruction. Bills only include compute, storage, and traffic.
The following cases are suitable for manual coordination:
- Large-scale cross-node clusters: Multiple machines working together for distributed training, with requirements on inter-node networking;
- Specified data center or region: Due to data compliance, access latency, or other reasons, you need to bind to a specific location;
- Dedicated reserved resources: You need to guarantee access to a fixed batch of GPUs over a period;
- Bulk compute procurement or batch quotas: Large GPU counts, or delivery in batches according to a schedule.
If you only want to run a fine-tuning job on two or three GPUs, or first verify whether an approach is feasible, self-service rental is faster. If you're not sure whether a single GPU is sufficient, you can first look at the several signals in When Consumer-Grade GPUs Become Insufficient.

Requirements Checklist to Prepare Before Consulting
1. Workload: What Will Run, and How Large Is the Model
First clarify whether it is large-scale training, batch fine-tuning, or continuous inference, then add the model parameter scale, the minimum VRAM requirement per GPU, and which parallelism strategy you plan to use (tensor parallelism, data parallelism, or pipeline parallelism).
These items determine whether you need single-machine multi-GPU (Scale-Up) or a cross-machine cluster (Scale-Out). For example, tensor parallelism is very sensitive to inter-GPU bandwidth and usually should be placed within the same machine. Pure data-parallel inference replicas can be distributed across multiple machines. For how to estimate VRAM and determine GPU count, refer to How to Choose GPUs for Large Model Training.
2. Interconnect: Is High-Speed Communication Needed Between GPUs or Machines?
This item has the greatest impact on price and deployment method, so be specific:
- Multi-machine distributed training: State whether intra-machine NVLink is needed, whether inter-node high-speed interconnect such as InfiniBand or RoCE is needed, and the expected bandwidth;
- Independent inference concurrency pools: Each instance handles requests independently, no cross-machine communication is needed, and standard Ethernet nodes are usually sufficient.
These two cases differ greatly in cost and architecture. If you submit an inference pool requirement as if it were a training cluster, you may pay extra for interconnect you don't need. Conversely, if you omit the interconnect required for training, you may find multi-machine scaling efficiency very low after receiving the nodes. For the minimum rental GPU count and existing data center distribution for high-speed interconnect clusters, please confirm directly with the other party during consultation.
3. Storage and Data: How Fast to Read, How Much to Store, How Often to Write
You need to specify the following:
- Training dataset size and required read throughput;
- How often checkpoints are written and how large each write is;
- How large a storage volume is needed, and whether it must be shared across multiple nodes.
Also plan in advance where data will be placed. On NexGPU, storage fees continue after an instance is stopped; all fees stop only after destruction, but data on the disk is also cleared. For long-term projects, it's best to decide in advance which data remains on the instance disk and which is migrated elsewhere. For details, see How to Save Data on Rented GPU Instances.
4. Runtime Environment: Use Existing Images or Your Own Containers
State the environment you depend on, such as common frameworks like PyTorch, vLLM, or TGI, or whether you need to use a privately built container image from your enterprise. If using a private image, also provide CUDA and driver version requirements. For parts that existing templates can satisfy, first confirm the names and versions on the Image Templates page and include them in the checklist, so the other party can evaluate more accurately.
5. Reservation Period and Utilization: Is the Load Continuous or Bursty
Reserved resources are usually based on relatively certain long-term usage. Before submitting a requirement, estimate two things:
- Load type: If it is a constant load running at full capacity around the clock, a fixed reservation is suitable; if the usual volume is small with occasional peaks, ask whether the portion exceeding the reservation can be temporarily scaled on demand, and how that portion is billed;
- Expected utilization: If reserved GPUs sit idle most of the time, you are paying for idle resources. In this case, reserving only the baseline and leaving peaks to on-demand rental is often more suitable.
6. Time, Scale, and Region
Clearly state the planned start date, duration, total GPU count, whether to roll out in batches, and whether there are hard requirements on data center location.
A Consultation Template You Can Directly Adapt
【用途】例:70B 级模型全参微调 / 在线推理服务
【模型与并行】参数规模、单卡显存最低要求、并行策略
【卡型与数量】期望卡型(可接受的替代卡型)、总卡数、单机卡数
【互联】是否需要 NVLink;是否需要跨机 IB/RoCE,期望带宽
【存储】数据集大小、读取吞吐、检查点频率与大小、是否需要共享存储
【环境】使用的模板(PyTorch/vLLM/TGI 等)或私有镜像,CUDA 版本
【周期与负载】开始日期、预留时长、恒定负载还是突发负载、预期利用率
【地域】是否需要指定机房或地域
【联系人】姓名、团队、方便沟通的时间It's okay if some items are uncertain. Writing "TBD" or giving a range is better than leaving them blank. When the other party sees the range, they can first provide several options.
Things to Clarify During Communication
After submitting the requirement, confirm the following items one by one and try to get written responses:
- Available nodes and delivery time: Are there nodes that meet the requirements now, how long will delivery take, and can delivery be split into batches;
- Whether interconnect matches the description: Type and bandwidth of cross-machine networking, and whether communication tests can be run before acceptance after delivery;
- Billing details: How the reserved portion is priced, how the overage is priced, and whether storage and traffic are billed separately under standard terms;
- Commercial terms: Prepayment ratio, early termination and refund rules, renewal method. These are subject to the actual signed commercial agreement;
- Support method: Whom to contact when faults occur, and what the response channels are.
While Waiting for a Reply, You Can First Run a Small-Scale Trial
The most easily misestimated items in the requirements checklist are VRAM, throughput, and utilization. While waiting for scheduling, you can first rent a few GPUs self-service and run a round with the same image and a small portion of data. Measure actual VRAM usage, throughput per second, and checkpoint write time, then add these numbers to the checklist. Subsequent estimates will be much more accurate. Self-service instances are metered by the second and have no contract, so remember to destroy them after the trial to stop all charges. To see which GPU models and nodes are currently available, check the Pricing and Available Nodes page.
Once the checklist is organized, submit it through the Multi-GPU and Enterprise Solutions page, or contact the Telegram bilingual Chinese-English customer service. They will confirm available nodes, delivery time, and billing details based on your checklist.
NexGPU-算力租赁,GPU服务器,GPU云算力,AI服务器租用-新闻博客
Comments(0)