GPU Cloud Pricing in India 2026: H100, H200, B200 and A100 Rates
GPU pricing in India is difficult to compare because one page may show a marketplace spot rate, another may show an India-hosted on-demand quote, and a third may show a reserved multi-GPU server. This guide puts those categories next to each other and explains what the number does and does not include.
The table below is a dated snapshot from GPUIndia's provider feed. It is useful for planning, not a guaranteed quote. Open the live GPU rental comparison before you commit because capacity and rates move daily.
Compare today's provider rates
Filter by GPU, billing mode and provider to see the current cheapest options.
Open the live GPU comparisonGPU cloud price snapshot for India
These are the cheapest tracked on-demand rates in the 14 September 2026 feed. INR is an approximate conversion at ₹83 per US dollar. A global provider can still be the cheapest option for an Indian buyer, while an India-hosted provider may win on latency, invoicing or data residency.
| GPU | Cheapest on-demand | Approx. ₹/hr | Spot when listed | Typical fit |
|---|---|---|---|---|
| A100 80GB | $1.67/hr | ₹139 | $0.61/hr | Fine-tuning and value inference |
| H100 | $2.50/hr | ₹208 | $0.99/hr | 70B-class training and inference |
| H200 | $2.80/hr | ₹232 | $2.00/hr | Memory-heavy and long-context jobs |
| B200 | $4.50/hr | ₹374 | $2.12/hr | Large-model throughput |
| RTX 5090 | $0.58/hr | ₹48 | Not listed | Small models and experiments |
The cheapest hourly rate is not automatically the cheapest workload. A GPU that cannot fit the model is not a bargain, and a slower card may need more hours to serve the same token volume. Use the workload cost calculator to compare cost per 1M tokens after VRAM fit.
What does a GPU cost per month?
A simple full-time estimate is hourly rate × 730 hours. At the snapshot rates, one month of the cheapest tracked on-demand capacity is approximately:
| GPU | Monthly compute | Approx. INR | Important limitation |
|---|---|---|---|
| A100 80GB | $1,219 | ₹1.01 lakh | Storage, egress and tax may be extra |
| H100 | $1,825 | ₹1.51 lakh | Cheapest marketplace rate, not every region |
| H200 | $2,044 | ₹1.70 lakh | Capacity is thinner than H100 |
| B200 | $3,285 | ₹2.73 lakh | Only a small number of providers list it |
Do not multiply a spot price by 730 unless you can tolerate interruptions and have checked the provider's actual reclaim policy. For a production endpoint, a higher on-demand rate with predictable capacity can be cheaper than repeated migrations and downtime.
Marketplace rate vs IndiaAI reference pricing
The public IndiaAI compute price list is a useful reference, but it lists different instance configurations and billing terms. It shows separate on-demand and reserved prices for H100, H200, B200 and other accelerators. That means it should be used as a benchmark, not merged into a marketplace table without checking GPU count, region, storage and reservation period.
For a buyer, the practical comparison is:
- Marketplace spot: lowest visible price, but preemption and supply risk.
- Marketplace on-demand: flexible, with provider and region differences.
- India-hosted or IndiaAI capacity: potentially better for INR billing, latency and residency, but configuration and access rules matter.
- Reserved capacity: lower effective rate when usage is predictable, with less flexibility.
Which GPU should an Indian team choose?
Choose A100 when cost matters and 80GB is enough
A100 remains a strong value option for fine-tuning, batch inference and workloads that do not need H100-class throughput. The lower hourly rate is useful when utilization is moderate and the job is not latency-sensitive.
Choose H100 for the default production path
H100 is the safe baseline for many 70B-class workloads because the software ecosystem is mature and the tracked supply is deeper than H200 or B200. Start with spot for checkpointed work, then move to on-demand when the workload becomes business-critical.
Choose H200 when the model is memory-bound
H200's 141GB of HBM3e and 4.8TB/s memory bandwidth make it useful when an 80GB H100 forces sharding or leaves too little headroom. The premium only makes sense when the extra memory changes the deployment shape or throughput.
Choose B200 when throughput is the bottleneck
B200 is the premium option for large-model training and inference. It is harder to find and costs more, so use it when the faster path reduces total job time or lets you use fewer GPUs—not simply because it is the newest card.
Hidden costs people miss
Compute is only one line in a real GPU budget. Add storage for model weights and checkpoints, egress for user traffic, CPU and RAM, networking, taxes, minimum billing, persistent disks and the engineering time required to recover a preempted job. The GPU TCO calculator lets you model these assumptions instead of treating the hourly sticker as the final cost.
For owned hardware, add electricity and cooling. The India electricity calculator estimates the monthly power bill from GPU draw, utilization, hours per day and state tariff. Cloud rentals usually bundle power, while dedicated and colocation deals may not.
Bottom line
For most Indian teams, H100 is the practical starting point, A100 is the value option, H200 is the memory upgrade and B200 is the throughput upgrade. Compare the workload first, then the hourly rate, then the contract and region. A dated, provider-level table is more useful than a single “cheapest GPU” number.
Sources and method: GPUIndia provider feed snapshot dated 14 September 2026; INR examples use ₹83/USD for readability. See the IndiaAI public compute price list for its own instance configurations and reservation terms. NVIDIA's H200 specifications explain the memory and bandwidth difference. Rates change daily and should be confirmed with the provider.