Simple, Transparent Pricing
Pay only for what you use. No monthly commitments.
Dedicated GPU
From $0.53/hr
Pay for actual running time
Serverless
From $0.008/req
No idle costs, pay per execution
Dedicated GPU Instances
Full GPU reserved for your workload. Charges follow actual running time, and you can stop anytime.
| GPU | VRAM | vCPU | RAM | Price/hr | Best For |
|---|---|---|---|---|---|
| CPU Only | — | 3 | 16 GB | $0.15/hour | Small LLMs (1B-3B), basic inference, API servers, testing |
| RTX 4090 | 24 GB | 6 | 41 GB | $0.92/hour | 7B-13B LLMs, SDXL Images, Flux, Most Popular GPU |
| L4 | 24 GB | 12 | 50 GB | $0.53/hour | 7B-13B LLMs, SDXL Images, Power Efficient |
| RTX A6000 | 48 GB | 9 | 50 GB | $0.66/hour | 30B+ LLMs (quantized 70B), Video Gen, 48GB VRAM |
| A100 80GB | 80 GB | 8 | 117 GB | $1.85/hour | 70B+ LLMs, Complex Workflows, High-VRAM Inference |
| H100 | 80 GB | 16 | 188 GB | $4.32/hour | Lowest-Latency Inference, Large Models, Complex Workflows |
CPU Only
$0.15/hourSmall LLMs (1B-3B), basic inference, API servers, testing
RTX 4090
$0.92/hour7B-13B LLMs, SDXL Images, Flux, Most Popular GPU
L4
$0.53/hour7B-13B LLMs, SDXL Images, Power Efficient
RTX A6000
$0.66/hour30B+ LLMs (quantized 70B), Video Gen, 48GB VRAM
A100 80GB
$1.85/hour70B+ LLMs, Complex Workflows, High-VRAM Inference
H100
$4.32/hourLowest-Latency Inference, Large Models, Complex Workflows
All prices include compute, memory, disk, and network.
Serverless — Pay Per Request
No idle costs. Your model scales to zero when not in use.
Text Models
| Model | Per Request |
|---|---|
| Qwen3 4BQuick Response | $0.010 |
| Qwen3 8BGeneral Purpose | $0.015 |
| DeepSeek R1 8BReasoning | $0.015 |
| GLM-Z1 9BBilingual Reasoning | $0.015 |
| DeepSeek R1 14BAdvanced Reasoning | $0.030 |
| Qwen3 32BPremium Quality | $0.050 |
| GLM-Z1 32BPremium Reasoning | $0.050 |
Image Models
| Model | Per Request |
|---|---|
| Flux Schnell | $0.008 |
| Flux Dev | $0.015 |
| Stable Diffusion XL | $0.005 |
| Z-Image Turbo | $0.008 |
No idle costs — pay only for the requests you run.
Estimate Your Costs
Daily
$2.12
Monthly
$46
Get Started Free
Try Before You Pay
Explore the free demo without a card. When you are ready to deploy, prepaid credits start at $5 with no subscription.
Frequently Asked Questions
- How does billing work?
- Dedicated GPU charges are calculated from actual running time — stop the instance when you are finished. Serverless is billed per request with no idle costs. No monthly minimums for either.
- Do I need a credit card to start?
- You can try the free demo without an account. To deploy your own GPU, sign up and add prepaid credits starting at $5.
- What if I forget to stop my GPU instance?
- Deployments don't auto-stop when idle — you stop them manually from the dashboard. As a safety net we do auto-stop when your credit balance reaches $0, so you'll never be charged beyond your prepaid balance.
- Can I use both dedicated GPU and serverless?
- Yes. Many teams use serverless for development and burst traffic, then switch to dedicated GPUs for sustained production workloads.
- What's included in the GPU price?
- Everything: GPU, vCPU, RAM, disk, and network. No hidden fees, no egress charges, no surprise bills.
- Which mode should I choose?
- Use serverless for low-volume, bursty, or development workloads (pay only when you generate). Use dedicated GPU for sustained production traffic where you need consistent latency and throughput.