Skip to main content

Pricing

GPU time and API requests

Explore for free. When you’re ready to run, choose the setup that fits your work and review the rate before launching.

Credits pay for usage. Creating an account does not launch a machine.

Dedicated GPU Instances

Choose compute reserved for your workload. Running time, including startup and idle time, is billed. Stopped machines may retain billable storage.

CPU Only

$0.15/hour
VRAM
vCPU3
RAM16 GB

Small LLMs (1B-3B), basic inference, API servers, testing

RTX 4090

$0.99/hour
VRAM24 GB
vCPU6
RAM41 GB

7B-13B LLMs, SDXL Images, Flux, Most Popular GPU

L4

$0.66/hour
VRAM24 GB
vCPU12
RAM50 GB

7B-13B LLMs, SDXL Images, Power Efficient

RTX A6000

$0.72/hour
VRAM48 GB
vCPU9
RAM50 GB

30B+ LLMs (quantized 70B), Video Gen, 48GB VRAM

A100 80GB

$1.85/hour
VRAM80 GB
vCPU8
RAM117 GB

70B+ LLMs, Complex Workflows, High-VRAM Inference

H100

$4.32/hour
VRAM80 GB
vCPU16
RAM188 GB

Lowest-Latency Inference, Large Models, Complex Workflows

RTX 6000 Ada Generation

$1.12/hour
VRAM48 GB
vCPU12
RAM62 GB

Video Gen, MiniMax H3, FP8/INT8 quantized models, 48GB VRAM

L40S

$1.31/hour
VRAM48 GB
vCPU12
RAM62 GB

Video Gen, MiniMax H3, high-bandwidth inference, 48GB VRAM

RTX PRO 6000 Blackwell Server Edition

$2.76/hour
VRAM96 GB
vCPU16
RAM188 GB

MiniMax H3, NVFP4/FP8 quantized models, 2K video, 96GB VRAM

All prices include compute, memory, disk, and network.

Serverless — Pay Per Request

Review supported models and per-request prices. First requests may wait for startup; text chat needs a configured endpoint.

Text Models

ModelPer Request
Qwen3 4BQuick Response$0.014
Qwen3 8BGeneral Purpose$0.019
DeepSeek R1 8BReasoning$0.019
GLM-Z1 9BBilingual Reasoning$0.019
DeepSeek R1 14BAdvanced Reasoning$0.047
Qwen3 32BPremium Quality$0.138
GLM-Z1 32BPremium Reasoning$0.138

Image Models

ModelPer Request
Flux Schnell$0.017
Flux Dev$0.017
Stable Diffusion XL$0.014
Z-Image Turbo$0.014
Kokoro TTS$0.009 / request
Chatterbox Turbo$0.017 / request

Video runs use dedicated workflows. Explore video workflows

Generation APIs use prepaid credits per successful request. Text chat requires a running deployment or configured endpoint; availability is checked when a request runs.

Estimate Your Costs

4h
1h24h
5d
1d7d

Daily

$2.64

Monthly

$57

Before you begin

Billing questions

Need a detail specific to your setup?

Contact support
What can I do before paying?

Browse the model catalog, inspect available workflows, and use the workflow checker without an account. When available, the free demo lets you try a supported model within its usage limits.

How do dedicated machines differ from serverless?

A dedicated machine is reserved for your workload and is billed by time. Serverless runs supported models per request. Their available models, configuration options and prices differ; compare the tables above.

What happens when I finish using a machine?

Manage dedicated machines from your dashboard. Closing the browser does not remove a machine. Stopping and removing are different actions: stopped machines can retain storage, so review the resource details and remove resources you no longer need.

Is the calculator my final bill?

No. It estimates use at the listed rate. Setup, storage, availability and actual usage depend on your configuration. Review the deployment details before launching and your recorded usage in Billing afterward.

Do I need a subscription?

Dedicated deployments use prepaid credits. You can review the available credit amounts in your account before checking out.

Workflow guides

Start with a workflow.

See its requirements before choosing a machine.

Explore workflows Create an account