Deploy LLaMA 4
Text & ChatLLaMA 4 is an open-weight model family from Meta. Scout uses a mixture-of-experts architecture and supports multimodal inputs. This catalog also includes LLaMA 3.3 70B for text applications; context and input support depend on the deployment.
Set up a LLaMA 4 workspace
Starting at $0.72/hr on dedicated GPU
Model configurations (2)
| Model | GPU | VRAM | Price | Action |
|---|---|---|---|---|
LLaMA 4 Scout Scout (109B MoE) | A100 80GB PCIe | 80 GB | $1.85/hr | Deploy |
LLaMA 3.3 70B Large (70B) | RTX A6000 | 48 GB | $0.72/hr | Deploy |
Prices include the service fee. Charges follow actual running time.
Requirements
The listed configuration specifies 48–80GB cloud GPUs across the listed variants. Actual local VRAM requirements vary with model variant, precision, quantization, resolution, and workflow settings.
On ModelPilot, deploy on a dedicated cloud GPU (up to 80GB VRAM) starting at $0.72/hr. Review the template’s required setup steps and settings before launch.
Use Cases
- General-purpose AI assistants
- Long-context document processing
- Multimodal understanding
- Enterprise AI applications
Related Models
Frequently Asked Questions
How much GPU memory is allocated for LLaMA 4?
The listed ModelPilot variants use 48–80GB cloud GPUs. Local memory needs vary with the variant, precision, quantization, and workflow settings.
How much does it cost to run LLaMA 4?
Starting at $0.72/hr on a dedicated GPU. Charges are calculated from actual running time, with auto-stop when credits run out.
How long does LLaMA 4 take to deploy?
Startup time depends on model downloads, container setup, cached files and GPU availability. Review the estimate for your selected template; it is not a guaranteed time to a finished result. Dedicated GPU time is billed while allocated, including startup.
Can I run LLaMA 4 on my local GPU?
It depends on the selected variant, precision, quantization, and workflow settings. Compare the variants below with your available VRAM; the table shows ModelPilot's cloud GPU allocation, not a universal local minimum.
Ready to deploy LLaMA 4?
Choose a GPU, review the estimated rate, and check the template’s setup requirements before launching.