Low-cost serverless inference for open models
DeepInfra is a serverless inference provider that runs open-source language, image, audio and embedding models on its own infrastructure. Requests go to a simple REST endpoint, and pricing is per token or per input unit, with no capacity planning involved. The catalogue focuses on popular open weights - Llama, Qwen, Mistral, Stable Diffusion and Whisper among them - so you can move from a self-hosted setup to a paid endpoint without rewriting anything. The differentiator is price: it competes largely on per-token cost for standard models.
Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.
cost-sensitive inference, open-model apis, batch generation, embeddings on a budget
If you're comparing similar products, check the alternatives below, or browse all tools in the AI Models & Platforms category.
Popular open-weight language models like Llama, Qwen and Mistral, plus image, speech and embedding models. The catalogue favours widely used open releases rather than exclusive models.
It runs optimised inference on its own infrastructure and competes mainly on per-token price. There is no reserved capacity to pay for, so cost tracks usage directly.
Free trial credit is limited, and rate limits apply at low usage tiers. High-volume customers can arrange custom deployments or higher throughput commitments.
Developer console and API keys for the Claude model family
Node-based local interface for running image and video diffusion models
Browser playground for prompting and prototyping with Gemini models
Hosted Jupyter notebooks with optional free GPU and TPU runtimes