DeepInfra

Low-cost serverless inference for open models

AI Models & PlatformsFree planOverseasβ˜…β˜…β˜…β˜†β˜† 3.0

What is DeepInfra?

DeepInfra is a serverless inference provider that runs open-source language, image, audio and embedding models on its own infrastructure. Requests go to a simple REST endpoint, and pricing is per token or per input unit, with no capacity planning involved. The catalogue focuses on popular open weights - Llama, Qwen, Mistral, Stable Diffusion and Whisper among them - so you can move from a self-hosted setup to a paid endpoint without rewriting anything. The differentiator is price: it competes largely on per-token cost for standard models.

Last updated: 2026-09-20. This site only provides an index; for exact features, pricing, and licensing, see the official website.

Key features

  • Serverless text, image, audio and embedding endpoints
  • OpenAI-compatible chat completions API for easy migration
  • Pay-per-use billing with no reserved capacity
  • Text-to-speech and speech-to-text models available
  • Simple REST calls plus Python and Node clients
  • Custom model deployment options for larger accounts

Pros & cons

Strengths

  • Per-token pricing is among the lowest for open models
  • No infrastructure or GPU management required
  • Broad model catalogue across several modalities

Watch out for

  • Latency varies more than on dedicated endpoints
  • Fewer governance features than enterprise clouds
  • Model selection depends on what the team has loaded

Best for & use cases

cost-sensitive inference, open-model apis, batch generation, embeddings on a budget

If you're comparing similar products, check the alternatives below, or browse all tools in the AI Models & Platforms category.

FAQ

What models does DeepInfra host?

Popular open-weight language models like Llama, Qwen and Mistral, plus image, speech and embedding models. The catalogue favours widely used open releases rather than exclusive models.

Why is it cheaper than some providers?

It runs optimised inference on its own infrastructure and competes mainly on per-token price. There is no reserved capacity to pay for, so cost tracks usage directly.

Are there usage limits?

Free trial credit is limited, and rate limits apply at low usage tiers. High-volume customers can arrange custom deployments or higher throughput commitments.