DeepInfra

Pay-as-you-go inference API for 100+ open models and rentable GPUs

Advanced API
Screenshot of DeepInfra, Pay-as-you-go inference API for 100+ open models and rentable GPUs

What is DeepInfra?

DeepInfra is a pay-as-you-go inference cloud with APIs for 100+ open models across text, speech, image, video and embeddings, plus on-demand GPU rental. It stresses low cost, zero data retention, and SOC 2 and ISO 27001 compliance.

DeepInfra is an inference cloud that serves machine-learning models through developer-facing APIs. Instead of provisioning GPUs and tuning serving stacks, a team picks a hosted model, calls it over an API and pays for the tokens or compute it uses. The catalog covers more than 100 models across text generation, embeddings, rerankers, speech recognition, text to speech, text to image, text to video, text to music, world models and zero-shot image classification. The text-generation shelf is the most visible part of the homepage. It lists large open-weight families from DeepSeek, Moonshot AI (Kimi), Qwen, Z.ai (GLM), Xiaomi (MiMo), Tencent and NVIDIA, each with per-million-token input and output rates. Some listings add a cached-input rate, a context window of up to roughly one million tokens, a precision tag such as fp8 or fp4, and a zero data retention marker. Temporary discounts are flagged on certain models. Beyond hosted models, the site offers on-demand GPU rental (a DGX B300 instance is listed by the instance-hour), a DeepCluster offering, a DeepStart program and a browser chat page for trying models. The audience is developers, startups and enterprises that want to run open models in production without owning hardware. The pitch centers on low cost, speed, simplicity and reliability, with scale from small experiments to very high token volumes and hands-on technical support. The company says it runs on its own inference-optimised hardware in US-based data centers, holds SOC 2 and ISO 27001 certifications, and applies a zero retention policy to inputs, outputs and user data. It also announced a $107M Series B in May 2026. In practice, DeepInfra sits among open-model inference providers and GPU clouds. It suits teams comparing per-token prices across many model families, or those who want to swap models without changing vendors. Buyers who need a single proprietary frontier model, a visual no-code builder or a fixed monthly plan will find it less direct, since the product is built around APIs and usage-based billing.

How do you use DeepInfra?

  1. 1Browse the model catalog
    Open the models page and filter by task, such as text generation, embeddings or text to image, to shortlist candidates and compare per-million-token rates.
    DeepInfra — Browse the model catalog
  2. 2Try a model in the chat
    Use the browser chat to test a text model's quality and tone before writing any code.
    DeepInfra — Try a model in the chat
  3. 3Create an account and open the dashboard
    Sign up or log in through the dashboard, where API access and usage are managed.
    DeepInfra — Create an account and open the dashboard
  4. 4Follow the docs to make a first API call
    Read the documentation for authentication and request format, then send a first request to your chosen model from your own application.
    DeepInfra — Follow the docs to make a first API call
  5. 5Review pricing and scale up
    Check the pricing page to estimate costs, and contact sales if you need dedicated capacity, GPU instances or a custom setup.

Pros and cons

Pros

  • Wide catalog of 100+ models spanning text, speech, image, video, music and embeddings behind one APIAI
  • Pay-as-you-go pricing with no long-term contracts or hidden fees, and per-token rates shown publiclyAI
  • Zero retention policy plus SOC 2 and ISO 27001 certification suit privacy-sensitive workloadsAI
  • Runs on its own inference-optimised hardware in US data centers, with GPU rental available as wellAI
  • Fast access to new open-weight releases, with large context windows and cached-input rates on many modelsAI

Cons

  • Aimed at developers; the API-first workflow offers little for non-technical usersAI
  • Catalog is open-weight models, so proprietary frontier models are not offeredAI
  • Usage-based billing makes monthly costs harder to predict than a flat planAI
  • Hosting is described only as US-based data centers, with no regional choice mentioned on the homepageAI
  • Per-token rates vary widely between models and change with promotions, so cost comparison takes effortAI

How much does DeepInfra cost?

Pricing

Pay-as-you-go with no long-term contracts or hidden fees. Models are priced per million input and output tokens (for example DeepSeek-V4-Flash at $0.09 in and $0.18 out), some with cached-input rates. On-demand DGX B300 GPUs are listed at $4.89 per instance-hour.

Learn more

Support

Hands-on technical support is mentioned, along with a documentation site, a contact-sales page for consultations and a trust center for security information.

Learn more

Features

APIs for 100+ models across text generation, embeddings, reranking, speech recognition, text to speech, text to image, text to video, text to music and world models. Also on-demand GPU instances, a browser chat, zero data retention and long context windows on many models.

Learn more

Frequently asked questions about DeepInfra

  • How much does DeepInfra cost?
    Pay-as-you-go with no long-term contracts or hidden fees. Models are priced per million input and output tokens (for example DeepSeek-V4-Flash at $0.09 in and $0.18 out), some with cached-input rates. On-demand DGX B300 GPUs are listed at $4.89 per instance-hour.
  • Does DeepInfra have an API?
    Yes, DeepInfra offers an API.
  • How do you use DeepInfra?
    The walkthrough on this page covers 5 steps: 1. Browse the model catalog 2. Try a model in the chat 3. Create an account and open the dashboard 4. Follow the docs to make a first API call 5. Review pricing and scale up.
  • What platforms does DeepInfra support?
    DeepInfra is available on Web App.
  • What are the limitations of DeepInfra?
    Aimed at developers; the API-first workflow offers little for non-technical users. Catalog is open-weight models, so proprietary frontier models are not offered. Usage-based billing makes monthly costs harder to predict than a flat plan.

Status

StatusActive
Views0
Outbound clicks0
Added10/7/2026

Platforms

Web App

Pricing

Paid