Free trial
The in-browser demo for image generation and chat is available without an account. A startup program offers credit to accepted startups. API keys require a card.
Learn moreOne OpenAI-compatible key for open-model inference, agents and private cloud

FlexAI is a coding tool. FlexAI offers managed inference on 30+ open models through one OpenAI-compatible API key, with a growth path from serverless calls to agent governance, dedicated GPU endpoints and private cloud deployment.
FlexAI is a managed inference platform that serves more than 30 open-weight models through a single OpenAI-compatible API. The catalog spans text generation, coding, reasoning, vision, embeddings, speech-to-text, text-to-speech and image generation, with families from Qwen, Gemma, Llama, DeepSeek, GLM, Mistral and OpenAI's open GPT-OSS models. Existing code written for the OpenAI SDK can be pointed at the FlexAI endpoint and run unchanged, and the vendor states that leaving again is a configuration change. The product is organised as a path rather than a single service. Token Factory covers serverless model calls priced by usage. AgentOS is the layer for agent loops, where teams bring skills, tools, scoped memory, evals and approvals, and FlexAI runs them across models with routing, governance and audit trails. Dedicated Endpoints move the same API onto reserved GPUs, and AI Factory deploys the stack in a VPC, on-premises or air-gapped. Compute runs on FlexAI-operated NVIDIA and AMD fleets, with an uptime SLA of up to 99.9 percent, and the company lists SOC 2 Type II and GDPR compliance. In practice, a developer can try image generation and chat completions in an in-browser demo without an account, then sign up for an API key. Billing is pay as you go with no commitment beyond usage, though a card is required to create a key. Text models are charged per token, while media models are charged per image, per minute of audio or per generated clip, and dedicated capacity is billed per GPU-hour. FlexAI suits agent-focused product teams, startups that expect to outgrow serverless, and regulated organisations that need private deployment. Compared with model API providers such as Fireworks or hyperscaler services such as AWS Bedrock, its pitch is a single account that scales from serverless calls to private infrastructure. It is limited to open-weight models, so teams that depend on proprietary frontier models will still need another provider alongside it.




The in-browser demo for image generation and chat is available without an account. A startup program offers credit to accepted startups. API keys require a card.
Learn morePay as you go with no commitment beyond usage. Serverless is priced per token for text and per image, speech character, audio minute or video clip for media. Dedicated endpoints are billed per GPU-hour and AI Factory per deployment.
Learn moreA public status page tracks uptime, and a security site covers compliance. A contact route exists for sales conversations.
Learn moreWorks with the OpenAI SDK by pointing it at https://api.flex.ai/v1. Runs on NVIDIA and AMD GPU fleets.
30+ open models behind one OpenAI-compatible API, with tool calls, streaming, structured output and vision. AgentOS adds routing, approvals and audit trails. Dedicated endpoints, fine-tuning and private AI cloud (VPC, on-prem, air-gapped) are available.
Learn more




