Support
A demo request page connects teams with an AI engineer. Documentation is available for MAX, and the site carries a blog, case studies and release notes.
Learn moreUnified AI inference stack from GPU kernels to cloud serving

Modular is an AI inference stack spanning custom GPU kernels to cloud serving on NVIDIA and AMD hardware. It offers a managed cloud for open models, in-your-cloud deployment, and an open source self-hosted option for teams running demanding inference workloads.
Modular is an inference platform that covers the whole path from low-level compute to production serving. It pairs an open source stack (MAX for model serving, Mojo as a programming language) with a managed cloud that hosts current open models, and the company positions it for running AI across GPUs, CPUs and other accelerators, including NVIDIA and AMD hardware. The homepage cites more than 15 supported CPU and GPU architectures. The product is aimed at engineering teams that run inference at scale and care about latency and cost per token. Customer stories on the site show the pattern: MiniMax describes serving its M3 model in production at large scale, Inworld reports a 70% improvement in time to first audio, Hippocratic AI cites sub-500ms time to first token for real-time patient conversations along with 70% total cost savings, and TensorWave credits Modular with savings of up to 70% on AMD compute. AWS is quoted as a partner bringing the MAX platform to its customers. In practice, a team can choose among several deployment routes. Modular's own cloud serves popular open models such as MiniMax M3, DeepSeek V3.2 and Kimi K2 coder variants. A second route runs the stack inside the customer's cloud account, and a third is self-hosted open source. A custom model path exists for teams that bring their own weights. Solution pages cover code generation, image generation, audio and agentic workloads. Signing up goes through a console, and a demo can be booked with an AI engineer. Among alternatives, Modular sits closer to inference engines and serving platforms than to model APIs or end-user AI apps. Its distinguishing claim is vertical integration, with kernels, compiler, runtime and serving layer built together, plus hardware portability across vendors. The company has also been acquired by Qualcomm, according to its own blog, which is relevant context for buyers weighing long-term roadmap and vendor independence.




A demo request page connects teams with an AI engineer. Documentation is available for MAX, and the site carries a blog, case studies and release notes.
Learn moreRuns across NVIDIA and AMD GPUs plus CPUs and other accelerators (15+ architectures). Partners and logos shown include AWS, Arm, Intel and Microsoft. Hosts open models such as MiniMax M3, DeepSeek V3.2 and Kimi K2 variants.
Unified inference stack from custom GPU kernels to cloud serving on NVIDIA and AMD, with support for GPUs, CPUs and ASICs. Includes the MAX serving framework, the Mojo language, hosted open models, custom model support, and deployment in Modular's cloud, your cloud or self-hosted.
Learn more




