Pricing
GPU compute is priced per hour and listed on the homepage, with examples such as H100 at $2.43/hr (spot $0.94/hr), B200 at $3.49/hr and B300 at $4.99/hr. Pricing for training, inference and dedicated capacity is not stated.
Learn moreTrain, evaluate and serve your own models on one open stack

Prime Intellect is a model-training platform that bundles GPU compute, RL environments, hosted training, evaluations, OpenAI-compatible inference and sandboxes, aimed at teams building their own task-specific and agentic models.
Prime Intellect is an integrated stack for teams that want to build their own models rather than rent a frontier API. It combines GPU compute, reinforcement learning environments, hosted training, evaluations, inference and code sandboxes behind a single account and a Python command line tool installed with pip. The pitch is ownership: a company trains a small, task-specific model on its own workflow and keeps improving it. The workflow starts with environments. Any task can be turned into an RL environment using the Prime CLI, with a loop of init, develop, eval and push, built on the open-source Verifiers library. A public Environment Hub lists more than 2,500 community environments that can be browsed, starred and reused. Hosted evaluations let users benchmark more than 100 open-source models without setting up infrastructure, and a public leaderboard shows the results. Hosted Training then runs large-scale training optimized for agentic workflows, with managed workflows that expose full visibility and control, plus support from the company's applied research team. For people who prefer to run things themselves, the Prime-RL framework handles asynchronous reinforcement learning at scale, and sandboxes provide secure code execution tuned for large RL runs. Inference and compute round out the offering. Prime hosts GLM-5.3 on its own infrastructure, and a gateway reaches third-party models through one OpenAI-compatible API. Dedicated serving capacity can be arranged for specific latency or reliability needs. On the compute side, users can rent from one to 256 GPUs on demand, with listed hourly prices for H100, H200, B200 and B300 hardware, SLURM and Kubernetes orchestration, Infiniband networking and Grafana dashboards. The platform suits ML engineers, research teams and AI product groups with enough technical depth to define rewards and evaluations. Customer examples on the homepage include Ramp, which trained a small subagent for spreadsheet questions, and Zapier, which uses evals as improvement loops. Compared with a plain GPU marketplace or a closed fine-tuning API, it bundles more of the loop, but it assumes comfort with reinforcement learning concepts.




GPU compute is priced per hour and listed on the homepage, with examples such as H100 at $2.43/hr (spot $0.94/hr), B200 at $3.49/hr and B300 at $4.99/hr. Pricing for training, inference and dedicated capacity is not stated.
Learn moreHosted Training includes hands-on support from the applied research team, and calls or demos can be booked through the contact page.
Learn moreOffers an OpenAI-compatible API, a gateway to third-party model providers, open-source Verifiers and Prime-RL on GitHub, and SLURM and Kubernetes orchestration for compute.
Learn moreRL environments via the Prime CLI and Verifiers, a 2,500+ environment Hub, hosted evaluations, Hosted Training, Prime-RL, secure sandboxes, OpenAI-compatible inference with a gateway, and on-demand GPUs with SLURM, Kubernetes, Infiniband and Grafana.





