Pricing
Billing is per token. Elastic endpoints give eligible models dedicated GPU performance with no GPU-hour commitment or idle-GPU bill, priced at 25% above the model's Serverless input and output token rates. Specific Serverless rates are not listed on the homepage.
Learn more









