Home / APIs & tokens / Cerebras / Tokens & top-ups

Cerebras tokens & top-ups

Cerebras

prices arrive at launch

About the service

Cerebras Inference runs open models on wafer-scale chips and delivers some of the highest token generation speeds available. Llama, Qwen, GPT-OSS and other models are exposed through an OpenAI-compatible API with Python and JS SDKs. It matters most for agent chains and reasoning models, where long outputs otherwise cost seconds of waiting.