Пополнение баланса
- Open models on wafer-scale hardware with record output speed
- Pay-per-token usage billing
- OpenAI-compatible API and official SDKs
Home / APIs & tokens / Cerebras
Cerebras
Plan contents as published by the vendor; seller prices arrive at launch.
Cerebras Inference runs open models on wafer-scale chips and delivers some of the highest token generation speeds available. Llama, Qwen, GPT-OSS and other models are exposed through an OpenAI-compatible API with Python and JS SDKs. It matters most for agent chains and reasoning models, where long outputs otherwise cost seconds of waiting.