Models
Which models you can call through Maple Proxy, how to list them with /v1/models, and how per-model API rates are charged per million tokens.
Maple Proxy can call any model that Maple offers through its API. Use the model’s ID in the model field of a request.
Get the current list
The model list changes as Maple adds and retires models. Ask the proxy for the list your key can use:
curl http://127.0.0.1:8080/v1/models \
-H "Authorization: Bearer $MAPLE_API_KEY" The response comes straight from Maple’s backend. See list models.
Chat models and rates
These are the models on Maple’s public rate table, with the same rates. Rates are in US dollars per million tokens.
| Model ID | Name | Input | Cached input | Output |
|---|---|---|---|---|
glm-5-3-flash | GLM-5.3 Flash | $0.80 | $0.20 | $2.50 |
glm-5-3 | GLM 5.3 | $3.60 | $0.90 | $11.50 |
kimi-k3 | Kimi K3 | $8.00 | $1.60 | $25.00 |
glm-5-2 | GLM 5.2 | $3.00 | $0.75 | $10.50 |
deepseek-v4-1-flash | DeepSeek V4.1 Flash | $1.30 | $0.26 | $2.90 |
kimi-k2-6 | Kimi K2.6 | $3.59 | $0.35 | $17.94 |
gpt-oss-120b | OpenAI GPT-OSS 120B | $0.30 | Standard input price applies | $1.20 |
gemma4-31b | Gemma 4 31B | $0.80 | Standard input price applies | $2.00 |
llama3-3-70b | Llama 3.3 70B | $3.50 | Standard input price applies | $5.50 |
gpt-oss-safeguard-120b | OpenAI GPT-OSS Safeguard 120B (API only) | $0.30 | Standard input price applies | $1.20 |
USD equivalent per 1 million tokens. Cached-input rates apply only to tokens reported as cached. — means the standard input price applies. Automatic model aliases are billed at the rate of the model selected for the request.
The plan price and API rates are separate: API usage is charged at the rate of the model used. See pricing for plans.
Embedding models
For /v1/embeddings, use nomic-embed-text. It is Maple’s embedding model and the default when a request names none.
Embeddings are billed on input tokens only; there are no output tokens. The embedding rate isn’t on the public rate table.
Model IDs to avoid
Use the IDs exactly as /v1/models returns them. Older guides used names such as llama-3.3-70b; the ID is llama3-3-70b. A request with an unknown or retired model ID usually fails with 400.