Chat completions
POST /v1/chat/completions through Maple Proxy. What differs from OpenAI's API, including streaming, extra parameters, size limits, timeouts and errors.
POST /v1/chat/completions Creates a chat completion. Send the same request body you would send to OpenAI’s Chat Completions API. This page covers only what is different through Maple Proxy.
curl -N http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $MAPLE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Write a haiku about privacy"}
],
"stream": true
}' What differs from OpenAI
The request is passed through unchanged
The proxy doesn’t parse or rewrite the request or response body. Fields are sent to Maple exactly as your client wrote them, including provider-specific fields that OpenAI doesn’t define (for example chat_template_kwargs). Whether a model honours a field depends on the model, not the proxy.
The query string and most request headers are forwarded too. The proxy removes Authorization (your key travels inside the encrypted envelope instead), cookies, hop-by-hop headers such as Connection, and forwarding headers such as X-Forwarded-For.
Models
Use Maple model IDs, not OpenAI ones. See models.
Streaming and non-streaming
Both are supported. Set "stream": true for Server-Sent Events, or "stream": false (or leave it out) for one JSON response. Streamed responses are forwarded byte for byte as they arrive.
Request size
Request bodies over 50 MiB are rejected with 413 Payload Too Large. That limit is large enough for image inputs from coding agents.
Timeouts
- Non-streaming: the whole response must arrive within
MAPLE_REQUEST_TIMEOUT_SECS(default 300 seconds). - Streaming: the stream must start within that time. After that, each new chunk must arrive within
MAPLE_STREAM_IDLE_TIMEOUT_SECS(default 300 seconds). There is no limit on the total length of a stream.
A timeout before the response starts returns 504. A timeout after the response has started ends the body with an error. The proxy doesn’t retry timed-out requests. See configuration.
Errors
Errors created by the proxy use OpenAI’s error shape:
{
"error": {
"message": "Failed to communicate securely with the Maple backend",
"type": "server_error",
"param": null,
"code": null
}
} | Status | From | Meaning |
|---|---|---|
401 | Proxy | No API key. See authentication. |
413 | Proxy | Request body over 50 MiB. |
502 | Proxy | The proxy couldn’t verify or talk to the Maple backend. |
504 | Proxy | The backend didn’t start responding in time. |
| Anything else | Maple | Passed through unchanged, with its body and safe headers. |
A 401 from Maple (not the proxy) usually means the key is invalid or has been deleted. A 400 usually means an old model ID. More in troubleshooting.