DocsMaple Proxy

Chat completions

POST /v1/chat/completions through Maple Proxy. What differs from OpenAI's API, including streaming, extra parameters, size limits, timeouts and errors.

HTTP
POST /v1/chat/completions

Creates a chat completion. Send the same request body you would send to OpenAI’s Chat Completions API. This page covers only what is different through Maple Proxy.

Shell
curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $MAPLE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Write a haiku about privacy"}
    ],
    "stream": true
  }'

What differs from OpenAI

The request is passed through unchanged

The proxy doesn’t parse or rewrite the request or response body. Fields are sent to Maple exactly as your client wrote them, including provider-specific fields that OpenAI doesn’t define (for example chat_template_kwargs). Whether a model honours a field depends on the model, not the proxy.

The query string and most request headers are forwarded too. The proxy removes Authorization (your key travels inside the encrypted envelope instead), cookies, hop-by-hop headers such as Connection, and forwarding headers such as X-Forwarded-For.

Models

Use Maple model IDs, not OpenAI ones. See models.

Streaming and non-streaming

Both are supported. Set "stream": true for Server-Sent Events, or "stream": false (or leave it out) for one JSON response. Streamed responses are forwarded byte for byte as they arrive.

Request size

Request bodies over 50 MiB are rejected with 413 Payload Too Large. That limit is large enough for image inputs from coding agents.

Timeouts

  • Non-streaming: the whole response must arrive within MAPLE_REQUEST_TIMEOUT_SECS (default 300 seconds).
  • Streaming: the stream must start within that time. After that, each new chunk must arrive within MAPLE_STREAM_IDLE_TIMEOUT_SECS (default 300 seconds). There is no limit on the total length of a stream.

A timeout before the response starts returns 504. A timeout after the response has started ends the body with an error. The proxy doesn’t retry timed-out requests. See configuration.

Errors

Errors created by the proxy use OpenAI’s error shape:

JSON
{
  "error": {
    "message": "Failed to communicate securely with the Maple backend",
    "type": "server_error",
    "param": null,
    "code": null
  }
}
StatusFromMeaning
401ProxyNo API key. See authentication.
413ProxyRequest body over 50 MiB.
502ProxyThe proxy couldn’t verify or talk to the Maple backend.
504ProxyThe backend didn’t start responding in time.
Anything elseMaplePassed through unchanged, with its body and safe headers.

A 401 from Maple (not the proxy) usually means the key is invalid or has been deleted. A 400 usually means an old model ID. More in troubleshooting.

Last updated