API reference
Call Maple Proxy with OpenAI-compatible requests. A short reference for chat, embeddings, model discovery, streaming and the limits that matter to clients.
Set your tool’s OpenAI base URL to http://127.0.0.1:8080/v1, use a real Maple API key, and choose a model ID. This also applies to custom OpenAI-compatible providers in LangChain, LlamaIndex and similar tools.
Supported endpoints
| Method | Path | Use |
|---|---|---|
POST | /v1/chat/completions | Chat, including streaming and model-supported tool/image inputs |
POST | /v1/embeddings | Text embeddings with nomic-embed-text |
GET | /v1/models | Model catalog and capability metadata |
GET | /health (or /) | Local process liveness and version; no key needed |
These endpoints are available in the released proxy 0.4.1. Responses, Assistants, audio, image-generation and file-upload APIs are not supported. A healthy /health response does not check your key or the Maple connection.
Chat completions
Create a key in Maple, then set it in the shell where you run the request:
export MAPLE_API_KEY=your-maple-api-key
curl -N http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $MAPLE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [{"role": "user", "content": "Write a haiku about privacy"}],
"reasoning_effort": "medium",
"stream": true
}' Set stream to false for a single JSON response. Streaming uses Server-Sent Events. Models that expose reasoning return it in delta.reasoning while streaming or message.reasoning in a complete response; answer text is in content.
Supported thinking levels and image/tool inputs depend on the model. See models and capabilities. An OpenAI-compatible request format does not mean every OpenAI feature or parameter is supported.
Embeddings
curl http://127.0.0.1:8080/v1/embeddings \
-H "Authorization: Bearer $MAPLE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"nomic-embed-text","input":"Text to search later"}' Embeddings use input tokens only. API usage and plan pricing are separate; see rates.
Limits and failures
Request bodies are limited to 50 MiB. The model’s token/context limit is separate. Timeouts default to 300 seconds for a response to start or a non-streaming response to finish, and 300 seconds between streaming chunks. Change them in proxy settings.
Check HTTP status before parsing an error: some failures have an OpenAI-style JSON error, while routing and oversized-body failures can have an empty or plain-text body. A stream can fail after its response starts. Use troubleshooting for keys, quotas, model errors and timeouts.