DocsMaple Proxy

API reference

Call Maple Proxy with OpenAI-compatible requests. A short reference for chat, embeddings, model discovery, streaming and the limits that matter to clients.

Set your tool’s OpenAI base URL to http://127.0.0.1:8080/v1, use a real Maple API key, and choose a model ID. This also applies to custom OpenAI-compatible providers in LangChain, LlamaIndex and similar tools.

Supported endpoints

MethodPathUse
POST/v1/chat/completionsChat, including streaming and model-supported tool/image inputs
POST/v1/embeddingsText embeddings with nomic-embed-text
GET/v1/modelsModel catalog and capability metadata
GET/health (or /)Local process liveness and version; no key needed

These endpoints are available in the released proxy 0.4.1. Responses, Assistants, audio, image-generation and file-upload APIs are not supported. A healthy /health response does not check your key or the Maple connection.

Chat completions

Create a key in Maple, then set it in the shell where you run the request:

Shell
export MAPLE_API_KEY=your-maple-api-key
curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $MAPLE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [{"role": "user", "content": "Write a haiku about privacy"}],
    "reasoning_effort": "medium",
    "stream": true
  }'

Set stream to false for a single JSON response. Streaming uses Server-Sent Events. Models that expose reasoning return it in delta.reasoning while streaming or message.reasoning in a complete response; answer text is in content.

Supported thinking levels and image/tool inputs depend on the model. See models and capabilities. An OpenAI-compatible request format does not mean every OpenAI feature or parameter is supported.

Embeddings

Shell
curl http://127.0.0.1:8080/v1/embeddings \
  -H "Authorization: Bearer $MAPLE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nomic-embed-text","input":"Text to search later"}'

Embeddings use input tokens only. API usage and plan pricing are separate; see rates.

Limits and failures

Request bodies are limited to 50 MiB. The model’s token/context limit is separate. Timeouts default to 300 seconds for a response to start or a non-streaming response to finish, and 300 seconds between streaming chunks. Change them in proxy settings.

Check HTTP status before parsing an error: some failures have an OpenAI-style JSON error, while routing and oversized-body failures can have an empty or plain-text body. A stream can fail after its response starts. Use troubleshooting for keys, quotas, model errors and timeouts.

Last updated