Developers
API Guide
Use every Milly Lab model from your own code — in any language, on any platform, from a server or a browser — through an OpenAI-compatible API that bills through your Milly Lab account. This guide covers the first request, SDK setup, streaming, live pricing, the billing rules, per-key controls, rate limits, error codes and the usage endpoint.
#Base URL
https://aria-web-production-a38d.up.railway.app/v1
Every endpoint below is relative to this base. The interactive OpenAPI reference (/docs) is available on non-production deployments; the same operations are documented here.
#Quickstart
- Create a key. Open the web app → Settings → Security → API keys → Create key (the Manage API keys button on the API page opens it directly). API access is included in Max and higher plans. The raw key (
mk_…) is shown once; only its last four characters are stored. - Point an OpenAI SDK at the base URL. Any client that speaks the OpenAI chat-completions protocol works with a base-URL change.
- Pick a model from
GET /v1/modelsand send your first request.
curl https://aria-web-production-a38d.up.railway.app/v1/chat/completions \
-H "Authorization: Bearer $MILLY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "claude-sonnet-5",
"messages": [{"role": "user", "content": "Three facts about Dushanbe."}]}'
The response is a standard chat.completion object with choices[0].message.content and usage (the tokens you were billed for).
#SDK setup
Python
from openai import OpenAI
client = OpenAI(
base_url="https://aria-web-production-a38d.up.railway.app/v1",
api_key="mk_...",
)
r = client.chat.completions.create(
model="claude-sonnet-5",
messages=[{"role": "user", "content": "Three facts about Dushanbe."}],
)
print(r.choices[0].message.content, r.usage)
Node.js
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://aria-web-production-a38d.up.railway.app/v1",
apiKey: process.env.MILLY_API_KEY,
});
const r = await client.chat.completions.create({
model: "claude-sonnet-5",
messages: [{ role: "user", content: "Three facts about Dushanbe." }],
});
console.log(r.choices[0].message.content, r.usage);
Plain fetch (any runtime)
const res = await fetch("https://aria-web-production-a38d.up.railway.app/v1/chat/completions", {
method: "POST",
headers: { Authorization: "Bearer " + MILLY_API_KEY, "Content-Type": "application/json" },
body: JSON.stringify({ model: "claude-sonnet-5",
messages: [{ role: "user", content: "Three facts about Dushanbe." }] }),
});
const data = await res.json();
if (!res.ok) throw new Error(`${data.error.code}: ${data.error.message}`);
Supported request fields: model, messages (roles system/developer, user, assistant; string content or text parts), stream, max_tokens / max_completion_tokens, reasoning_effort (low · medium · high). Sampling fields (temperature, top_p, …) are accepted for compatibility; the platform router decides sampling. tools, tool_choice and tool/function messages are accepted but not executed yet — see Coming next.
#Streaming
Set stream: true to receive server-sent events. Each event is a chat.completion.chunk; the final chunk carries finish_reason and usage, then data: [DONE].
stream = client.chat.completions.create(model="claude-sonnet-5", stream=True,
messages=[{"role": "user", "content": "Write a haiku about mountains."}])
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
if chunk.usage:
print("\n", chunk.usage)
If your client disconnects mid-stream, generation still runs to completion on the platform and the tokens produced are billed — the provider has already been paid for them.
#Models and live pricing
GET /v1/models is public (no key required) and returns every API-enabled model with the exact rates the platform bills:
{
"id": "claude-sonnet-5",
"display_name": "Claude Sonnet 5",
"modality": "chat",
"capabilities": { "vision": true, "tools": true, "streaming": true },
"context_window": 1000000,
"max_output_tokens": 64000,
"pricing": {
"unit": "per_1m_tokens",
"input_per_1m_usd": 2.0,
"output_per_1m_usd": 10.0,
"platform_fee_per_1m_usd": 1.0,
"effective_input_per_1m_usd": 3.0,
"effective_output_per_1m_usd": 11.0
}
}
Image models carry pricing.unit = "per_image" with per_image_usd, platform_fee_per_image_usd and effective_per_image_usd. The embeddings model is listed under modality: "embedding". When you call /v1/models with a key, each row also carries allowed_for_key — whether that key may call the model — and included_in_plan — whether the account's plan allocates the model (allowance first) or every call is balance-funded at 2×; the top-level plan names the plan. On a $0 balance, a model that is not included answers 402 from the first token.
Only listed models are callable. A retired key with no successor (gpt-4o, gpt-4o-mini) answers 404 model_not_found; a legacy alias that points at a current model (claude-sonnet → claude-sonnet-5) still resolves and is billed at the listed rate.
The live table on the API page inside the web app renders this endpoint.
#Billing rules
- Cost per request = provider rate × tokens + platform fee of $1 per 1M input tokens and $1 per 1M output tokens (images: $0.01 per image; embeddings: provider rate + $1 per 1M input tokens).
- Plan first. Your plan's monthly allowance for that model is spent first.
- Then balance at 2×. Anything beyond the allowance is charged to your top-up balance at twice the cost (the same top-up markup as the web app).
- Reserve before spend. A request that cannot be covered by remaining allowance plus balance is refused with 402 before any provider call; the platform holds the worst-case amount during generation and settles to the real usage afterwards. If the model's provider is not configured on the deployment, the answer is 503
provider_unavailable— also before any hold, so nothing is billed. A provider that fails before producing output is 502 with nothing billed; output that did stream before a failure is billed. - Per-key controls (below) add a spend cap and a model allow-list on top of the account-level rules.
Every API request writes a usage row tagged with the key that made it; those rows drive Settings → Billing, the per-key spent this month figure and GET /v1/usage.
#Per-key controls
Set on creation or later in Settings → Security → API keys (or PATCH /api/keys/{id} with your web session):
| Control | Behaviour |
|---|---|
Monthly spend cap (USD, whole cents — 0.05, not 0.001; a finer value is refused with 422) | Once month-to-date billed spend on the key reaches the cap, further calls return 402 spend_cap_reached. The request that crosses the line still completes (its cost is unknown until it ends), so the overshoot is at most one request. Resets on the 1st of each month (UTC). |
| Allowed models | A list of catalog keys. Any other model returns 403 model_not_allowed. Empty = every API model. Legacy aliases are resolved to their current key. |
Use both on any key that leaves your own infrastructure.
#Embeddings
POST /v1/embeddings — OpenAI-compatible. Model text-embedding-3-small (1536 dimensions). input is a string or a list of up to 256 strings (12,000 characters each). Billed on input tokens; embeddings are not part of any plan allowance, so they are balance-funded at 2×. encoding_format is float (default) or base64 (little-endian float32 — what the official SDKs request by default, so client.embeddings.create(...) works unmodified).
emb = client.embeddings.create(model="text-embedding-3-small", input=["Milly Lab", "public API"])
print(len(emb.data[0].embedding), emb.usage.prompt_tokens)
#Images
POST /v1/images/generations — OpenAI-compatible request and response. model is an image key from /v1/models (gpt-image-2, gemini-nano-banana, fal-ai/flux-pro, fal-ai/stable-diffusion-xl), n 1–4, size 1024x1024 · 1536x1024 · 1024x1536 (mapped to the model's nearest aspect ratio), quality low · standard · high (medium/hd accepted). The call is synchronous (up to 180 s) and returns hosted urls; usage.billed_usd is what was charged.
img = client.images.generate(model="gpt-image-2", prompt="A watercolor map of the Pamirs", size="1024x1024")
print(img.data[0].url)
Images produced through the API are not added to the Light Studio history.
#Usage endpoint
GET /v1/usage?from=YYYY-MM-DD&to=YYYY-MM-DD[&key=<id>] returns requests, tokens and billed dollars (plan + balance, fee included) for every key on the account — so one reporting key can watch a fleet. Default range is month-to-date; the maximum is 92 days. Web-app chat is never included.
{
"object": "usage",
"from": "2026-09-01T00:00:00", "to": "2026-09-10T12:00:00",
"totals": { "requests": 412, "tokens_in": 1830000, "tokens_out": 210000, "usd": 12.41 },
"by_key": [{ "key_id": "…", "name": "prod", "hint": "a1b2", "requests": 400, "usd": 12.10 }],
"by_model": [{ "model": "claude-sonnet-5", "requests": 300, "usd": 10.20 }]
}
#Rate limits
| Limit | Value |
|---|---|
| Requests per key | 120 per minute → 429 rate_limit_exceeded with Retry-After (seconds) |
| Active keys per account | 5 |
| Messages per request | 200 |
| Characters per request | 400,000 |
| Embedding inputs per call | 256 × 12,000 characters |
| Images per call | 4 |
Every /v1 response carries x-request-id (quote it when you write to support) and, once a key is authenticated, x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-window (seconds). A 429 adds Retry-After — wait that many seconds, then retry. Do not retry a 402 without changing something (top up, raise the cap, pick a cheaper model).
#Error codes
Every error is OpenAI-shaped: {"error": {"message": "…", "type": "…", "code": "…"}} — branch on code.
| HTTP | code | Meaning |
|---|---|---|
| 400 | invalid_body · invalid_messages · invalid_input | Malformed request, no user text, bad embedding input |
| 401 | missing_api_key · invalid_api_key | No bearer key, or a revoked/unknown one |
| 402 | insufficient_for_request · insufficient_balance | Allowance + balance cannot cover the request |
| 402 | spend_cap_reached | This key's monthly cap is reached |
| 403 | api_access_required | The account's plan has no API access |
| 403 | model_not_allowed | Model not on the key's allow-list |
| 403 | account_disabled | Account suspended |
| 404 | model_not_found | Unknown or retired model — only /v1/models rows are callable |
| 429 | rate_limit_exceeded | Per-key limit (120/min); Retry-After says how long to wait |
| 502 | upstream_error | The provider failed before producing output — nothing billed |
| 503 | provider_unavailable | The model's provider (chat, embeddings or image) is not configured on this deployment — nothing billed |
#Calling from a browser
CORS is open on /v1 for any origin (without credentials), so browser apps can call the API directly with Authorization: Bearer mk_…. A key shipped in browser code is visible to every visitor. For such keys: set a tight spend cap and an allowed-model list, rotate them regularly, or proxy through your own backend where the key stays private.
#Best practices
- Keep keys in environment variables or a secrets manager, one key per application or environment, and revoke keys you no longer use.
- Put a spend cap on every key; alert on
402 spend_cap_reachedin your logs. - Read
usagefrom every response (or the last streamed chunk) and reconcile withGET /v1/usage. - Choose the model per task: the price list is live, and a smaller model is often enough.
- Send
max_tokenswhen you know the answer size — it bounds the reserve and the bill. - Handle 429 by waiting
Retry-Afterseconds (exponential backoff on top is fine); treat 402 as a configuration signal, not a transient error.
#Coming next
- Tool-calling passthrough (
tools/tool_choiceare accepted today but not executed). - Video and 3D generation through
/v1. - Per-key IP allow-lists.