Начало работы
Plans, Billing & Usage
This guide explains what your plan includes, how the monthly allowance is spent, what happens when it runs out, and how to read the usage view. It is for anyone on a paid or free Milly Lab plan.
#How Milly Lab charges for model use
Every model costs a different amount to run. Rather than selling you a message count, Milly Lab gives each plan a monthly allowance in US dollars, divided across a specific set of models. Each model has its own portion of that allowance.
Two consequences:
- You can exhaust one model while others still have budget. Running out of your Claude Sonnet allowance does not stop you using GPT.
- Plans are not cumulative. A plan includes exactly the models listed for it — not the models of the plans below it. Moving up a plan can mean losing access to a model, so check the model list, not just the price.
#What each plan includes
Free
No monetary allowance. A registered user without a subscription gets chat only, on two models:
| Models | gpt-5.4-nano, deepseek-v4-flash |
| Limit | 10 messages per model per 4-hour window |
| Everything else | Locked — image, video, 3D, Swarm Mind, file generation |
The window is rolling. When you hit the limit, the error tells you when it resets.
Starter — $7.00 per month
| Model | Allowance | Approx. messages |
|---|---|---|
gpt-5.4-nano | $1.25 | 1,445 |
claude-haiku-4-5 | $3.60 | 973 |
deepseek-v4-flash | $1.40 | 4,545 |
gemini-2.5-flash | $0.75 | 466 |
Chat only. No file generation, no Swarm Mind.
Pro — $15.00 per month
| Model | Allowance | Approx. output |
|---|---|---|
gpt-5.4 | $4.00 | 281 messages |
claude-sonnet-5 | $4.50 | 300 messages |
deepseek-v4 | $2.00 | 1,576 messages |
gemini-3.6-flash-200k | $2.25 | 273 messages |
gemini-nano-banana | $2.25 | 57 images |
Adds file generation. No Swarm Mind.
Max — $27.00 per month
| Model | Allowance | Approx. output |
|---|---|---|
gpt-5.5 | $6.25 | 156 messages |
claude-sonnet-5 | $4.65 | 310 messages |
deepseek-v4-pro | $2.50 | 359 messages |
gemini-3.6-flash | $2.10 | 180 messages |
gemini-nano-banana | $3.25 | 83 images |
gpt-image-2 | $2.00 | 37 images |
fal-ai/flux-pro | $1.20 | 21 images |
veo-3.1 | $5.05 | 5 clips |
Adds Swarm Mind and video generation.
Advanced — $52.00 per month
| Model | Allowance | Approx. output |
|---|---|---|
gpt-5.5 | $8.75 | 219 messages |
claude-sonnet-5 | $5.40 | 360 messages |
claude-opus-4-8 | $6.75 | 193 messages |
deepseek-v4-pro | $3.75 | 539 messages |
gemini-3.6-flash | $3.50 | 300 messages |
gpt-5.3-codex | $7.50 | 226 sessions |
gemini-nano-banana | $3.05 | 78 images |
gpt-image-2 | $1.60 | 30 images |
fal-ai/stable-diffusion-xl | $0.60 | 100 images |
veo-3.1 | $5.00 | 5 clips |
gen4_turbo | $3.50 | 4 clips |
kling-3.0 | $2.00 | 4 clips |
meshy-api | $0.60 | 3 models |
Everything, including 3D generation.
Enterprise
Enterprise plans are arranged directly — get in touch through the Contacts page.
The message counts are estimates, calculated from a reference message of roughly 4,000 input tokens and 750 output tokens. A long conversation, a large attachment, or a run that reads files from a connected repository consumes more per message. Treat them as a guide, not a quota.
#What uses your allowance
Anything that calls a model:
- Chat messages, including the conversation history sent as context each turn.
- Image, video and 3D generation.
- Every model call inside a Swarm Mind run — an orchestrator plus several workers means several charges from one brief.
- File contents read from a connected GitHub repository or Google Drive, which are sent as input tokens.
- Web search, when enabled for a message.
Swarm Mind is the expensive one. A Project Mode run may call models many times. A run over a connected repository adds the file contents on top.
Web search adds cost and latency. It is a toggle in the composer — leave it off unless the answer genuinely depends on current information.
#When your allowance runs out
Your prepaid balance takes over, at twice the underlying provider cost. The markup is why the balance is a fallback rather than the main way to pay — a plan is the cheaper route for regular use.
If both the model's allowance and your balance are insufficient, the request is refused before it runs. You are not charged for a partial answer, and nothing is billed silently.
Milly Lab also checks in advance that a request can pay for its full input and a usable answer. This is why you may be blocked from a request that looks affordable — a long conversation with a large attachment can cost more than the balance remaining, and Milly Lab refuses rather than cutting the answer off mid-sentence.
#Reading the errors
All of these arrive as HTTP 402 Payment Required with a machine-readable code.
| Code | What it means | What to do |
|---|---|---|
free_limit | Free tier: 10 messages used for this model in the current 4-hour window | Wait for the reset time in the message, use the other free model, or subscribe |
free_model_locked | Free tier: this model needs a subscription | Use gpt-5.4-nano or deepseek-v4-flash, or subscribe |
plan_required | The feature needs any paid plan; the free tier cannot buy into it with balance | Subscribe |
feature_required | Your plan lacks this capability — file generation needs Pro or above | Move to a plan that includes it |
insufficient_balance | The model's allowance is spent and your balance cannot cover the next call | Top up, or use a model that still has allowance |
insufficient_for_request | This specific request is too large for what remains | Shorten the conversation, remove attachments, start a new conversation, or top up |
insufficient_for_request catches people out. It usually means the conversation has grown, not
that you are out of budget generally. Starting a fresh conversation often resolves it, because the
history is no longer being resent.
#Where to see your usage
Go to Settings → Billing (/app/settings/billing).
It shows your current plan, how much of each model's allowance is left, your balance, and your transaction history.
Every billable turn is recorded with the model, input tokens, output tokens, cost, and whether it came from your plan allowance or your balance. If a charge looks wrong, that record is what to check.
#Billing periods
Allowances are granted per subscription period and reset at the start of each one. Unused allowance does not carry over.
When a period ends without renewal, access falls back to the free tier. Your conversations, generated files and connected workspaces are not deleted — you keep everything, you just lose access to the paid models until you resubscribe.
#Paying
No payment gateway is connected yet. Milly Lab cannot currently take payment. A gateway with local banks is planned.
Until then the upgrade and top-up screens will not complete a purchase.
To arrange a plan in the meantime, get in touch through the Contacts page.
#Common questions
Can I move allowance between models? No. Each model's portion is fixed for the plan.
Why did a Swarm Mind run cost so much more than a chat message? It made many model calls — an orchestrator and several workers, each with its own input and output. Connected-workspace file contents add to that.
Why does the same prompt cost different amounts? The whole conversation is sent as context each turn, so cost grows as a conversation lengthens. Attachments and connected files add to it.
I upgraded and lost a model. Plans are not cumulative. Check the model list for the plan you moved to.
Does my team share my allowance? Every billable turn is recorded per user in the usage view. For how your plan meters team usage, ask through the Contacts page.