Pricing Terms
How Fronset charges for inference: no markup on model usage, a per-call fee on every call in both access modes, and a fee on credit purchases.
KAPUA Labs LLC — llm-bench inference and evaluation API
Version 2026-09-14 · Last updated 14 September 2026
These terms describe what you are charged and why. They are the Pricing Terms named in Terms of Service §14 and listed in §24’s entire-agreement clause, and they form part of that agreement. Where these terms and the Terms of Service differ on anything other than rates and charging mechanics, the Terms of Service govern.
They are written to be read, not skimmed.
1. No markup on inference
You pay the same rate for model usage as you would pay the provider directly. We do not add a percentage to token prices, and we do not adjust the token counts we bill you for.
When you make a request, we record the tokens the provider reports — input, output, cached input and cache creation — and charge them at that provider’s published rate for the model, context tier and service tier the request actually used. Your invoice shows those figures per model.
We make no money on the inference pass-through itself. Our revenue is the per-call fee in §3 — which applies to every call, on both access modes — and the credit-purchase fee in §2. Where a task has matured onto a published rate of ours instead of the pass-through, §1.3 says how that rate is set — from what the work costs us, aiming at no markup — and how it is corrected when the estimate behind it turns out to be wrong.
1.1 What “the same rate” means precisely
Provider prices vary by more than the model name, and we bill the rate that applied to your specific request. Input and output tokens are priced separately, at the provider’s own rates for each — as are cached input, cache-creation and reasoning tokens where the provider distinguishes them, and every other category the provider bills. Your invoice shows each separately rather than as one blended figure.
The other axes we bill on:
- Context tier. Several providers charge more above an input-size threshold. You are billed at the tier your request’s input token count fell into.
- Service tier. Batch requests are billed at the provider’s batch rate, which is lower than the synchronous rate.
- Time of day. One provider we route to prices some hours differently. You are billed at the multiplier in force when your request ran, not when it was invoiced.
- Cached, cache-creation and reasoning tokens are billed at their own rates where the provider distinguishes them.
Every one of these appears on your invoice as its own line, so you can see which rate applied and why.
1.2 What you are billed for when a task explores
Some tasks run more than one model while we establish which performs best for that task — this is how the benchmark is produced. While a task is in that state, every model call in the chain appears on your invoice, per model. That includes calls whose output was not returned to you: alternatives, and the evaluation calls that compare them. It also includes calls that failed, and our own retries after a provider error — the provider billed those to us, so at cost they reach you too.
We show this rather than averaging it away, so you can see what you paid for. Once a task reaches a stable published rate, you are billed that rate instead and the exploration is ours to fund.
1.3 How a published rate is set, and when it changes
For tasks where we publish our own rate rather than passing a provider’s through, that rate is set from measured cost: what the model calls behind that task have actually cost us, at the providers’ own prices, over the requests we have measured. It is set by a person, not by a job. An estimate is produced from that measurement, reviewed, and released deliberately; nothing reaches your bill because a background process recalculated something.
A published rate aims at cost, with no markup. We do not set one to earn money on the model usage behind it — as everywhere else in these terms, our revenue is the per-call fee (§3) and the credit-purchase fee (§2). But a published rate is a prediction: it is set from what the work has cost so far and charged on the work that follows, and a prediction can miss in either direction. Provider prices move, a different model may come to serve the task, and measurement lags what it measures. So we do not promise that a published rate equals cost at every instant, and we do not claim that it can never exceed what the same work would have cost at the provider’s own price.
What we promise is to keep measuring, and to correct. We compare the cost behind every published rate with what that rate is charging, on an ongoing basis. Where the prediction proves too low — the work costs more than the rate recovers — the rate goes up. Where it proves too high — a provider cut its prices, a cheaper model now serves the task as well, or the estimate overshot — the rate comes down. The correction runs in both directions, and by the same procedure as the original: a person reviews the measurement and releases a new rate.
A published rate is effective-dated and never edited in place. Changing one creates a new rate from a future date; usage already billed is never restated, in either direction. Which tasks are on a published rate, the current rate for each, and every change to one are published in the portal and through the API, each rate with the date from which it applies.
1.4 Failed requests and retries: the rule differs by mode
A request can fail, and a model call inside it can error and be retried. Whether either reaches your invoice depends on how that task is charged, so each of the three charging modes states its own rule rather than sharing one.
A task on a published rate (§1.3). A request that failed (Terms §1) is not billed — no token charge and no per-call fee. A retry after a provider error is ours, and so are the alternatives and evaluation calls. A published rate is one blended figure set from measured cost, so we carry that variance instead of itemising it to you.
A task that is still exploring (§1.2). This mode is billed at cost, and at cost means what the provider billed us — so it includes the model calls that did not work. Failed attempts, our retries after a provider error, the alternatives we try alongside the one we return and the evaluation calls that compare them all appear on your invoice. The per-call fee in §3 is charged on top, once, when the request completes.
A request of yours that failed (Terms §1) carries no per-call fee. Its pass-through still appears: the rule in this mode is what the provider billed us, and a provider bills for a model call whether or not we accepted its output. That includes a call whose output failed our format checks, and every call behind a request that never completed.
Your own keys (§3.2). We charge the per-call fee and nothing else, and only for a request that completed — a request of yours that failed carries no fee from us. What your provider charges you is a separate matter, and it is not limited in this way: see §3.2.
Our internal traffic is never on your invoice, in any mode. We are a customer of our own API; that usage is billed to us, not distributed across yours.
2. Credit purchases
llm-bench is prepaid. You buy credit, and usage draws it down.
A fee of 5.5% applies to each credit purchase, with a minimum of $0.80.
The fee covers payment processing and account servicing. It is charged on the purchase, never on your usage — which is what allows §1 to be true.
| you pay | credit you receive | fee |
|---|---|---|
| $5.80 | $5.00 | $0.80 (the minimum applies) |
| $10.80 | $10.00 | $0.80 (the minimum applies) |
| $21.10 | $20.00 | $1.10 |
| $105.50 | $100.00 | $5.50 |
| $527.50 | $500.00 | $27.50 |
The minimum binds on purchases below $14.55; above that the percentage governs. The two are never added together.
2.1 Credit is valid for 12 months
Credit expires 12 months after the date you buy it. Where you hold credit from more than one purchase, the oldest is always spent first, so a balance you keep using does not expire underneath you.
2.2 Automatic recharge is optional and off by default
You may set a balance threshold and a recharge amount, in which case we purchase credit on your stored payment method when your balance falls to that threshold. Doing so requires your explicit authorisation for off-session charges, set out separately from these terms in the Off-session payment mandate — turning automatic recharge on is how you give that authorisation, and turning it off withdraws it. It is off unless you turn it on, and you can turn it off at any time.
Each automatic recharge carries the same purchase fee as any other credit purchase, so the sum charged to your card is the recharge amount plus that fee. Three limits bind it: the per-purchase range in §2.3 ($500 is the most any single charge may be), your recharge amount must be at least twice your threshold, and your Account’s monthly ceiling on credit purchases (§2.3), which counts automatic and manual purchases together. You may set a lower daily limit of your own. Reaching a limit stops the charge; it never produces a larger or a deferred one.
If your card is declined, we do not present it again for that recharge. No credit is added and nothing further is charged. A declined card is not re-presented repeatedly — that is how a card issuer locks a card out. Where a payment fails for a reason that is not a decline — the processor was unavailable, or returned a temporary error — it may be attempted again a small number of times over the following days.
We do not send you a message when this happens, and the reason for a particular failure is not always recorded on your account. Where it is, it is readable in your billing settings. If you rely on automatic recharge, check your balance rather than waiting to be told.
2.3 Purchase limits, and refunds of unused credit
Each credit purchase is between $5 and $500, whether you make it yourself or automatic recharge makes it for you.
Your Account carries a monthly ceiling on credit purchases, counting automatic and manual purchases together. The ceiling is set per Account: it bounds payment exposure while an Account is new, and it is not a limit on how much of the Service you may use — Terms of Service §14 says no tier carries one. Your Account’s ceiling is shown on the Billing page of the portal, which also says how to have it raised. Reaching it stops the purchase; it never produces a larger or a deferred one. A change to your Account’s ceiling takes effect at once and is not an amendment of these terms.
Credit you have bought but not used is refundable. Ask us at the address in Terms of Service §25. We refund to the payment method that bought the credit, less the purchase fee, once any work in flight on your Account has settled — and with automatic recharge turned off first, so the refund is not immediately bought back. We process a refund within 14 days of the request. Credit that has been used is not refundable: the usage records on your Account are the evidence of what was delivered. Expired credit (§2.1) is not restored.
3. The per-call fee
Every call carries a per-call fee, whichever access mode served it. It does not matter whether we supplied the model access (§1) or you supplied the provider credential (§3.2) — the fee is the same, and it is charged once for the request however many model calls we make behind it.
What differs between the modes is only who pays the provider for the tokens: we do in the operator-supplied mode — the mode the service, the portal and the documentation call managed (Terms of Service §1 and §5.2) — and pass that through at cost under §1; you do in bring-your-own-key mode, directly. If you bring your own keys, read §3.2 — your provider bills you for every model call we make on your behalf, not only the one that produced your answer.
The per-call fee is graduated, and identical on every plan:
| calls per month | fee |
|---|---|
| first 10,000 | included |
| 10,001 – 50,000 | $2.00 per 1,000 |
| 50,001 – 250,000 | $1.50 per 1,000 |
| 250,001+ | $1.00 per 1,000 |
The bands are marginal: calls in each band are charged at that band’s rate, so crossing a threshold never increases the cost of the calls below it.
3.1 What counts as a call
One request you make is one call, regardless of how many model calls we make to serve it. Exploration alternatives, evaluation calls and our own retries never add a call — they are further model calls inside the one request, not further requests. A batch of 10,000 items is 10,000 calls; submitting it is not.
A request that failed carries no fee, in any mode. We charge for a call once, when it completes — a request completes when a model returned an output and it passed our format checks (Terms §1); a request that failed is not a completion. §1.4 states what the pass-through does in each mode.
That is how our fee is counted. It is not what your provider counts — see §3.2.
3.2 Bring your own keys: your provider bills you for the whole chain, not only the answer
Every model call we make to serve your request runs on your keys, and your provider bills you for all of them. That includes calls whose output you never see: the alternatives we try alongside the one we return, and the evaluation calls that compare them. It includes repairs when a model returns malformed output, and our own retries after a provider error.
We charge one per-call fee for the request. Your provider charges for the work.
A single request against a task that is still exploring can therefore look like this on the two bills:
| your provider’s bill | our bill | |
|---|---|---|
| the model call we returned to you | charged | — |
| 4 alternative model calls | charged | — |
| 1 evaluation call comparing them | charged | — |
| the request itself | — | 1 call |
§1.4’s rule for this mode is about our invoice only. A failed request and a retry cost you nothing from us; if your keys served them, your provider still charged for them.
You control whether this happens. Exploration is a setting on the task, not something we impose: you can turn it off for a single request or for the task as a whole. With it off there are no alternatives and no evaluation calls, so your request runs one model — though a retry after a provider error, or a repair of malformed output, can still add a call.
Work we schedule for ourselves is never on your keys. Separately from your requests, we run comparisons of our own to produce the public benchmark. Those are paid for by us, even when they sample a task of yours, because their output is discarded and returns you nothing.
We hold the chain to what the measurement needs. When alternatives do run, we do not run every model eligible for the task. We choose them by what would actually improve the comparison — favouring the models we have least evidence about, and weighing what each one costs — so the chain is a fraction of the eligible set rather than all of it. This is the same rule we apply when we are paying, and we do not widen it because you are.
And a failure never widens the chain. If the machinery that chooses the alternatives cannot run, we serve your request with one model rather than falling back to running all of them.
4. Plans
Your plan determines privacy and support, not your usage rates. Every plan pays the same rates in §1 and the same fees in §2 and §3.
5. Changes
We will publish a change to these terms before it takes effect. A change to a provider’s own rates reaches you as it reaches us, because §1 passes those rates through — that is not a change to these terms. Nor is a change to a rate we publish ourselves under §1.3: those are set, corrected and announced as §1.3 describes, in the portal and through the API.
6. Questions
Billing questions: llmbench@kapualabs.com
Version 2026-09-14. Published by KAPUA Labs LLC as part of the Terms of Service (§24).