Models and Usage
Compare Bkper AI models, capabilities, usage rates, monthly allowances, and usage visibility.
Bkper AI includes access to selected AI models with eligible Bkper plans. Sign in with Bkper instead of setting up separate provider accounts or API keys. Use the models through the Bkper CLI Agent or another compatible client, with usage tracked against one monthly allowance.
To configure authentication, endpoints, and compatible clients, see Bkper AI Provider.
Cost control by design
Bkper AI combines ready-to-use access with a centralized allowance:
- Ready to use. The Bkper CLI Agent connects with your Bkper login. No separate provider setup is required.
- Included, not billed separately. AI usage draws from a monthly allowance included in eligible plans. There is no separate AI bill.
- Choose by task. Switch among selected models based on the work, capability, and usage rate you need.
- Centralized allowance enforcement. Bkper AI meters provider-reported usage and blocks new requests once the recorded monthly allowance is exhausted. No automatic paid overages are charged.
- Bring another provider when needed. External providers remain available where supported and do not consume the Bkper AI allowance.
Requests already in flight can settle after the allowance snapshot used for admission. This means recorded usage can exceed the allowance slightly under concurrency, but Bkper does not automatically bill that difference as a paid AI overage.
Choose a model
Choose a model based on its workload, cost, and capabilities. Clients can request lower output budgets and any reasoning effort listed for the selected model.
| Model | Best for | Capabilities |
|---|---|---|
Gemini 3.5 Flash google/gemini-3.5-flash | Fast Gemini model balancing multimodal reasoning, tool use, and cost. | Reasoning efforts: low, medium, highMaximum context: 1,048,576 tokens Maximum output: 65,536 tokens |
GPT-5.6 Luna openai/gpt-5.6-luna | Cost-efficient GPT-5.6 model for fast, high-volume workloads. | Reasoning efforts: none, low, medium, high, xhighMaximum context: 1,050k tokens Maximum output: 128k tokens |
GPT-5.6 Terra openai/gpt-5.6-terra | Balanced GPT-5.6 model for capable, cost-efficient everyday work. | Reasoning efforts: none, low, medium, high, xhighMaximum context: 1,050k tokens Maximum output: 128k tokens |
Grok 4.5 xai/grok-4.5 | xAI’s latest Grok for chat, coding, agentic tools, and lower hallucination risk. | Reasoning efforts: low, medium, highMaximum context: 500k tokens Maximum output: 500k tokens |
The table shows each model’s maximum supported capabilities. Bkper CLI uses a 200,000-token managed context window and a 32,000-token maximum output for included models. Other compatible clients can use the settings supported by each model.
Usage rates
Usage rates reduce the included monthly allowance. They are not billed separately by Bkper.
USD of included usage per one million tokens
| Model | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Gemini 3.5 Flash | $1.50 | $0.15 | $0.00 | $9.00 |
| GPT-5.6 Luna | $1.00 | $0.10 | $1.25 | $6.00 |
| GPT-5.6 Terra | $2.50 | $0.25 | $3.125 | $15.00 |
| Grok 4.5 | $2.00 | $0.50 | $0.00 | $6.00 |
Input means tokens sent without a cache match. Cache read means reused input already stored by the provider. Cache write means input added to a provider cache. Output includes generated response and reasoning tokens reported by the provider.
Monthly allowance
For paid plans:
Monthly AI usage allowance = 50% of normalized monthly software subscription value.
The calculation works as follows:
- Monthly plans use the active recurring software subscription value.
- Annual plans divide the annual recurring value by 12, then apply 50%.
- Professional and custom plans use only the recurring software subscription component.
- Professional services, implementation, consulting, taxes, unrelated one-time charges, credits, refunds, and prorations do not increase the allowance.
- Free users receive a separately configured trial allowance. The 50% formula does not apply to Free.
The allowance is an inference entitlement. It is not cash, refund value, a separately billed balance, or transferable account credit.
The allowance resets monthly. Unused allowance does not roll over. The authenticated Bkper AI usage dashboard is authoritative for your exact current allowance.
Individual and pooled usage
Allowance scope follows the subscription:
- Free and Standard usage is assigned to the individual user.
- Business and Professional usage can be pooled when the subscription has domain-wide scope.
- Everyone sharing a pooled allowance reduces the same monthly total.
A pooled allowance does not make every user’s request history visible to everyone. Visibility depends on the viewer’s billing role.
Usage visibility and privacy
The Bkper AI usage dashboard separates allowance visibility from request attribution:
- Regular users see the shared allowance remaining and their own requests and usage.
- The billing or subscription administrator sees domain-wide usage attributed by user, AI model, and app or source.
- The dashboard does not expose prompts or responses.
Usage value is an estimate based on the published rates above. It shows how much of the included allowance a request consumed; it is not a separate Bkper charge.
When the limit is reached
Bkper AI blocks new allowance-backed requests once the recorded monthly allowance is exhausted. There are no automatic paid Bkper AI overages at launch.
Bkper AI is not a lock-in: where supported, you can connect an external model provider at any time. External subscriptions, API keys, charges, privacy terms, and limits are governed by that provider and do not use the included Bkper AI allowance.
How Bkper selects models
We build Bkper with the Bkper CLI Agent and use it every day. We test many models through real work and include only those that consistently work well for us within our cost and control constraints. The catalog is a practical, opinionated shortlist—not a directory of every available model.
Bkper prioritizes strong results at controlled cost — the efficient frontier of capability per dollar — rather than pursuing the highest benchmark score at any price. Selection also considers:
- results and reliability in daily agent workflows;
- provider capabilities and tool use;
- observed usage cost;
- model capabilities and controls;
- public benchmarks.
Explore the live DeepSWE leaderboard.
DeepSWE measures long-horizon software-engineering work. It is one input into model selection, not a measure of accounting accuracy or a guarantee of performance in Bkper workflows.
Model references
| Model | Provider | Released |
|---|---|---|
| Gemini 3.5 Flash | 2026-05-19 | |
| GPT-5.6 Luna | OpenAI | 2026-07-09 |
| GPT-5.6 Terra | OpenAI | 2026-07-09 |
| Grok 4.5 | xAI | 2026-07-08 |
Sources
Last synchronized: 2026-07-24
Models.dev metadata is provided under the MIT License. Provider names and logos remain trademarks of their respective owners.