Inference pricing that
stops at the token

One API for 500+ models across 70+ providers. Pay provider list price, add one flat platform fee, and bring your own keys or GPUs whenever it works out cheaper.

Free

Kick the tires on open models — no card required.

$0platform fee

100 requests / day

  • 20+ open models, 3 shared providers
  • Playground and OpenAI-compatible API
  • Latency and throughput metrics
  • 7-day activity log
Most Popular

Pay-as-you-go

Every model, every provider, billed per token.

4%platform fee on inference spend

No minimum, no commitment

  • 500+ models across 70+ providers
  • $50K/mo of BYOK traffic with no fees
  • Budgets, prompt caching and spend alerts
  • Email + shared Slack support

Enterprise

Committed capacity, private routing, contractual SLAs.

Customvolume pricing

Invoicing and AWS Marketplace

  • Dedicated throughput and committed-use pricing
  • Policy-based routing and residency controls
  • SSO/SAML, SCIM and managed policy enforcement
  • Negotiated uptime SLA and a named account team
FeatureFreePay-as-you-goEnterprise
Access
Models20+ open models500+ models500+ models
Providers3 shared providers70+ providers70+ providers
Playground & OpenAI-compatible API
Bring your own provider keys (BYOK)$50K/mo free, 4% afterCustom limits
Bring your own GPUsSelf-serve connectManaged fleets
Rate limits100 req/dayHigh shared limitsDedicated capacity
Cost control
Platform feeNone4% on inference spendNegotiable / volume tiers
Token pricingFree models onlyAt-provider list priceCommitted-use discounts
Prompt & prefix caching
Budgets and spend alerts
Per-key and per-team cost attribution
Invoicing & committed spend
Routing & reliability
Automatic failover between providers
Price and latency-aware routing
Preferred-provider ordering
Region and data-residency pinning
Policy-based routing
Uptime SLABy contract
Observability
Activity logs7 days90 daysCustom retention
Log export & webhooks
Latency, TTFT and throughput metrics
Evaluation and A/B traffic splits
Security & governance
Zero data retention on request
Management API keys
SSO / SAML & SCIM
Managed policy enforcement
VPC / self-hosted control plane
Support
Support levelCommunityEmail + Shared SlackShared Slack + SLA
Onboarding & migration help
Dedicated account manager

Token prices are set by the underlying providers and passed through at list price — the platform fee is the only markup. Ask about volume discounts.

Pricing FAQ

Billing and pricing

Usage and rate limits

Routing and latency

Privacy and security

Models and reliability

Ready to get started?

Start on free models in minutes, then scale onto committed capacity when the traffic shows up.