LAT   37.7749° N
LON   122.4194° W
ROUTE M-217
P95 Latency
612 ms
Tokens routed
12.4M
Success rate
99.97%
Cost savings
$1.2M+

Cut your AI spend in half.
Keep the same quality.

Route every request through one gateway — metered by team, redacted before egress, and sent to the cheapest model that clears your quality bar.

OpenAI-compatible · Self-hosted or VPC · No prompt retention
48%
lower cost
34 ms
overhead
100%
policy coverage
Live request Routed 34 ms
Incoming prompt
Refund the order for A. Schmidt██████████ on card ···7753███████. Item arrived damaged.
PII detected & redacted
Meter 412 tokens
Redact 2 spans
Route Small tier
Model evaluation Bar 92+
Model Score Cost
Cache below quality bar 88 $0.0000
Small cheapest above bar 94 $0.0074
Mid-tier 31% of traffic 96 $0.0180
Frontier 8% of traffic 97 $0.0412
Frontier default
$0.0412
Routed cost
$0.0074
Same answer
−82%
The problem

Enterprise AI budgets tripled last year. Most finance teams still can't say what they bought.

Inference arrives as one invoice line per provider. Nobody can attribute it to a team, an application or a customer — so nobody can govern it. Underneath that opacity sit three compounding costs: spend you can't explain, spend you didn't need, and prompts you can't prove were safe.

No visibility 01
73%
of enterprises cannot attribute AI spend to a team or application.

Provider invoices arrive as a single monthly line. Without per-request attribution there is no budget owner, no forecast and no chargeback — only a number that grows.

Cost: unforecastable budget
Structural waste 02
1 in 3
requests go to a frontier model that a small model would have answered.

Teams hardcode the strongest model during prototyping and never revisit it. Identical prompts are re-billed instead of cached. The waste is invisible, so it compounds quietly.

Cost: 40–60% of the invoice
Unmanaged exposure 03
1.24M
personal-data spans reached prompts in a single month, pre-gateway.

Customer names, account numbers and national IDs are pasted into prompts and sent to third-party providers. Prompts are data egress, and they are the least governed channel in the enterprise.

Cost: unprovable compliance
Source: Meridian internal benchmark.
The solution

One endpoint in front of every provider. Three controls behind it.

Change one base URL. Your applications keep the OpenAI-compatible calls they already make.

01 Meter

Every request is attributed to a team, application and key before it is forwarded. Finance gets chargeback; engineering gets a budget they can see.

Per-team and per-app cost attribution
Budget thresholds with alerts at 80% and 90%
Full request and response audit log
Attribution coverage
100% of routed traffic
02 Route

Each prompt is scored for complexity and sent to the cheapest model that clears the quality bar. Repeat prompts never reach a provider at all.

Complexity-scored tier selection
Semantic cache ahead of every provider
Automatic escalation when confidence is low
Cache hit rate
61% was 5%
03 Redact

Personal data is detected and replaced on-device, before the prompt leaves your network. Policy is set per team, and every redaction is logged.

On-device detection, no third-party call
Locale packs: US, EU, India
Per-team strict or standard policy
Policy violations
0 in 12 months
Outcomes · first 90 days

Usage up 3.1×. Spend down 48%.

Saved per month
$86,400
48% below pre-gateway baseline
Monthly spend
$94,200
down from $181,000
Tokens processed
84.2B
up 212% year over year
Cache hit rate
61%
from 5% before deployment
PII entities redacted
1.24M
zero policy violations
Added latency
34 ms
p50, including redaction
Based on estimated values.
Deployment & assurance

Runs inside your perimeter. Nothing about it is a black box.

Redaction happens before egress, on your infrastructure. Prompts are never retained by us, because they never reach us.

Self-hosted or single-tenant VPC Available
SOC 2 Type II · ISO 27001 In progress
Prompt retention by Meridian None
Source available for audit Open
OpenAI-compatible

Swap one base URL. No SDK change, no rewrite.

Bring your own keys

Your provider contracts and rates stay yours.

Provider failover

Automatic retry on a second provider when one degrades.

Local models

Route sensitive workloads to on-premise weights.

Policy as code

Redaction and routing rules reviewed in your repo.

Cost simulation

Replay last month against a new routing policy.

SSO and SCIM

Okta, Entra ID, and automated deprovisioning.

Exportable audit log

Streams to your SIEM in real time.

Bring us your last invoice. We'll show you the half you didn't need to pay.

A two-week shadow deployment reads your live traffic and reports what routing and caching would have saved — with no change to your applications.

Contact