Owned demonstration · September 30, 2026

Compare the outputs. Reproduce the failures.

Two runnable examples show what the integration pilot delivers. The application classifies 20 fictional support messages. These are our demonstrations, not customer results.

1. A measured two-model comparison

All 40 live requests completed. Both models ran through OpenRouter; this measures a model change within one gateway.

ModelExpected labels matchedMean latencyEstimated token cost
GPT-4.1 mini20 / 20847 msUS$0.0006756
Gemini 2.5 Flash Lite18 / 20252 msUS$0.0001472

The two disagreements are retained in the raw report. These simple examples do not establish general accuracy or customer savings. Costs are usage-based estimates for the successful run, excluding setup, failed qualification attempts, hosting and subscriptions.

Earlier local-Mistral, hosted-Mistral and NIM qualification attempts encountered timeouts, rate limits or an unavailable route. Their reports are included so the selection is inspectable.

2. A provider failure you can reproduce

The offline failure lab injects a rate limit and verifies exactly one approved fallback. Sixteen tests cover provider errors, stalled bodies, output validation and budget checks. Authentication/configuration errors stop; they do not start a retry loop.

npm test
node failure-lab.ts

Download and inspect

The MIT source archive includes the original application, exact patch, configuration example, 20 cases, recorded reports, tests and rollback instructions. Node.js 22.18 or newer; no dependencies. Live comparisons require your own approved provider key and may incur usage charges.

Download the free examples

Verify the archive checksum

See the US$495 implementation pilot