Owned demonstration · September 30, 2026
Compare the outputs. Reproduce the failures.
Two runnable examples show what the integration pilot delivers. The application classifies 20 fictional support messages. These are our demonstrations, not customer results.
1. A measured two-model comparison
All 40 live requests completed. Both models ran through OpenRouter; this measures a model change within one gateway.
| Model | Expected labels matched | Mean latency | Estimated token cost |
|---|---|---|---|
| GPT-4.1 mini | 20 / 20 | 847 ms | US$0.0006756 |
| Gemini 2.5 Flash Lite | 18 / 20 | 252 ms | US$0.0001472 |
The two disagreements are retained in the raw report. These simple examples do not establish general accuracy or customer savings. Costs are usage-based estimates for the successful run, excluding setup, failed qualification attempts, hosting and subscriptions.
Earlier local-Mistral, hosted-Mistral and NIM qualification attempts encountered timeouts, rate limits or an unavailable route. Their reports are included so the selection is inspectable.
2. A provider failure you can reproduce
The offline failure lab injects a rate limit and verifies exactly one approved fallback. Sixteen tests cover provider errors, stalled bodies, output validation and budget checks. Authentication/configuration errors stop; they do not start a retry loop.
npm test
node failure-lab.tsDownload and inspect
The MIT source archive includes the original application, exact patch, configuration example, 20 cases, recorded reports, tests and rollback instructions. Node.js 22.18 or newer; no dependencies. Live comparisons require your own approved provider key and may incur usage charges.
Download the free examples