Performance, Latency & Economic Benchmarks
Comprehensive comparison of System One decision models against traditional autoregressive frontier LLMs across real-world workloads, automated evaluations, and cost metrics.
| Model / Pipeline | Input / MTok | Output / MTok | P50 Latency | Sampling Method | Type Error Rate | Cost / 100k Decisions |
|---|---|---|---|---|---|---|
| TypeSafe Jev 1.1 | $0.042 ($42/BT) | $0.00 (FREE) | 114ms | Parallel decision | 0.0% (Guaranteed) | $0.008 |
| OpenAI GPT-4o | $2.50 | $10.00 | 3,450ms | Autoregressive | 1.2% | $1.250 |
| Claude 3.5 Sonnet | $3.00 | $15.00 | 3,200ms | Autoregressive | 0.9% | $2.400 |
| DeepSeek V3 | $0.14 | $0.28 | 1,950ms | Autoregressive | 1.5% | $0.080 |
| Google Gemini 1.5 Flash | $0.075 | $0.30 | 1,400ms | Autoregressive | 2.1% | $0.045 |
System One Evaluation Pipeline
Jev evaluates structured decisions in parallel without autoregressive token generation, enabling dramatic cost reductions for high-volume pass/fail assertions compared to invoking frontier LLMs.
High-Throughput Content Classification
The dual-system architecture processes bulk content categorization workloads by delegating filtering and classification to Jev, reserving frontier LLMs for high-value synthesis on qualified candidates only.
Dual-System Agent Architecture
Autonomous agents using a dual-system approach delegate high-frequency state evaluation decisions to Jev while retaining frontier LLMs for strategic planning and complex reasoning tasks.
Verified Benchmark Sources
"Jev is the first frontier model built for automation rather than chat. Input tokens are priced at $0.042 per million tokens ($42 per billion). Output tokens are free because decisions are sampled in parallel."
Context: Official launch paper and system announcement for the first System One Model.