Qwerky 72B
Linear models offer a promising approach to significantly reduce computational costs at scale, particularly for large context lengths. Enabling a >1000x improvement in inference costs, enabling o1 inference time thinking and wider AI accessibility.
Pricing
Auto routing · per 1M tokens- Input
- $0.50
- Output
- $0.50
Specifications
- Context window
- 32K
- Max output
- 8.2K
- Parameters
- 72B
Benchmarks
No public benchmark scores for this model yet.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…