Qwerky 72B

Linear models offer a promising approach to significantly reduce computational costs at scale, particularly for large context lengths. Enabling a >1000x improvement in inference costs, enabling o1 inference time thinking and wider AI accessibility.

Pricing

Auto routing · per 1M tokens
Input
$0.50
Output
$0.50
Compare provider prices

Specifications

Context window
32K
Max output
8.2K
Parameters
72B

Benchmarks

No public benchmark scores for this model yet.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…