DeepInfra

DeepInfra

US
69 models available

Models served by DeepInfra

ModelContextInput /MOutput /MCache read /MCache hitLatencyThroughput
128k$0.53/M$2.26/M$0.37/M59.6%0.9s22 tps
128k$0.25/M$0.94/M$0.14/M20.8%1.1s19 tps
128k$0.26/M$1.00/M$0.14/M34%0.7s15 tps
163k$0.27/M$0.40/M$0.14/MN/A1.7s13 tps
163k$0.27/M$0.40/M$0.14/MN/A1.7s13 tps
1.0M$0.09/M$0.19/M$0.02/M0%0.8s24 tps
1.0M$0.09/M$0.19/M$0.02/M1%0.8s24 tps
1.0M$0.08/M$0.19/M$0.02/M19.3%0.8s29 tps
1.0M$0.08/M$0.19/M$0.02/M4%0.8s29 tps
1.0M$0.23/M$0.68/M$0.07/M75%2.7s35 tps
1.0M$1.36/M$2.73/M$0.10/M0%1.1s35 tps
1.0M$1.36/M$2.73/M$0.10/M80.9%1.1s35 tps
1.0M$1.36/M$2.73/M$0.10/MN/A0.9s69 tps
1.0M$1.36/M$2.73/M$0.10/M32.8%0.9s69 tps
262k$0.07/M$0.36/M$0.01/M0%0.6s24 tps
262k$0.14/M$0.40/M$0.02/M3.7%2.7s61 tps
200k$0.53/M$2.10/M$0.10/M84.2%0.6s47 tps
200k$0.53/M$2.10/M$0.10/M75%0.6s47 tps
200k$0.42/M$1.84/M$0.08/M36.1%0.9s21 tps
200k$0.42/M$1.84/M$0.08/M28.5%0.9s21 tps
200k$0.06/M$0.42/M$0.01/MN/A0.4s20 tps
200k$0.06/M$0.42/M$0.01/MN/A0.4s20 tps
200k$0.63/M$2.18/M$0.13/MN/A0.8s30 tps
200k$0.63/M$2.18/M$0.13/MN/A0.8s30 tps
200k$1.10/M$3.68/M$0.22/MN/A5.9s8 tps
200k$1.10/M$3.68/M$0.22/MN/A5.9s8 tps
1.0M$0.51/M$1.64/M$0.10/M66.7%1.2s59 tps
1.0M$0.51/M$1.64/M$0.10/M66.7%1.2s59 tps
1.0M$1.26/M$4.20/M$0.13/MN/A2s45 tps
1.0M$1.26/M$4.20/M$0.13/MN/A2s45 tps
1.0M$0.08/M$0.26/M$0.02/M7%1.2s34 tps
128k$0.21/M$1.00/MN/A0%1.9s87 tps
128k$0.03/M$0.15/MN/AN/A0.4s81 tps
131k$0.06/M$0.26/M$0.02/M73.4%0.4s110 tps
66k$0.73/M$0.73/MN/A0%0.8s32 tps
524k$0.47/M$1.26/M$0.10/MN/A0.9s75 tps
524k$0.47/M$1.26/M$0.10/MN/A0.9s75 tps
256k$0.47/M$2.36/M$0.07/M10.1%0.8s26 tps
256k$0.47/M$2.36/M$0.07/M10.1%0.8s26 tps
256k$0.79/M$3.68/M$0.16/MN/A1.4s28 tps
262k$0.71/M$3.57/M$0.14/M0%6.4s42.5 tps
1.0M$2.99/M$14.96/M$0.30/M17.2%1.4s22 tps
131k$0.10/M$0.34/MN/A0%0.5s16 tps
1.0M$0.21/M$0.84/MN/AN/A0.4s40 tps
1.0M$0.14/M$0.68/M$0.03/MN/A2.1s25 tps
1.0M$1.05/M$3.15/M$0.21/M24.3%0.7s41 tps
1.0M$1.05/M$3.15/M$0.21/M0%0.7s41 tps
1.0M$0.14/M$0.68/M$0.03/MN/A2.1s25 tps
205k$0.26/M$1.05/M$0.05/M1.4%0.6s27 tps
512k$0.29/M$1.16/M$0.06/M95.5%0.9s48 tps
512k$0.29/M$1.16/M$0.06/M0.4%0.9s48 tps
262k$0.09/M$0.42/MN/A0%1.4s54 tps
262k$0.09/M$0.42/MN/A0%1.4s54 tps
256k$0.53/M$2.31/M$0.11/M38.3%4.2s3 tps
256k$0.53/M$2.31/M$0.11/M0%10.1s307.4 tps
29k$0.08/M$0.21/M$0.04/M19.6%0.9s142.7 tps
29k$0.08/M$0.21/M$0.04/MN/AN/AN/A
131k$0.38/M$0.42/MN/A0%0.6s15 tps
260k$0.34/M$3.36/MN/AN/A1.7s23 tps
260k$0.34/M$3.36/MN/AN/A1.7s23 tps
262k$0.10/M$1.00/MN/AN/A1s105 tps
262k$0.10/M$1.00/MN/AN/A1s105 tps
262k$0.09/M$0.58/MN/AN/A0.5s13 tps
41k$0.08/M$0.29/MN/A0%0.5s25 tps
256k$0.09/M$1.16/MN/A0%0.6s71 tps
258k$0.47/M$3.15/M$0.23/MN/A0.7s35 tps
258k$0.47/M$3.15/M$0.23/MN/A0.7s35 tps
262k$2.10/M$6.30/M$0.21/M82.4%0.6s75.5 tps
262k$2.10/M$6.30/M$0.21/M82.4%0.6s75.5 tps

Prices are per million tokens and include the 5% provider-selection markup. Availability and pricing refresh continuously; each model page shows the live provider comparison.

Using DeepInfra via the API

Append :deepinfra to the model ID, set "provider": "deepinfra" in the request body, or send an X-Provider: deepinfra header. Explicit provider selection adds a 5% markup over that provider's base price; the prices in the table above already include it.

curl https://nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-R1-0528:deepinfra",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the API documentation for provider preferences, price-aware routing, and error behavior.