DeepInfra

DeepInfra

US
65 models available

Models served by DeepInfra

ModelContextInput /MOutput /MCache read /MCache hitLatencyThroughput
128k$0.53/M$2.26/M$0.37/M59.9%0.9s26 tps
128k$0.25/M$0.94/M$0.14/M38.1%0.5s26 tps
128k$0.26/M$1.00/M$0.14/M42.1%0.9s10 tps
163k$0.27/M$0.40/M$0.14/M0%0.8s28 tps
163k$0.27/M$0.40/M$0.14/MN/A0.8s28 tps
1.0M$0.09/M$0.19/M$0.02/MN/A1.4s19 tps
1.0M$0.09/M$0.19/M$0.02/M48.2%1.4s19 tps
1.0M$0.08/M$0.19/M$0.02/M49.6%0.7s34 tps
1.0M$0.08/M$0.19/M$0.02/M51.9%0.7s34 tps
1.0M$1.36/M$2.73/M$0.10/M43.1%1s18 tps
1.0M$1.36/M$2.73/M$0.10/M51.5%1s18 tps
1.0M$1.36/M$2.73/M$0.10/M50.8%2.3s22 tps
1.0M$1.36/M$2.73/M$0.10/M23%2.3s22 tps
262k$0.07/M$0.36/M$0.01/M0%0.4s26 tps
262k$0.14/M$0.40/M$0.02/M13.5%2.1s19 tps
200k$0.53/M$2.10/M$0.10/M30.8%0.6s45 tps
200k$0.53/M$2.10/M$0.10/M23.9%1.6s38.1 tps
200k$0.42/M$1.84/M$0.08/M23.1%1.1s22 tps
200k$0.42/M$1.84/M$0.08/M20.7%1.1s22 tps
200k$0.06/M$0.42/M$0.01/MN/A0.7s38 tps
200k$0.06/M$0.42/M$0.01/MN/A0.7s38 tps
200k$0.63/M$2.18/M$0.13/MN/A0.9s22 tps
200k$0.63/M$2.18/M$0.13/MN/A0.9s22 tps
200k$1.10/M$3.68/M$0.22/MN/A0.9s13 tps
200k$1.10/M$3.68/M$0.22/MN/A0.9s13 tps
1.0M$0.51/M$1.64/M$0.10/MN/A1.1s40 tps
1.0M$0.51/M$1.64/M$0.10/M91.9%1.1s40 tps
1.0M$1.26/M$4.20/M$0.25/M4%0.7s38 tps
1.0M$1.26/M$4.20/M$0.25/M5.1%0.7s38 tps
1.0M$0.08/M$0.26/M$0.02/M63.3%1.7s31 tps
128k$0.16/M$0.63/MN/A0%0.3s136 tps
128k$0.03/M$0.15/MN/AN/A0.3s102 tps
66k$0.73/M$0.73/MN/A0%0.5s29 tps
256k$0.47/M$2.36/M$0.07/M0%0.6s33 tps
256k$0.47/M$2.36/M$0.07/M80.7%0.6s33 tps
256k$0.79/M$3.68/M$0.16/MN/A1s28 tps
262k$0.71/M$3.57/M$0.14/M0%0.8s53.5 tps
1.0M$2.99/M$14.96/M$0.30/M5.3%1.4s41 tps
131k$0.10/M$0.34/MN/A0%0.5s9 tps
1.0M$0.21/M$0.84/MN/AN/A0.3s52 tps
1.0M$0.14/M$0.68/M$0.03/M29.2%2.6s20 tps
1.0M$1.05/M$3.15/M$0.21/M17.7%1.5s32 tps
1.0M$1.05/M$3.15/M$0.21/M0%1.5s32 tps
1.0M$0.14/M$0.68/M$0.03/M0%2.6s20 tps
205k$0.26/M$1.05/M$0.05/M12.8%0.5s32.5 tps
512k$0.29/M$1.16/M$0.06/MN/A1.1s20 tps
512k$0.29/M$1.16/M$0.06/MN/A1.1s20 tps
262k$0.09/M$0.42/MN/A0%2.9s50 tps
262k$0.09/M$0.42/MN/A0%2.9s50 tps
256k$0.53/M$2.31/M$0.11/M46.6%1.9s68.2 tps
256k$0.53/M$2.31/M$0.11/M0%N/A389.6 tps
29k$0.08/M$0.21/M$0.04/M24.2%1.2s106.2 tps
29k$0.08/M$0.21/M$0.04/M0%1.6s126.9 tps
131k$0.38/M$0.42/MN/A0%0.3s22 tps
260k$0.34/M$3.36/MN/AN/A0.3s55 tps
260k$0.34/M$3.36/MN/AN/A0.3s55 tps
262k$0.10/M$1.00/MN/AN/A0.8s68 tps
262k$0.10/M$1.00/MN/AN/A0.8s68 tps
41k$0.09/M$0.58/MN/A0%0.5s7 tps
41k$0.08/M$0.29/MN/A0%0.5s22 tps
256k$0.09/M$1.16/MN/A0%0.4s68 tps
258k$0.47/M$3.15/M$0.23/MN/A0.3s37 tps
258k$0.47/M$3.15/M$0.23/MN/A0.3s37 tps
262k$2.10/M$6.30/M$0.21/MN/A1.9s77 tps
262k$2.10/M$6.30/M$0.21/M7.1%1.9s77 tps

Prices are per million tokens and include the 5% provider-selection markup. Availability and pricing refresh continuously; each model page shows the live provider comparison.

Using DeepInfra via the API

Append :deepinfra to the model ID, set "provider": "deepinfra" in the request body, or send an X-Provider: deepinfra header. Explicit provider selection adds a 5% markup over that provider's base price; the prices in the table above already include it.

curl https://nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-R1-0528:deepinfra",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the API documentation for provider preferences, price-aware routing, and error behavior.