Sail Research

Sail Research

GLB
15 models available

Models served by Sail Research

ModelContextInput /MOutput /MCache read /MCache hitLatencyThroughput
262k$0.08/M$0.36/M$0.02/M47%1.4s27 tps
262k$0.08/M$0.36/M$0.02/M47%1.4s27 tps
1.0M$0.73/M$3.11/M$0.03/M62.5%2.1s16 tps
1.0M$0.73/M$3.11/M$0.03/M33.5%2.1s16 tps
262k$0.15/M$0.42/M$0.07/MN/AN/AN/A
Gemma 4 31B
Standard
FP4
262k$0.07/M$0.42/M$0.05/MN/AN/AN/A
262k$0.15/M$0.42/M$0.07/MN/AN/AN/A
262k$0.07/M$0.42/M$0.05/MN/AN/AN/A
GLM 5.2
ASAP
FP8
1.0M$1.47/M$4.62/M$0.27/MN/AN/AN/A
GLM 5.2
Priority
FP8
1.0M$0.73/M$3.15/M$0.19/MN/AN/AN/A
GLM 5.2
Standard
FP8
1.0M$0.53/M$2.63/M$0.13/MN/A1.9s16 tps
1.0M$1.47/M$4.62/M$0.27/MN/AN/AN/A
1.0M$0.73/M$3.15/M$0.19/MN/AN/AN/A
1.0M$0.53/M$2.63/M$0.13/MN/A1.9s16 tps
GLM 5.3
Standard
FP8
1.0M$1.27/M$4.06/M$0.24/M55.7%1.2s53 tps
1.0M$1.27/M$4.06/M$0.24/M2.9%1.2s53 tps
GLM 5.3 Flash
Standard
FP8
1.0M$0.15/M$0.50/M$0.03/M85.9%1.7s11 tps
128k$0.06/M$0.42/M$0.03/MN/AN/AN/A
GPT OSS 120B
Priority
128k$0.04/M$0.31/M$0.02/MN/AN/AN/A
Kimi K2.6
ASAP
FP8
256k$1.05/M$4.20/M$0.21/MN/AN/AN/A
Kimi K2.6
Priority
FP8
256k$0.47/M$3.15/M$0.21/MN/AN/AN/A
256k$1.05/M$4.20/M$0.21/MN/AN/AN/A
256k$0.47/M$3.15/M$0.21/MN/AN/AN/A
Kimi K3
Standard
FP4
975k$2.25/M$11.30/M$0.26/M92.2%2s51 tps

Prices are per million tokens. Selectable routes include the applicable provider-selection markup; fixed routes show the standard model price. Availability and pricing refresh continuously; each model page shows the live provider comparison.

Using Sail Research via the API

For models with provider selection, append :sailresearch-standard to the model ID, set "provider": "sailresearch-standard" in the request body, or send an X-Provider: sailresearch-standard header. Explicit provider selection adds a route-specific markup over that provider's base price; the prices in the table above already include it. Models without provider selection use the displayed model ID without a provider suffix or selection surcharge.

Sail Research offers multiple service lanes, each pinned with its own id: sailresearch-standard (Standard), sailresearch-asap (ASAP), sailresearch-priority (Priority). The table shows which lane serves each row.

curl https://nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4-flash-0731:sailresearch-standard",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the API documentation for provider preferences, price-aware routing, and error behavior.