Aoru AI

Aoru AI

KR
33 models available

Models served by Aoru AI

ModelContextInput /MOutput /MCache read /MCache hitLatencyThroughput
262k$0.13/M$0.40/M$0.06/M11.3%5.1s15.4 tps
262k$0.13/M$0.40/M$0.06/M40.6%5.6s19.5 tps
262k$0.11/M$0.47/M$0.05/M17%25s15.6 tps
1.0M$0.16/M$0.63/M$0.08/M96.7%2.1s90.7 tps
Fabled
NVFP4
262k$0.11/M$0.47/M$0.05/M50%10.5s16.4 tps
Garnet
NVFP4
262k$0.11/M$0.47/M$0.05/M10.5%15.8s13.4 tps
262k$0.11/M$0.47/M$0.05/M8.6%9.1s18.7 tps
131k$0.05/M$0.26/M$0.03/M9.7%1.3s12.5 tps
131k$0.05/M$0.26/M$0.03/M0%1.1s16.5 tps
131k$0.05/M$0.26/M$0.03/M14%12.4s10.2 tps
262k$0.13/M$0.40/M$0.06/M8.6%10s14.2 tps
262k$0.13/M$0.40/M$0.06/M2.6%8.4s14.3 tps
262k$0.13/M$0.40/M$0.06/M0%18.9s14.3 tps
262k$0.11/M$0.47/M$0.05/MN/AN/AN/A
262k$0.11/M$0.47/M$0.05/M17.7%11.1s15.2 tps
262k$0.11/M$0.47/M$0.05/M15.8%8.7s15.7 tps
262k$0.11/M$0.47/M$0.05/M20.9%7.4s17.3 tps
262k$0.11/M$0.47/M$0.05/M17.3%7.9s15.9 tps
262k$0.11/M$0.47/M$0.05/M9%12.5s14.5 tps
262k$0.11/M$0.47/M$0.05/M0%12.4s16.2 tps
262k$0.11/M$0.47/M$0.05/M4.1%12.2s14.3 tps
262k$0.11/M$0.47/M$0.05/M3.7%10.5s14.5 tps
262k$0.13/M$0.40/M$0.06/M38.7%9.8s17 tps
262k$0.13/M$0.40/M$0.06/M0%4.5s19.1 tps
262k$0.13/M$0.40/M$0.06/M0%39.9s2.9 tps
262k$0.11/M$0.47/M$0.05/M0%33.3s14.7 tps
262k$0.13/M$0.40/M$0.06/M0%17.3s21.3 tps
262k$0.11/M$0.63/M$0.05/M82.3%4.1s39.4 tps
262k$0.11/M$0.63/M$0.05/M67.2%2s43.4 tps
262k$0.11/M$0.63/M$0.05/M8.6%1.8s60.7 tps
262k$0.11/M$0.63/M$0.05/M75%1.4s58.7 tps
262k$0.13/M$0.40/M$0.06/M0.9%9.3s17 tps
262k$0.11/M$0.47/M$0.05/M25.4%8.4s15.6 tps

Prices are per million tokens. Selectable routes include the applicable provider-selection markup; fixed routes show the standard model price. Availability and pricing refresh continuously; each model page shows the live provider comparison.

Using Aoru AI via the API

For models with provider selection, append :aoru to the model ID, set "provider": "aoru" in the request body, or send an X-Provider: aoru header. Explicit provider selection adds a route-specific markup over that provider's base price; the prices in the table above already include it. Models without provider selection use the displayed model ID without a provider suffix or selection surcharge.

curl https://nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-26b-a4b-it-chimerax:aoru",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See the API documentation for provider preferences, price-aware routing, and error behavior.