Nano GPT logo
NanoGPT

Private AI

Explore Text Models

Browse and discover the best AI language models for conversations, coding, and creative writing.

qwen

Qwen3.8 Max

Qwen3.8 Max is Qwen's 2.4T-parameter flagship model for advanced reasoning, coding, knowledge work, full-stack development, data analysis, and long-running agent workflows. It supports text, image, and video input, selectable thinking effort, tool calling, structured output, and a near-million-token context window.

Features

Context

991.0K

Max Output

131.1K

Date Added

Aug 3, 2026

Pricing

Input:

$2.00/1M

Output:

$6.00/1M

Cache:

Read $0.25/1M · Write $2.50/1M

Est./msg:

$0.0050

Subscription

Not included in subscription

Try Qwen3.8 MaxDetails
deepseek

DeepSeek V4 Flash Latest

Compatibility alias that routes to the newest dated DeepSeek V4 Flash release. Currently routes to DeepSeek V4 Flash 0731. ⚠️ This route goes directly to DeepSeek, so privacy and logging guarantees are limited.

Features

Context

1.0M

Max Output

384.0K

Parameters

284B / 13B

Date Added

Aug 2, 2026

Pricing

Input:

$0.14/1M

Output:

$0.28/1M

Cache:

Read $0.0028/1M

Est./msg:

$0.0003

Subscription

Included in subscription

Try DeepSeek V4 Flash LatestDetails
deepseek

DeepSeek V4 Flash 0731 (Thinking)

DeepSeek V4 Flash 0731 Thinking enables reasoning by default on the re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.

Try DeepSeek V4 Flash 0731 (Thinking)Details
deepseek

DeepSeek V4 Flash 0731 Cheaper

DeepSeek V4 Flash 0731 Cheaper is the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route goes directly to DeepSeek to use its lower cached-input pricing. ⚠️ Privacy and logging guarantees are limited.

Try DeepSeek V4 Flash 0731 CheaperDetails
deepseek

DeepSeek V4 Flash 0731 Cheaper (Thinking)

DeepSeek V4 Flash 0731 Cheaper Thinking enables reasoning by default on the same re-post-trained Mixture-of-Experts model with a 1M-token context window. This route goes directly to DeepSeek to use its lower cached-input pricing. ⚠️ Privacy and logging guarantees are limited.

Try DeepSeek V4 Flash 0731 Cheaper (Thinking)Details
google

Gemma 4 12B Instruct

Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.

Try Gemma 4 12B InstructDetails
google

Gemma 4 E2B Instruct

Google's Gemma 4 E2B Instruct is a compact open-weight multimodal model with 2B active parameters, supporting text, image, audio, video, tools, and structured output.

Try Gemma 4 E2B InstructDetails
google

Gemma 4 E4B Instruct

Google's Gemma 4 E4B Instruct is an efficient open-weight multimodal model with 4B active parameters, supporting text, image, audio, video, tools, and structured output.

Try Gemma 4 E4B InstructDetails
deepseek

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a re-post-trained Mixture-of-Experts model with a 1M-token context window, built for coding, reasoning, and agent workflows.

Try DeepSeek V4 Flash 0731Details

Inkling Small

The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.

Try Inkling SmallDetails

Inkling Small Thinking

The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.

Try Inkling Small ThinkingDetails
moonshot

Kimi K3 TEE

Kimi K3 is Moonshot AI's open-weight, always-thinking multimodal model for long-context reasoning, coding, tool use, and native image/video understanding. Running inside a TEE (Trusted Execution Environment), with provider attestation support.

Try Kimi K3 TEEDetails

Celeris 1

Celeris 1 is a diffusion language model built for ultra-low-latency classification, extraction, judging, query rewriting, and other short structured responses.

Try Celeris 1Details
qwen

Qwen3.7 Flash

Qwen3.7 Flash is Qwen's fast multimodal model for coding, search and computer-use agents, visual understanding, object recognition, spatial reasoning, and stable end-to-end task execution.

Try Qwen3.7 FlashDetails
qwen

Qwen3.7 Flash Thinking

Qwen3.7 Flash with thinking enabled for deeper multimodal reasoning, coding, search and computer-use agents, spatial reasoning, and multi-step task execution.

Try Qwen3.7 Flash ThinkingDetails
anthropic

Claude Opus 5

Claude Opus 5 is Anthropic's flagship model for advanced coding, long-running agentic tasks, research, and complex knowledge work.

Try Claude Opus 5Details
sakana

Fugu Ultra v1.1

Sakana AI's upgraded Fugu Ultra release with stronger coding, agentic task execution, and advanced reasoning through dynamic orchestration of frontier models.

Try Fugu Ultra v1.1Details
inclusionai

Ling 3.0 Flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Try Ling 3.0 FlashDetails