Add Suno AI
It would be great to have suno available for music generation in nano-gpt. I love your work!
Share ideas and upvote the ones you want most. Bug reports stay private between you and our support team.
It would be great to have suno available for music generation in nano-gpt. I love your work!
I'm primarily using PayPal, so having it as a payment option would be nice
We're working on this but this is proving quite a bit more difficult than we had hoped. Sorry - we'd love to add this in ourselves as well! - Milan
I remember that NanoGPT used to advertise "zero data retention" agreements with their LLM providers. At some point this language has been removed (or I'm misremembering) and replaced with "minimum data retention", which is a far more nebulous term, though even "zero data retention" does not have a standardized definition among providers (especially regarding abuse detection). Can you disclose which providers you have special data retention agreements with and what those terms are? It would go a long way to help the transparency of this platform.
We’ve added provider-level privacy information to the model provider lists, including whether providers retain prompts or may use them for training, along with links to their policies. We also offer Zero Data Retention routing for text requests.
ZDR reflects provider commitments; it isn’t an independent audit of their systems. Our privacy explanation is at https://nano-gpt.com/privacy.
As a user, I'd like to see more detailed information for each log in the Usage page, including provider, latency, throughput of each log/usage.
We’ve added more detail to individual Usage entries. Open an entry to see time to first token and full-response latency when those measurements are available, alongside its billing details. We also display the provider, except if you use auto routing.
Plus let me attach my referral for new account creations via this module
Thanks for the suggestion, this is now live! Add "Sign in with NanoGPT" to your app with our OAuth flow and pass your referral code as invitation_code (the part after /r/ in your referral link). People who create a new NanoGPT account through your sign-in get you as their referrer. They see this on the consent screen and can untick it, they get 5% off, and you earn the standard 10% referral share of their eligible spending. Existing NanoGPT accounts keep their current referrer. Docs: https://docs.nano-gpt.com/api-reference/miscellaneous/oauth-pkce#referral-attribution
Hello NanoGPT team! A suggestion I have is to provide vouchers like Mullvad does so people can deposit money anonymously without needing crypto. I didn't want to deal with the tax headache of cryptocurrencies, so I used the stripe process to add funds. I was under the impression that personal information would not be knowable by NanoGPT, but I can see my personal information under billing so you surely must have access, right? Vouchers would enable a clean separation for those like myself which do not want to rely on trust.
We’ve reached out to services such as Bitrefill about offering NanoGPT vouchers, but unfortunately haven’t heard back yet. We don’t have an availability date to share at this point.
We really do want to minimize the information that we see, but unfortunately, Stripe does not actually have a setting or option to see even less than we do now. We do see, indeed, the name you fill in and such there (within Stripe, it does not reach our own database).
Many providers do offer prompt caching nowadays and utilizing that yields great discounts. Unfortunately, through NanoGPT, we can't benefit from those (except with Claude, which is the most inconvenient to implement). That remains a big advantage of OpenRouter. For me personally, the most interesting providers are Moonshot AI, Google (Gemini) and DeepSeek. I am aware there are some suggestions like this already, but they can't be voted on, so this is my way of expressing demand. I'm also aware this might hurt your profits, if you only pay discounted rates while you charge base costs, but that would also mean the pricing isn't fair.
We now haev this for Gemini as well - Milan
if the 2x token use on model name is not possible, please add a double price tier of sub where 2 x token models are 1x. can be the same limits as the current, but it would eliminate the uncertainty of selecting models, I am sure many of us will happily pay the extra.
Would be nice to add Google drive files straight to the model instead of downloading first
We have this working - you can actually already try it now. We however need to do extra verification with Google, since otherwise every user that tries to use this is shown a security warning which is quite a pain (essentially Google's way of saying "this company did not yet prove what they want to use this for". This might take a while, we've sent in our side but are now dependent on Google reviewing. We're leaving it up in the meantime for people to try out.
Or gift cards/voucher codes that can be purchased with cash from somewhere.
Unfortunately, accepting cash by mail isn’t something we want to offer, so we won’t be implementing this payment method. We’ve reached out to services such as Bitrefill about offering NanoGPT vouchers, but unfortunately haven’t heard back yet. We don’t have an availability date to share at this point.
Primarily: Text LLM models already have the ability to be Favorited/Star'd. Add the same functionality to the UI for Image models as well. And videos as well, while you are at it. Appears to be the same UI element (?) so should be fairly simple. Secondarily, it's tough to search through the Image model menu for the 4 or so models that are included with the sub. Add the toggle to Show Only Subscription Models like you already have for Text LLMs. :)
hi, I don't know why it took us this long to do it, but we are doing the favoriting for image and video models now as well. I think that the image model selector now also respects the subscription setting for hiding paid models.
I noticed your site and API are behind the Cloudflare proxy/CDN. Cloudflare inspects and routes all visitor traffic, which can expose users’ IP addresses and request data to a third party and centralizes metadata that could be retained or accessed. For better privacy and to avoid third‑party interception of user requests, please consider removing the Cloudflare proxy (or configure DNS only / self‑hosted CDN) so visitors connect directly to your origin servers. Thank you.
You’re right to flag the privacy implications of reverse proxies/CDNs and we’re evaluating options to reduce third-party exposure where feasible.
For clarity: NanoGPT is not currently behind a Cloudflare reverse proxy. We use our hosting provider’s global edge network (running on AWS) for performance and DDoS protection, which means AWS/our edge provider may still process similar connection metadata.
title says it all, wondering if we'll get it in the sub now that it's opensourced, since we have glm 5.1 but not yet mimo, wanted to test that it's allegedly the current top dog
Hi,
So we'd love to, but there are no open source providers actually hosting the model so far. So it is open source, but no one seems very enthusiastic about offering it to us.
Kind regards,
Milan
Allow users to import lorebooks, and attach them to specific conversations. Lorebooks are a very simple form of RAG where each data entry specifies a list of keywords. The algorithm checks the last n messages for these keywords, where n is a user-adjustable parameter. If a keyword is found, the corresponding data entry is inserted into the context, typically below the system prompt but above the chat content. Use case: Most lorebooks are created for roleplaying uses, so it's appropriate to add support for them since the site already provides many RP-oriented models.
I understand that you rely on many different providers for models, including open-source models in the subscription. However, it isn't clearly written anywhere which models are hosted by which providers. (Or maybe I'm just not looking hard enough.) Any chance you could disclose this information somewhere on the site? Like, "Primary provider for X is Fireworks, fallback Chutes", as an example. Adding a note about it for the "Available Models" list under the "</> API" tab would feel intuitive, for example.
Hi - we have added provider selection for (most of) the open source models now, so you can select your provider of choice.
It's also still possible to use the "auto" routing still, which remains the same as it was before. For that, we do not display the provider and likely will not.
The discord room should be bridged to an XMPP or Matrix group chat. Lots of privacy and freedom conscious people stay away from user-hostile software & company like discord, and thus, they are unable to participate in the user community discussion with the nano-gpt developers.
Sometimes a reasoning model gets interrupted, and does not finish the task. Currently i make them continue by sending a message that just says "continue". Maybe it's better to provide a separate button to continue generating.
I read situation about those 3-5% users who caused abuse of service and unlimited tokens, and understand introduction of weekly limits fo 30mil tokens. But even among those 95% people can be few % of loyal users who may need extra tokens per week (especially when using agents or reasoning models). I would like to suggest a feature to introduce weekly limit size depending on how long user has been subscribed. It may be flat bonus after X months (i.e. 60m after 3-6 months) or incremental (i.e. +3-5mil with each month). Kinda a reward system for loyalty. And abusers usually will create a lot of new accounts and probably would not patient enough to cause issues.
Hi,
Sorry, but we do not intend to do this. It essentially just makes the subscription even more unprofitable for us, sorry.
Would be nice if there was an anonymous way to access this via Tor. Probably would have to disable the free model for that, though.
We are still testing the implementation, but you can now access the NanoGPT API directly through the Tor network.
Onion address
http://nanogpt5wwal7shnauuilouj3gaewmk7lskqxdmxlbfz5ho4tdqeh2id.onion
API base URL
http://nanogpt5wwal7shnauuilouj3gaewmk7lskqxdmxlbfz5ho4tdqeh2id.onion/api/v1
Your existing NanoGPT API key works as usual. Connections to the onion service remain within the Tor network and do not pass through Tor exit nodes.
Please test it with your usual clients, models, and streaming requests. If something breaks, let us know which client and endpoint you used.
Is there a way to see supported TPS for each model like OpenRouter displays?
We now do this for most (if not all?) models. - Milan
More filter options for models (all of them) would be useful, such as "can be coaxed easily into NSFW", "does not guarantee privacy", etc...
We’ve added privacy information to the model browser, including a “ZDR available” filter and provider badges covering retention and training policies. The filter shows models with a zero-data-retention provider available; it doesn’t itself change routing. This covers the privacy part of your suggestion, although we don’t currently offer a filter rating how easily a model can be prompted into NSFW responses.
Anima is a very tiny (2B) text-to-image model from CircleStone Labs. (https://huggingface.co/circlestone-labs/Anima) It is small but quite powerful at 2D image generation. But, it has a non-commercial license. So you need to request a license from them. I know this could be difficult, but Anima is really good (even in preview). I hope you add the Anima model in NanoGPT.
Since NanoGPT is privacy-centric it would be nice to have the additional option of being able to use an encrypted messenger (or two) for contact information in addition to email or Discord
Thanks, but we are not planning to do this, mostly because it gives us even more channels we need to keep up with. We would recommend either a ticket on the website, which is fully anonymous, or some anonymous email to contact us from.
Kind regards,
Milan
They're open source so it would make sense.
We want to do this and hopefully can do so longer term, but for now there just are not enough providers yet that do TEE to make this economically viable. TEE models also tend to be more expensive in general, so we'd likely have to charge more for this.
Please consider adding generation settings such as Min-P, this is really helpful since most of prompts and presets for roleplay comes along with generation settings and its bad that nanoGPT API does not support Min-P setting.
Hi! We already supported most of those but didn't have it in the documentation properly. https://docs.nano-gpt.com/api-reference/endpoint/chat-completion has now been updated!
Can you contact the OpenCode team so we can use all our subscription models, as they aren't currently listed.
Could it be that you are using api.nano-gpt.com? As far as we know they just use the v1/models endpoint which should return all. But the URL should be nano-gpt.com/api/v1. - Milan
As open source models are becoming an alternative for commercial models, especially for coding, the reliability needs are also increasing. The subscription provides a good way to have access to such models. For normal chat completion requests, "cheap" providers (your backend providers) are usually sufficient. But for coding, high reliability is needed (reasoning, tool calling). Some providers apply settings to the models that make them nearly impossible to use for coding. Therefore I suggest a pro or coding subscription with the same / similar limits like the normal subscription, but with providers who don't heavily modify the model settings (usually the model dev providers).
We are unlikely to do this or if we do it would be much more expensive - frankly we'd suggest looking at other providers for this because it's not something we are likely to do in the short term.
Qwen3 Embedding 0.6B A powerful and cheap Embedding model that is also open source and available here: https://huggingface.co/Qwen/Qwen3-Embedding-0.6B, it would also be nice to obtain vectors without spending API requests.
Sorry, but we're not going to add embedding models to the subscription. - Milan
Hello. Sorry to bother you, but when I roleplay, I try to budget my usage by setting a limit on the api key. But I use X2 multiplier models and because of that, I end up using double what I want. Could it be possible to have double multiplier counted in the token limit? Thank you.
We've done this now!
I'm not sure how good it will be, as the current proposed models are already pretty cheap
I think it would be a good addition to include details about when a model was last updated or uploaded. Why? For example, the recent appearance of the Mistral Large 3 model. It appears in the available models list, but I'm not sure if it's the new version or an older one. While it might have "2512" at the end of the model name, this doesn't guarantee a specific date, as the other existing models don't have anything like that.
We've now also added the date-added into the hover-over description, as well as in all other places that could be relevant!
Allow finer control of API key usage, including daily RPD and daily NANO consumption allowed.
It is now also possible to set a daily token limit
You guys are a little too fast at adding new features to the website/API. I often only find very useful features (like this new suggestion board) months after they launch because they don't seem to be announced anywhere other than your Discord. Can you please merge what you announce on Discord and the updates section on the website into a unified changelog? Ideally this would be backdated to include everything you've added thus far as well.
https://nano-gpt.com/updates does something like this work for you or were you thinking more changelog also of sort of, code changes/backend changes? - Milan
Hello! Z-Image Turbo is super cheap to run and is good enough quality for most of my needs, where the other models in the subscription don't quite meet my requirements. It would be super awesome if that model could be included in the subscription please! Thank you
This is now included in the subscription! - Milan
Would it be possible to add embedding models to the subscription?
Hi - we're not going to do this, there are not enough providers hosting embedding models for this to make sense for us to do. Sorry :/ - Milan
i want to use glm 5.1 as subscription. otherwise there is no reason to subscribe. i u
GLM 5.1 and GLM 5.1 Thinking are now included in the subscription model selection. You can find both in the model picker with subscription models selected. Normal subscription usage limits still apply.
For those of us less organized, it would be nice to have option in all conversations tab to select multiple chats from a given page and have the option to apply a function from the three dot menu that you would normally only be able to apply for a chat one at a time and not even be able to select another without returning back to all conversations screen.
Yes, this is perfect thank you so much!
It would be cool to have a badge in the model list for open-source / open-weight models, similar to the (old?) badges for vision-enabled models, or maybe another way to highlight them (colored model names perhaps), as the description for models can’t really be seen on a mobile device. Why? Because I like to run open models most of the time and consider them as having better privacy, but still want to be able to use proprietary models should I want to. It’s a bit of a pain to know whether they are open source or not. This might clutter the UI, so I guess it can also be an opt-in/opt-out setting?
We’ve added an “Open source” badge and an “Open source only” filter in the text-model visibility settings. This makes it easier to identify those models without relying on hover descriptions. Open weights don’t by themselves guarantee provider privacy; the provider privacy labels give separate information about retention and training.
Gemini natively supports grounding/google search api and is arguably more powerful than third-party search tools such as the supported tavily or linkup tools. Currently, it cannot be used through the nano-gpt api Supporting the native grounding api would allow for better search performance additionally, supporting more tools like google's code execution tool would allow for more use cases, matching the official api's capabilities
Models like GLM 4.7 and MiniMax M2/M2.1 now recommend passing `reasoning_content` or `reasoning_details` back into the context for enhanced reasoning performance. See https://docs.z.ai/guides/capabilities/thinking-mode and https://platform.minimax.io/docs/guides/text-m2-function-call. Is this possible with NanoGPT's OpenAI-compatible API? Or does it perhaps need an Anthropic-compatible API? And if so, can I suggest some clarity in the docs on how this might work?
Hi,
Really very late reply, but from everything we can see it should now work properly.
I would love to see Ministral 8b and Ministral 3b added, as well as the Mixtral 8x7b & 8x22b Models, if that would be possible
Ministral 3B, Ministral 8B and Mixtral 8x22B are now available on NanoGPT. The Ministral entries are the December 2025 versions. Mixtral 8x7B is still missing, so this request is only partly covered, because we can't find a host for it.
Your application already supports over 500 LLMs, and the Custom CivitAI update got me thinking - it would be incredible if users could access HuggingFace models through NanoGPT. Because RAM and GPU prices are rising, an AI rig is impossible for the average consumer, myself included. Adding this feature would be perfect because users could access more models and use them without getting expensive hardware. The basic idea would be to find a HF URL, like https://huggingface.co/mradermacher/Huihui-GLM-4.7-Flash-abliterated-i1-GGUF (a currently inaccessible model) that contains GGUFs and then to paste and submit it. Like Custom CivitAI, the LLM would have to be resolved to be used, but once finished you could start using it. Likely all users would enjoy this feature.
Thanks! We'd love to add something like this but as far as we know this is not possible to add in a cheap way - as in while this is possible for image models there is no "load on demand" for text models. Maybe we are wrong - if you did find it somewhere we'd definitely love to know! - Milan
Some of the subscription models take forever to start replying. Would be nice to know how fast they are in advance.
Model details now show average generation speed (TPS) and time to first token (TTFT) where we have performance data. You can use these to compare subscription models before sending a message. Provider lists also show speed and latency for routes with available data. Actual performance can vary with load and your request.
Maybe people will pay for it, monthly limits can be toned down for those two models.
We will not do this, sorry. It's simply way too hard to figure out how much we would need to charge, and we'd have a lot of risk that people use it too much. We understand it's very attractive for users hah, but it's not for us. - Milan
It would be great to add all Qwen3 embedding and reranking models to the subscription: Qwen3-Embedding-8B Qwen3-Embedding-4B Qwen3-Embedding-0.6B Qwen3-Reranker-8B Qwen3-Reranker-4B Qwen3-Reranker-0.6B Currently, such models are becoming increasingly necessary for codebase indexing and RAG tasks. As far as I have seen, competitors already offer these models within their subscriptions.
Hi - we're not going to do this, there are not enough providers hosting embedding models for this to make sense for us to do. Sorry :/ - Milan
Sup.Ai is doing something really cool that I bet you guys could implement. They give the user the ability to submit to multiple models, which you have, but instead of just returning the results, they pass all three results through a model to evaluate the responses and construct a consensus using the best aspects of each response. Then, they return the final consensus to the user. The idea is that while one model might hallucinate or write poorly, it's unlikely that they all would at the same time. I would love to, for example, be able to combine Deepseek, GLM and Kimi K2 or something like that. The biggest trick, here, would be how you prompt the final model to generate a useful consensus, but if you pulled it off, it would be incredible. Another advantage is that if a model is unavailable or unwilling to complete the task, another model will, so the user never experiences a dropped response.
We’ve added this as Fusion. Enable multi-model mode, choose the models you want to compare, then switch on Fusion and choose a Synthesizer model. It combines the model responses into a final answer. This can help compare and reconcile responses, but the combined answer can still make mistakes.
Either on a per-conversation level (where the web-search/thinking/memory toggles are), a per-preset level, globally, or through prompt tag eg <CURRENT_TIME> (in the manner you already have <CURRENT_DATE>) Plz I beg, I need my character to stop hallucinating timestamps and arguing when I say it's X time but they think it's Y. And it's a big boost for immersion in general when prompt engineering ❤️
Sorry, but we don't really like to put this into the Settings/system prompt and such. Mostly because it feels like something that is quite.. specific in a way, and every option we add adds complexity and takes up space. Sorry :/ - Milan
Please add an option to protect past conversations with pincode or 2FA
You can now protect individual conversations with a passkey. Open the conversation title menu and choose “Protect with passkey.” Viewing the conversation then requires passkey verification, and it locks again after the window has been inactive for five minutes. This uses your passkey rather than a separate NanoGPT PIN.
Automatically choose the aspect ratio for generated image based on the original image.
Thanks! This should happen now. - Milan
I want to make the API pricing easier to understand. This includes showing the price for image inputs and for video generation with or without audio. Some models do not have pricing, so I also want to clearly show which features each model can access, such as whether image input is allowed.
Suggest it and let the community vote on it.