Share your ideas about our platform
Add Country Philippines to the Country list
I’ve been trying to add my billing details inside the Nebius Token Factory, but it won’t let me since Philippines was not an option in the country of residence. I’m doing this while taking the course for Nebius Agentic AI. So if the team could make this an option, it would be very much appreciated. Thanks!
Cilium Cluster Mesh
Based on Nebius team While ClusterMesh is not currently supported in Nebius Managed Kubernetes as a standard product feature, Nebius MK8s uses Cilium as the managed CNI, and the Cilium ConfigMap can be inspected and edited in the kube-system namespace. ClusterMesh would be treated as an add-on, changes to cilium-config or deploying the cilium-clustermesh-apiserver daemon are treated as custom configuration that might affect the managed upgrade path for Cilium on MK8s. It would be great to officially support ClusterMesh, thanks! ☺️
Report CO2 Emissions for Token Factory
We need to report the co2 emissions caused by our AI models usage, therefore we need to get this information on a monthly basis (or yearly).
Usage and Balance API on Token Factory
Please add an API for viewing live usage with daily breakdowns and the ability to query our current balance as I can’t seem to find an auto top-up function on Token Factory. Thank you!
Support for newer PostgreSQL versions (17 and 18)
Nebius currently offers PostgreSQL 16 as the latest version in its Managed Service for PostgreSQL. Adding support for PostgreSQL 17 and 18 would significantly lower the barrier for teams migrating existing instances to Nebius and keep the platform aligned with the upstream release cadence.
Account Balance API endpoint for Nebius Token Factory
Please add an API endpoint to retrieve the current account balance for a Nebius Token Factory account
Offer separate short-term and long-term support LLM endpoints
I understand that there might be different needs among users of the token factory: build stable applications with reliable LLM performance → requires long-term support (LTS) for the model endpoint (even if the model is somewhat outdated) use the latest and greatest model as soon as it has been released, because it offers the best performance per dollar/euro → only needs short-term support (STS) for the model endpoint Public model endpoints should clearly state “STS” or “LTS”. This way the users can pick depending on individual needs. On Nebius’ side this would allow superseding STS model endpoints without further notice as soon as a superior version of that model gets released - ultimately offering a better service and keeping the “zoo” of model endpoints manageable in the long run. LTS model endpoints on the other hand could have a predetermined EOL date, so developers can plan for necessary model updates.
Support Implicite/Explicite Prompt Caching
Many other providers offer prompt-caching with discounted pricing for cache hits, either implicitly (e.g., OpenAI, DeepSeek, Google, DeepInfra, NovitaAI, Fireworks) or explicitly (most notably Anthropic). This capability can significantly reduce costs in agentic workflows, where a single session often re-sends the same context repeatedly (for example, when the model performs multiple tool calls in sequence and the shared conversation/context is included each time). Today, Nebius Token Factory is at a cost disadvantage in these repeated, input-token-heavy scenarios compared to providers that support prompt caching and pass the savings through to customers. Please add support for prompt caching (implicit or explicit), including discounted pricing for cached prompt tokens, to improve cost-efficiency for agentic and tool-using applications.
Support Latest Top 5 LLM Models
Out of the top 5 models (based on LLM and artificalanalysis) token factory still lacks: - MiniMax M2.5 - GLM 5 - Qwen3.5-397B-A17B - Step 3.5 Flash Additionally MiMo-V2-Flash would be welcome as Step 3.5 Flash only supports 64k token window. I’m quite confident that providing these models in the token factory catalog would provide users frontier proprietary level models, which would greatly help adoption. Additional note: On the main website (nebius.com) when someone hovers over the token factory menu the models that appear in the popup are all relatively old. K2.5 is already provided in token factory, but only K2 is listed. (1) https://llm-stats.com/leaderboards/open-llm-leaderboard (SWE-bench Verified) (2) https://artificialanalysis.ai/models/open-source (intelligence ranking)
/v1/responses OpenAI compatible endpoint
The /v1/responses endpoint from OpenAI is used by tools like Kilo Code and Opencode, and it is particularly useful for agent development and handling large files. I noticed that Nebius does not currently support this, but I believe it would be a valuable addition. Maintaining compatibility with the official OpenAI API would enable Nebius to serve as a backend for tools like Opencode and similar applications.
Larger FAST models with SO.
Currently SO by constrained decoding is supported by large models (>80B) only with tps <50. This is insanely slow for agentic architectures based on SGR, and you keep slowing down inference. Please provide at least one fast, large, good model with SO
Support of PostGIS
Hello, I noticed that your PostgresSQL managed service supports many extensions, but PostGIS isn’t listed (https://docs.nebius.com/postgresql/databases/extensions). Would it be possible to add support for PostGIS (https://postgis.net/)? Thanks!
Be nimble with new open source models that lead the benchmarks: Support Kimi K2.5 in Token factory
If a new open weights / open source models drops and it leads the benchmarks (look https://artificialanalysis.ai/models in the open section), be nimble and quick to offer it in the token factory
Token Factory API: Lowercase model IDs
In the Token Factory API, the model selector should be able to be specified as lowercase. For example: Required today: {"model":"Qwen/Qwen3-Coder-480B-A35B-Instruct", … } Wanted: {"model":"qwen/qwen3-coder-480b-a35b-instruct", … } Why? Currently, OpenCode does not work with any Nebius model, even if OpenCode have a built-in integration with Nebius. Nothing works. Any request says “Not found”, and that is an error from your API because they (mistakenly) force lowercase on all model IDs. If I edit the OpenCode config file and Pascal Case the Nebius model IDs, then your API accepts it. This is also easy to test with curl. This is really an OpenCode issue, but maybe you could also edit your API and allow lowercased IDs, for us developers (customers) sake 🙏 Regards, Christoffer
Serverless Lora for Gemma3
Would it be possible to support fine-tuning and serverless LoRA for Gemma3-9b/27b?
Cloud DNS
Hello! I would like to be able to use Cloud DNS, primarily internal DNS for internal resources and for the team that accesses internal resources via VPN (i.e. there should be a way to specify DNS servers).
Add Country Ghana to the Country list
Under the Billing setup, there is no option for my country Ghana.