![]()
vCluster Labs, the platform operators use to build and run their own AI cloud like a hyperscaler, today announced the general availability of Stacks in vCluster Platform v4.12, which automate the long list of ordered setup steps required before a managed service can go live, so an operator can deploy the same environment to every tenant.
The limit on how many managed services an operator can offer is not GPU supply. It is the setup work between an empty tenant cluster and an environment that is ready to use, stitched together by hand and repeated for every tenant.
“Every operator runs into the same wall,” said Lukas Gentele, Co-Founder and CEO of vCluster Labs. “Once the cluster exists, the real project starts: registry credentials, ingress, certificates, GPU components, a control plane, and registration, all wired together by hand, in the right order, again for every tenant.”
A Stack is that environment described once: which applications, in what order, what has to be healthy before the next thing starts, and which parameters the person deploying it is allowed to set. Deploy it and a tenant gets a running environment. Deploy it a hundred times and the hundredth costs what the first one did. A new managed service becomes something an operator ships rather than something they staff.
Stacks make new services faster to launch and easier to operate
A new managed service can go live in days instead of months, because the work is authoring one template rather than staffing an install for every tenant. Each additional tenant is a deployment rather than a project, which is how a GPU fleet turns into higher-margin managed products instead of raw capacity sold by the hour.
Operating them is the other half. Health, logs, upgrades, and clean tenant removal are handled by the platform, and because every tenant runs an identical environment, the team supporting the fiftieth tenant is troubleshooting the same configuration it learned on the first. Operators add new tenants without adding people to keep each environment running.
Available today
Stacks ship with ready-made examples covering the environments operators are standing up now:
- NVIDIA Run:ai, in four certified configurations covering both a dedicated control plane for a single cluster and a shared control plane serving many tenant clusters.
- NVIDIA Dynamo, standing up a distributed inference environment inside a tenant cluster so an operator can offer serving for large models that span many GPUs and nodes.
- Saturn Cloud, installing the AI token factory platform inside a tenant cluster so an operator can offer per-token inference, fine-tuning, and usage billing on GPUs they already own.
“Each of these is a service an operator can put in front of customers without building the integration first,” Gentele added. “That is the difference between selling capacity and selling products, and you add one without adding a team to run it. The repository is open, so the next one does not have to come from us.”
“Adding a service to an AI cloud shouldn’t mean starting a new integration project each time,” said Hugo Shi, co-founder and CTO of Saturn Cloud. “vCluster built the Saturn Cloud Stack, so an operator can deploy it like any other service. Each customer gets private inference in an environment of their own, ready to use right away, and a small team can keep adding services without adding headcount.”
Anyone can author a Stack
Stacks are not limited to integrations vCluster Labs builds. Any team whose software runs inside customer clusters can describe a correct installation once and have it land identically everywhere it is deployed. A public repository is open for exactly that, so a team that has already solved this can hand the next team a starting point.
Stacks are generally available today in vCluster Platform 4.12 with vCluster 0.37.
Supporting resources
About vCluster Labs
vCluster Labs is the platform operators use to build and run their own AI cloud, like a hyperscaler. vCluster turns raw GPUs into cluster products they can sell, from bare metal provisioning up through managed Kubernetes, Slurm and inference clusters, with tenant isolation at the infrastructure layer. The result is more margin per GPU, less dependence on a few big customers, and new managed services live in days instead of months. vCluster is trusted by fast-growing AI cloud providers including Nebius, Groq, Firmus and Corvex across 100K+ GPUs, and by enterprises including Adobe, Samsung and Deloitte.
View source version on businesswire.com: https://www.businesswire.com/news/home/20261008868645/en/
Media gallery