Nebius Group N.V. (Nasdaq: NBIS) has acquired inference optimisation startup Inferize, adding its technology and engineering team to Nebius Token Factory as the AI cloud company expands its production inference capabilities.
The acquisition brings Inferize’s technology into Token Factory, where Nebius is building infrastructure to help customers launch and scale large AI models more efficiently.
Inferize focuses on reducing cold starts, the time AI models need to load before they can begin serving requests. These delays can leave GPUs idle when demand increases, new instances are launched, or model weights are updated during reinforcement learning, forcing platforms to maintain spare computing capacity.
Its technology is designed to reduce this idle capacity, allowing computing resources to scale more closely with actual demand while improving GPU utilisation and token economics.
Commenting on the acquisition, Danila Shtan, Chief Technology Officer of Nebius, said, “Running inference well takes more than fast GPUs and optimized models. The whole system needs to respond when demand changes, including how quickly additional capacity is ready to serve customers. Inferize brings technology that accelerates that process and a team with deep expertise in GPU systems. We’re bringing both into Token Factory to make it more responsive to customer demand and get more useful work out of our infrastructure. The team’s contribution will extend well beyond this first integration.”
The deal builds on Nebius’s broader push to expand its production inference stack. Eigen AI previously brought optimisation capabilities across model, kernel and system levels into Token Factory, while Clarifai’s core team and licensed technology added inference and compute orchestration.
Founded in January 2026, Inferize developed a working prototype within three months. Its engineers will now work across Nebius Token Factory, beginning with the integration of their technology.
Speaking about the acquisition, Guy Bortnikov, Co-founder and Chief Executive Officer of Inferize, said, “Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform. Our team will work across the stack with one objective: serving more customer demand from every GPU.”

