Neocloud darling Nebius has acquired an inference optimization company whose technology it claims can shorten the time to launch large-scale AI models.
Only founded in January 2026 by Granulate alums, Inferize will join the neocloud’s managed inference platform, Nebius Token Factory.
No financial details were disclosed.
“Running inference well takes more than fast GPUs and optimized models. The whole system needs to respond when demand changes, including how quickly additional capacity is ready to serve customers,” Nebius CTO Danila Shtan said. “Inferize brings technology that accelerates that process and a team with deep expertise in GPU systems.”
Inferize was co-founded by Lior Gorbonos and Guy Bortnikov, who previously founded an Israel-based software optimization firm sold to Intel in 2022.
Granulate’s software optimizations reduced CPU utilization and application latencies to boost performance, with the pair taking their learnings and applying them to AI inference.
Inferize’s tech will now be integrated into Nebius Token Factory, which the neocloud has sought to iterate through acquisitions. Earlier this year, it snapped up Eigen AI, another optimization firm that added autoscaling endpoints and fine-tuning pipelines.
Among Inferize's innovations are means to cut cold starts – the time models need to load before serving a single request, thereby improving capacity utilization.
“Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform,” Bortnikov said in a statement. “Our team will work across the stack with one objective: serving more customer demand from every GPU.”
Comments