The Zhitong Finance App learned that AI cloud infrastructure service provider Nebius (NBIS.US) announced the acquisition of Israeli artificial intelligence startup Inferize to further enhance its AI inference infrastructure capabilities. Inferize focuses on reducing graphics processing unit (GPU) idle time and speeding up the deployment of large AI models. According to estimates by Israeli tech media Calcalist, the transaction amount was between 100 million and 150 million US dollars, but Nebius did not disclose specific financial terms.
The core of this acquisition is to improve the efficiency of GPU resource utilization, especially in the face of rapid changes in AI inference requirements, to help Nebius allocate computing resources more quickly and reduce the idle time of expensive computing power devices.
Target AI inference efficiency bottlenecks to reduce GPU idle time
Inferize's core technology mainly addresses the “cold start” problem during AI model deployment. Cold start means that the AI model needs to complete the process of loading the model and initializing related resources before starting to process user requests. When a new inference instance is launched, the GPU may need to wait for the model to load before it can be officially put into operation, resulting in a short period of idle resources.
This problem is particularly prominent when the demand for AI inference suddenly increases. As user requests surge, cloud service providers need to quickly launch more inference instances, but if the model takes a long time to load, the newly added GPU resources cannot be immediately put into use, affecting service response speed and overall computing power utilization efficiency.
Inferize's technology aims to shorten this waiting process, enable new inference resources to be put into operation faster, and make computing power expansion closer to actual changes in demand.
For Nebius, this means that the company is expected to meet customer needs more flexibly and improve the efficiency of the use of existing infrastructure without having to maintain a large number of idle GPUs as backup resources for a long time.
Danila Shtan, Chief Technology Officer of Nebius, said that the efficient operation of AI inference services requires not only faster GPUs and optimized models, but the entire system must also be able to respond in a timely manner to changes in demand, including quickly providing additional computing power when customer demand increases. He pointed out that Inferize not only brings technology to accelerate this process, but also has an engineering team experienced in GPU systems.
Nebius plans to include Inferize engineers in its inference business team and integrate related technology into its AI inference platform Token Factory to improve the platform's ability to respond to customer needs and allow existing infrastructure to undertake more practical computational tasks.
At the same time, Shtan said that the Inferize team's future contributions will not be limited to this initial technology integration.
Reduce backup computing power costs and improve the efficiency of AI infrastructure operations
Guy Bortnikov, co-founder and CEO of Inferize, said that in order to respond to customer needs at any time, cloud service providers usually need to maintain a certain amount of backup GPU resources, and these standby computing power devices themselves represent additional costs. He pointed out that Inferize was founded to reduce this part of the cost, and adding Nebius can directly apply related technology to the AI cloud platform in actual operation.
In the AI inference business, the scheduling efficiency of computing power resources is directly related to infrastructure operating costs. Due to the high cost of GPU devices, if enterprises need to keep a large amount of idle computing power for a long time to cope with potential spikes in demand, it may affect the overall return on investment.
By shortening model startup time and speeding up the use of computing resources, Nebius is expected to reduce dependence on backup GPU capacity while improving service elasticity.
Notably, Bortnikov also previously co-founded Granulate, a computing infrastructure optimization software company. The company mainly develops technology to improve the operating efficiency of computing resources, and was acquired by Intel (INTC.US) in 2022. This background also reflects that the Inferize team has relevant experience in the field of computing infrastructure optimization.
Successive acquisitions of AI technology company Nebius to continue strengthening the inference platform
The acquisition of Inferize is one of a series of recent acquisitions carried out by Nebius in the field of AI inference and model optimization. Previously, Nebius had completed the acquisition of Eigen AI, with a transaction value of US$643 million. Additionally, the company recently acquired AI technology company Clarifai, but did not disclose the exact amount of the deal.
These acquisitions are all aimed at enhancing Nebius' AI inference and model optimization capabilities and further improving its AI infrastructure services.
Unlike infrastructure investment, which mainly focuses on the computing power required for AI model training, AI inference focuses more on operating efficiency after the model is put into actual use, including request processing speed, computational resource scheduling, model deployment, and unit computing power costs.
As Nebius continues to integrate related technologies, its business layout is also being further extended from providing GPU computing resources to technologies and services that improve the actual operation efficiency of AI models.
The acquisition of Inferize will further complement Nebius' technical capabilities in dynamic computing power scheduling and rapid model deployment. However, the extent to which the integration of related technologies can bring about cost savings and efficiency improvements has yet to be verified.