The launch of Cloudera AI Inference follows the agreement announced a few months ago with NVIDIA, reinforcing Cloudera's commitment to AI innovation for businesses at a time when all sectors are facing the challenges of digital transformation and the integration of this technology.

This allows developers to build, customize, and deploy enterprise-level large language models (LLMs) with up to 36 times faster performance using NVIDIA Tensor Core GPUs, and nearly 4 times faster performance compared to CPUs.

Because the user experience is integrated, it connects the graphical interface and APIs directly to NVIDIA's NIM microservices containers, eliminating the need for separate interfaces and monitoring systems. Integrating the service with Cloudera's AI Model Registry also enhances security and governance, enabling access controls for both model endpoints and operations. Users thus benefit from a unified platform where all models—whether LLM deployments or traditional models—are seamlessly managed under a single service.

“We are thrilled to partner with NVIDIA to launch Cloudera AI Inference, providing a single AI and ML platform that supports virtually every model and use case. This allows businesses to build powerful AI applications with our software and run them directly on our platform,” adds Dipto Chakravarty, Chief Product Officer at Cloudera.

“Today, businesses need to seamlessly integrate generative AI with their existing data infrastructure to achieve better business outcomes,” adds Kari Briski, vice president of AI software, models, and services at NVIDIA. “By incorporating NVIDIA NIM microservices into Cloudera’s AI Inference platform, we’re giving developers more tools to easily build high-quality generative AI applications.”.