Open source and hybrid cloud: the hidden engine of new AI

The combination of open source and hybrid cloud is revolutionizing artificial intelligence, delivering flexibility and scalability for modern enterprises.
The combination of open source and hybrid cloud is revolutionizing artificial intelligence, delivering flexibility and scalability for modern enterprises.

 


By: Chris Wright, Chief Technology Officer and Senior Vice President, Global Engineer, Red Hat.

“Any workload, any application, anywhere” was the mantra of Red Hat Summit 2023. It is true that, over the past two years, we have seen some changes in IT. But Red Hat’s vision has not changed; it has evolved.

This is the message of hybrid cloud for the AI era. And the best part is that, just like the “old” hybrid cloud, open source innovation drives everything.

At Red Hat Summit, we demonstrated how AI ecosystems structured around open source and open models can create new options for businesses. Openness brings the possibility of choice, and this freedom leads to greater flexibility: from the model that best meets a company’s needs to the underlying accelerator and the location where the workload will actually run.

Successful AI strategies will follow the data, wherever it resides within the hybrid cloud.

The Role of Open Source.

So, what drives hybrid cloud? Open source.

In my view, we need to start looking beyond models. Yes, models are very important for AI strategies. But without inference — the “practical” aspect of AI — models are merely datasets that do not “do” anything.

Inference refers to how quickly a model responds to user input and how efficiently decisions can be made on accelerated computing resources.

Ultimately, slow responses and inefficiency cost money and erode customer trust.

Innovations in AI Inference.

That is why I am very excited that Red Hat prioritizes inference in our work with open source AI, beginning with the launch of Red Hat AI Inference Server.

This solution, built on the leading open source project vLLM and optimized with Neural Magic technologies, offers AI deployments an inference server with full support, a complete lifecycle, and production readiness.

Best of all, it can track your data wherever it resides, as the solution works on any Linux platform, any Kubernetes distribution, whether Red Hat or otherwise.

The flagship enterprise IT application is not a single unified workload or a new cloud service: it is the ability to scale quickly and efficiently.

This also applies to AI. However, AI has a unique challenge: the accelerated computing resources that underpin AI workloads must also scale.

This is no easy task, given the costs and skills required to properly deploy such hardware.

Open Source Projects for AI’s Future.

What we need is not only the ability to scale AI but also to distribute massive AI workloads across multiple accelerated computing clusters.

Chris Wright, Chief Technology Officer and Senior Vice President, Global Engineer, Red Hat.
Chris Wright, Chief Technology Officer and Senior Vice President, Global Engineer, Red Hat.

This challenge is compounded by the scaling inference time required by reasoning models and agentic AI.

By sharing the load, performance issues can be reduced, efficiency improved, and ultimately the user experience enhanced.

With the open source project llm-d, Red Hat has taken steps to address this problem.

The llm-d project, led by Red Hat and supported by leaders in AI hardware acceleration, model development, and cloud computing, combines the proven power of Kubernetes orchestration with vLLM, uniting two open source leaders to solve a very real need.

Along with technologies like AI-driven network routing and KV cache offloading, among others, llm-d decentralizes and democratizes AI inference, helping businesses optimize their computing resources and manage more effective, cost-efficient AI workloads.

Llm-d and vLLM, included in Red Hat AI Inference Server, are open source technologies ready to meet today’s enterprise AI challenges.

However, development communities are not limited to addressing only current needs.

AI technologies have a particular way of shortening timelines: the vertigo of innovation means something once thought to be a challenge years away suddenly demands immediate attention.

Advances in Generative AI and Agents.

That is why Red Hat is allocating resources to the early development phase of Llama Stack, the Meta-led project offering core components and standardized APIs for the lifecycle of generative AI applications.

Moreover, Llama Stack is ideal for creating agentic AI applications, representing a new evolution of the powerful generative AI workloads we see today.

Beyond its initial development, Llama Stack is available as a developer preview within Red Hat AI for companies eager to engage with the future today.

Regarding AI agents, we still lack a common protocol for how other applications provide them with context and information.

This is where the Model Context Protocol (MCP) comes into play.

Developed and open sourced by Anthropic at the end of 2024, it is a standardized protocol for agent-application interactions, akin to client-server protocols in traditional computing.

Most importantly, current applications can suddenly leverage AI without requiring large-scale reimplementation.

This is crucial and would not be possible without the power of open source.

Like Llama Stack, MCP is available as a developer preview on the Red Hat AI platform.

Proprietary AI models may have gained an early advantage, but there is no doubt that open ecosystems have surpassed them, especially regarding the software that underpins next-generation AI models.

Thanks to vLLM and llm-d, together with enterprise-grade open source products fortified with security, the future of AI is promising—regardless of model, accelerator, or cloud—and it is driven by open source and Red Hat.

Share:
Hosting Web
Most Read