Red Hat introduces open stack for Generative AI

At its annual Red Hat Summit, Red Hat unveiled a series of key innovations within its artificial intelligence portfolio , aimed at accelerating enterprise adoption of generative AI.
Among the headline announcements are the launch of Red Hat AI Inference Server, new validated models for Red Hat AI, and the integration of the Llama Stack APIs and the Model Context Protocol (MCP).
Infrastructure for AI.
With these additions, the company reinforces its commitment to delivering flexible, high-performance solutions in hybrid cloud environments.
These innovations address the technical challenges of generative AI by providing tools that enable greater control, efficiency, and scalability for IT leaders, data scientists, and developers.
According to a Forrester study cited by the company, open-source software will play a pivotal role in advancing enterprise AI.
Built on this premise, Red Hat AI Inference Server positions itself as a robust solution to run AI models faster, more consistently, and cost-effectively.
This server is now available both as a standalone solution and integrated into the latest releases of OpenShift AI and Enterprise Linux AI (RHEL AI).
Meanwhile, third-party validated models for Red Hat AI, accessible via the Hugging Face platform, provide proven options tailored to specific use cases.
The company has optimized several of these models through compression techniques to enhance inference speed and reduce operational costs.
Ongoing validation processes ensure that enterprises stay at the forefront of generative AI developments.
Open Standards for Agents.
As part of this strategy, Red Hat is also integrating Llama Stack, originally developed by Meta, and Anthropic’s MCP, both offered as standardized APIs to build AI agents and applications.
Currently in developer preview, Llama Stack provides a unified API that includes features such as inference with vLLM, retrieval-augmented generation (RAG), model evaluation, and security tools.
MCP facilitates connecting models to external data sources and tools through standardized workflows.
OpenShift AI version 2.20 introduces enhanced capabilities for creating, training, and monitoring large-scale predictive and generative models.
Key improvements include:
– A catalog of optimized models (technology preview) that simplifies deployment and management of validated models directly from the web console.
– Distributed training via the Kubeflow Training Operator, leveraging GPUs and RDMA networks to reduce costs and accelerate workloads.
– A feature store based on Kubeflow Feast that centralizes data management for training and inference, improving model accuracy and reuse.
Multilingual and Cloud AI.
Concurrently, Enterprise Linux AI 1.5 brings new features such as:
– Availability on Google Cloud Marketplace, complementing AWS and Azure to expand deployment options on public cloud platforms.
– Enhanced multilingual capabilities optimized for Spanish, German, French, and Italian via InstructLab, enabling more granular model customization. The company plans to add Japanese, Hindi, and Korean in future releases.
Additionally, the Red Hat AI InstructLab service on IBM Cloud is now generally available.
This service facilitates model customization through an improved user experience, cloud scalability, and increased control over organizational data.
Joe Fernandes, Vice President and General Manager, AI Business Unit, Red Hat, states:
“Faster and more efficient inference is emerging as the critical decision point for innovation in generative AI . Red Hat AI, with optimized inference capabilities via Red Hat AI Inference Server and a new collection of validated third-party models , empowers enterprises to deploy intelligent applications where they need them, how they need them, and with components that best suit their specific requirements.”
The multinational company reaffirms its vision of a future without technological barriers, where organizations can deploy any AI model, on any accelerator, and on any cloud.
The goal: maximize performance and reduce costs.

