Db2 for z/OS and its ecosystem

Db2 for z/OS and its ecosystem

Connect with Db2, Informix, Netezza, open source, and other data experts to gain value from your data, share insights, and solve problems.

 View Only

Simplified and enhanced: Introducing the rebranded AI Optimizer for IBM Z and IBM LinuxONE

By Guanjun Cai posted 05/05/26 02:33 PM

  

By Guanjun Cai, Db2 for z/OS and Ecosystem Products Content, IBM Software

AI Optimizer for IBM Z and IBM LinuxONE (AI Optimizer for Z and LinuxONE) is the rebranded evolution of IBM AI Optimizer for Z. This release transforms the way you deploy large language models (LLMs) and route inference requests on mainframe platforms. With a streamlined architecture, intelligent tag‑based routing, and enhanced appliance management, it delivers more secure, efficient, and powerful LLM deployment and inference routing than ever before.

What is AI Optimizer for Z and LinuxONE?

In the simplest term, AI Optimizer for Z and LinuxONE is an enterprise-grade appliance that serves as your central gateway for deploying LLMs and routing inference requests on IBM Z and IBM LinuxONE systems. Think of it as a unified control plane that makes AI inference as straightforward as any other enterprise workload — but with all the security, reliability, and performance you expect from mainframe platforms.

Built for the IBM Secure Service Container (SSC) platform, the appliance is specifically optimized to work with IBM Spyre accelerator cards. This means you get purpose-built AI inference capabilities while keeping your data exactly where it belongs - the secure, compliant platform where your mission-critical business logic already runs.

What's new in the rebranded release?

This release represents a significant evolution. The biggest change is the architectural shift from Red Hat OpenShift to SSC. This isn't just a technical change — it's a fundamental improvement that provides a hardened, purpose-built infrastructure specifically designed for secure AI deployments and inference workloads on IBM Z and IBM LinuxONE.

Here's what this architectural transformation means in terms of new and enhanced functionality for you:

  • Simplified setup and management: The new appliance-based architecture dramatically simplifies setup. Following a straightforward installation of the containerized appliance, you can complete a one-time configuration — loading, provisioning, and enabling all required components — with a single click. No more complex multi-step deployments or ongoing infrastructure management overheads.
  • Direct user management: This release introduces the ability to create and manage users directly within the appliance through the new Appliance Manager UI. You no longer need to rely on external identity providers. With administrative privileges, you can create user accounts, set passwords, and manage profiles all in one place.
  • Enhanced role-based access control (RBAC): The new RBAC system lets you assign fine-grained permissions that control exactly what each user can access and do within the appliance. This ensures security boundaries are maintained while giving users the access they need to be productive.
  • Native IBM Granite support: The appliance now includes native support of IBM Granite 3.3-8B-Instruct, an LLM specifically optimized for Spyre accelerators. You can deploy this model directly through the UI without manual downloads or complex configuration steps.
  • Remote model gateway (Tech preview): The new remote model gateway lets you register, tag, and manage remote models alongside your locally deployed ones. You get unified management and routing across your entire AI infrastructure, with performance metrics included in your monitoring dashboard where supported.
  • Built-in troubleshooting tools: The Appliance Manager now provides direct access to logs and system dumps for monitoring and troubleshooting. When issues arise, you have the diagnostic information you need right at your fingertips.
  • Enhanced security features: This release includes comprehensive security improvements, including external server certificate management for secure communications and SSL/TLS implementation for data encryption across all aspects of the appliance.

How does AI Optimizer for Z and LinuxONE work?

The beauty of AI Optimizer for Z and LinuxONE lies in its simplicity. At its core is the inference router that acts as a single point of entry for all your AI inference requests. The router exposes API endpoints built on the OpenAI API specification, which has emerged as the de facto industry standard for interacting with LLMs. This means that if you're already using OpenAI tools and libraries, you can integrate with the appliance with minimal code changes.

Here's how it flows: You deploy and tag LLMs based on your use cases in the appliance's web UI and then send inference requests by using the REST API. The router intelligently examines each request and uses tag-based routing to direct it to the most appropriate model. Maybe you've tagged certain models for financial data, others for customer service, and still others for different languages or performance requirements. The router handles all that complexity behind the scenes, automatically selecting the best model for each request.

What makes this particularly powerful is the integration with Spyre accelerator cards. When you deploy models on Spyre hardware, the inference engine automatically leverages hardware acceleration to maximize performance. You get the speed you need without sacrificing the data gravity and security that mainframe platforms provide.

Equally powerful is the native support of Granite LLMs. The appliance comes preinstalled with Granite 3.3-8B-instruct, a model designed for advanced instruction-following, reasoning, and coding and specifically optimized for the Spyre accelerators. You can easily deploy the model without manual downloads or complex configurations.

Throughout all of this, a comprehensive observability stack built on Prometheus and Grafana collects metrics on everything from inference latency to resource utilization. You get real-time visibility into how your AI inference routing workloads are performing, making it easy to identify bottlenecks and optimize your deployment strategy.

Ready to get started?

AI Optimizer for Z and LinuxONE represents a significant step forward in making enterprise LLM deployment and inference routing simpler on mainframe platforms. The new appliance-based architecture simplifies deployment and management, while enhanced features, such as intelligent routing, native Granite support, tight integration with Spyre accelerator, and the remote model gateway, provide the flexibility you need to build sophisticated AI solutions.

Ready to explore what AI Optimizer for Z and LinuxONE can do for you? Visit IBM Documentation and IBM Products to learn more about how to get started.

A special thank-you to Alexander Merschel, AI Optimizer Technical Lead, IBM Software, for his contributions.

0 comments
25 views

Permalink