AI mainframe cost is now the fastest-growing line item in most IBM Z budgets. AI workloads on z/OS have moved from pilots into production: IBM’s CEO has confirmed that customers running IBM watsonx Code Assistant for Z are scaling MIPS capacity three times faster than those who have not.
If you are a mainframe leader, the important question is how AI will change your costs under the current IBM pricing model, and whether you need to act now. Let’s talk about it.
AI on the mainframe is already here
Over the past year, the conversation has gone from whether or not enterprises should deploy generative AI in z/OS environments to what is the best way to manage the capacity and cost of mainframe AI workloads.
In its September 2025 mainframe survey, BMC reported that 65% of mainframe organizations are already using GenAI in their z/OS environment for code analysis, documentation, fraud detection, customer-experience workloads, or operational automation. Essentially, the majority of the market has already adopted GenAI.
Source: BMC Mainframe Survey 2025
The question now is: how should mainframe managers plan and budget for cost and capacity as AI adoption expands?
3x faster MIPS capacity growth trajectory
The most consequential statement on AI and mainframe in 2026 came from IBM CEO Arvind Krishna in IBM’s first-quarter earnings call on April 22, 2026:
“Clients who have deployed watsonx Code Assistant for Z are growing MIPS capacity three times faster than those who have not.”
— Arvind Krishna, IBM CEO, 1Q26 earnings call
CFO Jim Kavanaugh repeated the same figure later in the same call, framing it as a “3x differential on growth and capacity.”
This does not mean AI workloads are consuming more MIPS in absolute terms, but rather that the growth trajectory is three times faster than the rest of the IBM Z installed base.
That growth rate translates into cost differently depending on which IBM pricing model you are on. For AWLC, where software cost tracks the monthly billing peak, a 3x faster capacity growth leads to a monthly bill that keeps growing every month. For TFP, the above-baseline consumption that accumulates across the twelve-month contract period is reconciled at year-end. It also translates into a higher TFP baseline for the next year’s contract, which locks you into a higher cost floor for the following term.
This cost increase is already visible in the IBM numbers: IBM has now reported four straight quarters of more than 100% growth in new MIPS shipped on the z17 platform — the strongest hardware momentum on the mainframe in years. The growth is explicitly tied by IBM management to AI workload adoption.
Source: IBM 1Q26 Earnings Call Transcript
A new third type of compute is coming to z/OS, and it is structural
In the same earnings call, Arvind Krishna pointed to a structural shift:
“AI is adding a third kind of compute capacity into the mainframe.”
— Arvind Krishna, IBM CEO
The mainframe historically has had two types of compute capacity: classic MIPS (transactional workloads) and Linux MIPS (sparser Linux workloads). AI is now adding what he called a “third kind” of compute capacity.
The example he gave is immediately recognizable to anyone running mainframe workloads in financial services or insurance: today, credit card fraud detection runs a few rules against a sampling of transactions. With sufficient on-platform inference capacity, customers can run a 20–30 billion parameter model against every single transaction in milliseconds. IBM’s current statement is that a fully populated z17 system can process approximately 450 billion inferences per day on the mainframe itself.
What mainframe managers need to realize is that this is not a one-time uplift. AI workloads do not displace existing mainframe consumption, but rather add to it. IBM is explicitly positioning the mainframe as an AI inference destination, not a place customers should migrate away from. And that positioning has direct consequences for cost trajectory.
Don’t expect AI to be the migration escape hatch
A reasonable question at this point is: if AI is going to make the mainframe more expensive, can AI also make migration to the cloud finally tractable?
Many vendors and consultancies have spent 2025 and 2026 pitching that GenAI-powered code conversion, automated documentation, and compressed timelines have changed the economics of mainframe exit. In April 2026, Gartner published a research note titled “Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI.” Its central forecast is that more than 70% of mainframe exit projects initiated in 2026 will fail to produce their intended benefits because they overestimate GenAI capabilities. Gartner identifies three reasons for this:
- The marketing-to-reality gap. GenAI vendor claims about automated code conversion outpace delivered capability, particularly for the non-standard languages (Assembler, Easytrieve, CA Ideal, PL/1) and undocumented business logic that sit at the periphery of most mainframe applications.
- Performance and throughput equivalence is not guaranteed. AI-converted code can pass functional testing and still fail to match the performance, throughput, and resilience guarantees of the mainframe equivalent. The differences typically only emerge under production load.
- Data complexity defeats wholesale conversion. For most large enterprises, the volume and interdependency of mainframe transaction data makes a single-pass automated migration physically and financially impractical.
Gartner further forecasts that by 2030, 75% of vendors operating in the “mainframe exit” market will pivot their business models or cease to exist. That is a strong indicator that the current AI-led migration boom is unlikely to translate into a stable industry capable of supporting customers through multi-year programs.
Stepping back, two things become apparent: AI is making mainframe workloads more expensive. At the same time, it is not making mainframe exit cheaper or safer. The practical thing to do is to control the cost curve where the workloads actually live.
Source: AI-powered Mainframe Exits Are a Bubble Set to Pop — The Register
Where the controls are
AI is pushing mainframe consumption upward, but the cost mechanics — and therefore the right control — differ significantly between AWLC and TFP. Each pricing model needs its own approach.
For AWLC customers — Zetaly Automated Capacity (ZAC)
Under AWLC, MLC is calculated on the peak Rolling 4-Hour Average (R4HA) of MSU consumption each month, which means a single high-demand period sets the cost for the entire month. Accelerating AI workloads capacity translates directly into rising R4HA peaks, and that is how AWLC AI cost reaches your monthly bill. ZAC manages that billing peak in real time. It continuously monitors MSU consumption across LPARs and adjusts defined capacity limits and soft capping settings based on workload priority and current billing exposure in order to prevent R4HA peaks. This adjustment is based on workload priority: lower-priority work is adjusted first, while mission-critical transactions and SLA-sensitive workloads are never impacted. Across 60+ enterprise deployments, ZAC customers achieve a 5–20% reduction in MLC costs without any change to their applications. For fast-growing AI implementation in an AWLC environment, ZAC keeps the costs under control.
For TFP customers — Zetaly Data Platform (ZDP) and ZSI
Under TFP, the cost mechanics are different, and the AI costs compound in two distinct ways. Your TFP baseline is calculated from the last twelve months of SCRT data; any consumption above that baseline accumulates across the twelve-month contract period and is reconciled at year-end. That is the first component. The second one is structural: the next year’s contract baseline is calculated from the current year’s actual MSU consumption, meaning the AI-driven overage becomes the cost floor for the following term. AI MSU growth is reflected in the bill twice: as a one-time year-end bill and a permanent baseline reset.
The practical problem for any mainframe leader on TFP is that without workload-level visibility, you cannot see which workloads are driving your above-baseline consumption — AI versus business-as-usual. And you cannot manage what you cannot see. IBM’s native monitoring tools provide some insights, but they consume meaningful MSU themselves, which contributes to the very baseline you are trying to manage.
ZDP resolves the visibility problem at the data layer. It collects mainframe operational data, such as SMF records, system logs, DCOLLECT, and IMS logs, continuously and in near real time, with under 0.2% mainframe resource overhead. ZSI resolves it at the readability layer: it generates dashboards and reports on top of ZDP data, including reports organized by métier, so non-specialist audiences inside the business (including finance, line-of-business owners, or even the COMEX) can read and act on mainframe consumption without needing a mainframe specialist to translate. For TFP customers absorbing growing AI workloads, ZDP and ZSI together are the difference between discovering the year-end overage retroactively and managing the trajectory toward it as it accumulates.
What you can do to start reducing costs
What “controlling the AI cost curve” actually looks like depends on which IBM pricing model you are under. If your environment is on AWLC:
- Model the AWLC cost trajectory. Take your current AWLC consumption baseline and apply the 3x capacity growth curve IBM is now describing for Watsonx-engaged customers (the capacity side of watsonx Code Assistant cost). The resulting eighteen-to-twenty-four-month cost figure is the conversation you need to have with finance before AI workloads ramp up.
- Get a ZAC simulation. Zetaly offers a free simulation that models your optimization potential against your actual baseline — your data, not a vendor estimate. It is the best way to see how you could control AI cost growth in your own environment.
If your environment is on TFP or is being moved toward TFP under IBM pressure:
- Get workload-level visibility of TFP AI consumption early. Under TFP, AI workloads tend to add up to above-baseline consumption without notice. You need to know which workloads are driving the overage, so that you can reduce the year-end bill and avoid locking in an inflated baseline for next year.
- Decide your observability posture deliberately. Native IBM monitoring tools have material overhead that goes directly into your TFP baseline calculation. ZDP collects the same operational data (SMF, DCOLLECT, IMS logs) at under 0.2% overhead, so monitoring does not inflate the very baseline you are trying to manage. ZSI puts the resulting data in front of finance and line-of-business owners in the form they actually need.
AI is becoming a powerful mainframe tool that allows you to do more and faster than it ever has before. The AI conversation that mainframe leaders need to be having in 2026 is about cost trajectory and operational visibility that will help them make the most out of this transition while keeping the costs in check.
Frequently Asked Questions
What is AI mainframe cost, and how is AI affecting IBM mainframe costs?
AI workloads are driving up mainframe consumption faster than any other workload category in 2026. According to BMC’s 2025 Mainframe Survey, 65% of mainframe organizations are already using generative AI in their z/OS environment. IBM’s CEO confirmed in the company’s Q1 2026 earnings call that customers running Watsonx Code Assistant for Z are scaling MIPS capacity three times faster than mainframe customers without it. The cost impact differs by IBM pricing model: under AWLC, faster MIPS growth pushes the monthly Rolling 4-Hour Average billing peak up directly. Under TFP, AI workloads add MSU consumption above the contract baseline that accumulates across the 12-month period and resets next year’s baseline higher.
Why is mainframe MIPS capacity growing three times faster for AI customers?
On April 22 2026, IBM CEO Arvind Krishna stated that “clients who have deployed Watsonx Code Assistant for Z are growing MIPS capacity three times faster than those who have not.” The driver is structural: AI workloads represent a third type of compute on the mainframe, alongside classic transactional MIPS and Linux MIPS, which accumulates on top of existing workloads. A typical AI-driven shift involves running 20–30 billion parameter models inline against every transaction — for example, full-population fraud detection rather than statistical sampling. These are workloads that did not exist on the mainframe twelve months ago.
Will AI-powered tooling help me migrate off the mainframe?
Probably not. In April 2026, Gartner published a research note titled “Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI” forecasting that more than 70% of AI-led mainframe exit projects initiated in 2026 will fail to produce their intended benefits. Three structural reasons: GenAI code conversion tools have not solved the “exotic component” problem (Assembler, Easytrieve, CA Ideal, PL/1, undocumented business logic); performance and throughput equivalence between converted code and the mainframe original is not guaranteed and surfaces only under production load; and the volume and interdependency of mainframe transaction data makes wholesale automated migration impractical. Gartner additionally forecasts that 75% of vendors operating in the “mainframe exit” market will pivot or cease to exist by 2030.
How does AI cost impact differ between AWLC and TFP customers?
Under AWLC, software cost is calculated on the peak Rolling 4-Hour Average (R4HA) of MSU consumption each month. Faster AI-driven capacity growth pushes that monthly billing peak up directly, every month. Under TFP, the cost mechanics are different: any consumption above the contract baseline accumulates across the 12-month period and is reconciled at year-end, then rolls forward into the following year’s baseline calculation. The same AI workload growth therefore produces a recurring monthly cost increase under AWLC and a year-end reconciliation plus permanent baseline reset under TFP. While the mechanisms are different, the magnitude of impact is similar.
How do I control AI-driven mainframe cost growth?
The right control depends on the IBM pricing model. AWLC customers can reduce the R4HA billing peak using automated capacity management tools. Zetaly Automated Capacity (ZAC) achieves 5–20% MLC cost reduction across 30+ enterprise deployments by managing LPAR defined capacity in real time. TFP customers need workload-level visibility to identify which workloads are driving above-baseline consumption and act on it before year-end reconciliation. Zetaly Data Platform (ZDP) provides this kind of visibility for TFP at under 0.2% mainframe resource overhead, with Zetaly Service Intelligence (ZSI) generating dashboards organized by line of business. Visibility and active management are essential for both pricing models.
What is “AI MIPS” on the IBM mainframe?
“AI MIPS” is the term used by IBM CEO Arvind Krishna in IBM’s Q1 2026 earnings call to describe a third category of mainframe compute capacity, alongside classic transactional MIPS and Linux MIPS. AI MIPS represents inference workloads running directly on the mainframe — for example, real-time fraud detection running large language models against every transaction. IBM has stated that a fully populated z17 system can process approximately 450 billion inferences per day. The term reflects IBM’s strategic positioning of the mainframe as an AI inference platform, not a system to migrate AI workloads off of.
Sources & further reading
BMC Mainframe Survey 2025 — industry survey on mainframe trends and GenAI adoption (September 2025).
IBM 1Q26 Earnings Call Transcript — CEO Arvind Krishna and CFO Jim Kavanaugh, April 22, 2026. Source for the 3x MIPS capacity growth statistic and the “third kind of compute” framing.
Gartner: AI-Powered Mainframe Exits Are a Bubble Set to Pop — The Register — summarising Gartner’s April 2026 report “Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI.” Primary report is paywalled.
Zetaly: Mainframe to Cloud Migration — Pros, Cons, and What Actually Happens — companion guide covering migration strategies, real risks, and the role of mainframe FinOps.

