Skip to content
Aug 14, 2026·8 min read

Mainframe cost per year in a defensible budget

Build a defensible mainframe cost per year from software capacity, hardware, labor, facilities, and the costs a migration can actually remove.

Mainframe cost per year in a defensible budget

A mainframe cost per year is not the number on a hardware invoice. It is a stack of contracts, capacity measurements, labor obligations, facilities charges, and risk provisions that behave differently when the workload changes. Put them into one blended rate and you will make a bad migration case: either the savings look miraculous or the machine looks impossible to leave.

I have seen cost exercises fail because finance asked for a single price per MIPS and engineering supplied one. MIPS can describe relative processor capacity, but most important software bills do not simply equal MIPS multiplied by a public rate. The useful model starts at each invoice and contract, identifies the driver for that line, and then asks whether moving one workload changes that driver at all.

There is no honest universal mainframe rate

An annual mainframe budget needs separate cost pools because each pool responds to a different event. Software may react to a measured four-hour peak. Hardware maintenance may stay fixed until a machine leaves service. Staff cost changes only when responsibilities and on-call coverage change. A data center allocation may remain on the books after the floor space is empty.

At minimum, separate these pools:

  • Capacity-priced software, including products whose charge follows MSUs or another measured capacity metric.
  • Fixed or tiered software, including annual licenses, support, and products priced by machine class or environment.
  • Hardware ownership, lease, maintenance, storage, network equipment, and attached devices.
  • People, external support, facilities, disaster recovery, security operations, and compliance work.

Do not allocate all of them to applications with the same denominator. CPU consumption can be reasonable for one software product and absurd for storage, third-party support, or a specialist who spends most of the week on release coordination. Choose a driver that explains why the cost exists: measured peak contribution, storage occupied, tickets handled, environments used, or dedicated effort. Leave a shared cost unallocated if any allocation would pretend to precision you do not have.

The first artifact should be a contract register, not a grand total. For every line item record the supplier, product, contract term, renewal date, billing metric, committed minimum, current tier, cancellation condition, and evidence source. Add an owner who can defend the interpretation. A spreadsheet cell called "mainframe software" cannot tell you which charge falls when a workload moves.

MIPS estimates capacity but rarely reproduces the invoice

MIPS means millions of instructions per second, but it is not a stable unit of business work and it is not IBM's general invoice currency. Instruction mixes differ. Processor generations do more useful work for a nominal capacity level. I/O-heavy batch, transaction processing, compression, Java, and database work stress the machine in different ways. A MIPS figure without its source and measurement method is a label, not evidence.

Organizations still use MIPS for planning because it gives teams a familiar scale. Vendors also use MIPS bands in some contracts. That makes the number relevant, just not universal. Ask whether each quoted MIPS value is a machine rating, an observed use level, a peak, an average, or a vendor's contractual band. Those figures cannot substitute for one another.

The dangerous shortcut looks like this:

annual mainframe cost = total MIPS x assumed price per MIPS

That formula hides minimum commitments, software tiers, specialty engines, development machines, storage, and labor. It also assumes the next MIPS costs the same as the first. Contracts often make that false. A marginal capacity increase can cross a tier; a modest decrease can save nothing until the committed floor or band changes.

Use MIPS as a reconciliation check instead. If an inventory claims that one application consumes half the estate while SMF-derived reports show a much smaller share of general-purpose processor use, investigate the boundary. The application estimate may include database and middleware work charged elsewhere, or the measurement may omit batch jobs submitted under another identifier. That argument is useful because it exposes ownership. Multiplying an uncertain MIPS total by an equally uncertain rate only gives uncertainty a currency symbol.

MSU billing follows measured capacity and contract rules

An MSU, or million service units, is a capacity measure used in IBM mainframe software pricing. It is closer to the billing machinery than MIPS, but an MSU is still not a price. The charge depends on the product, pricing model, eligible machine, contract, country, committed level, and reported use. Treat anyone offering one universal dollar-per-MSU number with suspicion.

For many sub-capacity arrangements, the operational number people watch is the rolling four-hour average, commonly called R4HA. Workload Manager records service consumption, and the Sub-Capacity Reporting Tool processes the relevant data into reports used for eligible software charging. IBM's SCRT documentation is explicit about collecting complete, valid input for the reporting period. That requirement matters: missing intervals are not proof that the workload was free. They are a reporting defect that can affect how the report is accepted or calculated.

The field routinely blurs three different values:

  • Instantaneous demand tells operators what is happening now.
  • A rolling average smooths that demand over the defined window.
  • The billed value applies product eligibility and contractual rules to reported capacity.

Confusing them creates imaginary savings. Removing a job that runs outside the peak window may cut CPU hours and electricity while leaving the software charge unchanged. Removing work inside the peak can still save nothing if another workload becomes the new peak, a minimum applies, or the product remains required on the same machine.

Build a peak-contribution view from the underlying intervals. For each product and logical partition, record the timestamp of the monthly high, the workloads active through that window, the reported MSUs, the contractual floor or tier, and the next lower economic boundary. Then test workload removal against the whole time series. Do not subtract an application's average MSUs from the reported peak. Peaks move.

Specialty engines also complicate the story. Eligible work on zIIP or other specialized capacity can change the economics, but it does not make the surrounding software, storage, operations, or fallback capacity disappear. Document which work is eligible, where it actually ran, and what happens during contention or failover. Capacity eligibility is an engineering property; the invoice result is a contract property. You need both.

The reporting chain deserves its own reconciliation. Map each central processor complex and LPAR to the machine identifiers in the report, then map each charged product to the places where it is licensed and used. Check that development, test, recovery, and production environments appear under the correct pricing treatment. An LPAR missing from the application inventory can still contribute to a product bill; an application listed on an LPAR may not use the product at all.

Keep the raw interval records for the months you model. A monthly maximum without the surrounding intervals cannot show what would happen if one job moved, ended earlier, or ran on a different day. For a candidate workload, replay the arithmetic across every interval after removing only the consumption you can attribute to it. Recalculate the rolling average and find the new high. This is still an estimate because product rules and contract terms sit above the capacity data, but it is far better than subtracting an annual average from one peak.

The replay also exposes load displacement. Suppose a close process owns Tuesday's high and a customer statement run is only slightly lower on Thursday. Removing the close process does not create savings equal to Tuesday's full contribution. Thursday becomes the capacity high, so only the gap between the old and new peaks might affect the measured level. If both peaks remain inside the same contractual band, the immediate invoice change can be zero.

Ask the software asset manager for the entitlement and invoice, not just the product list from system programmers. A product can be installed but uncharged under one arrangement, charged under a suite, subject to a minimum, or included in a wider agreement. Conversely, a component that looks incidental in the code inventory can carry its own support or usage charge. Procurement language decides what can terminate; technical discovery decides whether it is safe to terminate. Neither team can complete the row alone.

Development and test capacity needs separate attention. Teams often allocate all nonproduction cost in proportion to production CPU, although a release-heavy application may consume far more test time than its production share suggests. Record which environments are dedicated, which products must be licensed there, how test data is held, and whether disaster recovery uses active or standby rights. A migration may remove production demand yet leave a test image alive for defect history, tax queries, or a delayed downstream cutover.

Do not treat list prices as invoice data. List material can help identify the shape of a metric, but negotiated agreements, bundles, caps, and minimums determine actual cash. If contract access is restricted, keep the model under the same controls rather than replacing it with a public estimate. A redacted executive view can show cost categories and removal dates while the controlled schedule retains supplier and price detail.

Currency and accounting periods can also distort comparisons. Normalize invoice currency using finance's approved method, place prepaid support into the period it covers, and distinguish tax from supplier revenue if the business case treats tax differently. Reconcile credits and one-time adjustments instead of silently using a favorable month. A twelve-month view should explain every material difference between contracted annual value, invoiced cash, and the ledger.

Finally, assign confidence to the claim, not to the whole workbook. A signed termination clause may support high confidence in a license saving. A peak reduction inferred from incomplete workload tags deserves low confidence until measurement improves. A hardware exit that depends on three later migrations is conditional. Put the missing evidence and the person responsible beside each uncertain row. Decision makers can tolerate uncertainty when they can see its source and the work required to resolve it.

Hardware cost is larger than the purchase order

Hardware cost includes acquisition or lease payments, vendor maintenance, disk and tape systems, network gear, environmental capacity, spares, and the second site. Some organizations own the processor outright and pay rising maintenance. Others refresh under a financing arrangement. The accounting presentation differs, but the cash obligation still has a term and an exit condition.

A processor upgrade can also pull software into a higher capacity band even when the business workload barely changes. Conversely, moving workload away may not create cash savings while the current lease runs or while the smaller configuration would breach resilience requirements. Separate accounting depreciation from avoidable cash. Depreciation may remain after a migration decision; a maintenance renewal may be preventable.

Facilities deserve measured inputs, not folklore. Use metered power where available, contracted floor and cage charges, cooling allocation, remote-hands costs, and disaster-recovery expenses. Yet do not lead the business case with electricity. On many estates, software and scarce labor dominate the decision. Power matters most when leaving the platform lets you close or materially shrink a facility commitment.

Record the earliest removal date for every physical charge. If the machine supports twelve systems and you migrate one, none of the chassis, maintenance, or recovery-site cost may move. This is why application-level "savings" and estate-level cash savings often diverge.

The expensive specialist is usually carrying several jobs

Modernize beyond a source translation
CodeHero rebuilds the architecture in Go, Rust, TypeScript, and Postgres while preserving observed behavior.

Mainframe labor cannot be modeled by counting COBOL developers and multiplying salaries. The hard-to-replace people often cover production control, JCL, scheduler behavior, RACF administration, CICS or IMS operations, Db2 recovery, performance analysis, storage, release mechanics, incident diagnosis, and knowledge that never reached a runbook. One name on an organization chart may hide five operational roles.

Split labor into capability coverage. For each capability, record the primary person, backup, weekly planned effort, on-call duty, external dependency, and what system still requires it. This reveals concentration risk without inventing a financial number for "tribal knowledge." It also stops a common mistake: booking a full salary as a saving even though the employee will run the replacement, support another mainframe workload, or remain through decommissioning.

Contractors and vendor retainers need the same treatment. A retainer may be cancellable only at renewal. A specialist may cover several applications. An outsourcer may charge for a minimum team or service tower. Read the statement of work and notice period before calling any amount variable.

I argue against using replacement hiring cost as the mainframe's annual labor cost. The approach is popular because retirement risk is real and recruiters can supply a large number. It is wrong because a hypothetical emergency hire is not the current run rate. Keep two columns: recurring labor cash and quantified transition or resilience exposure. A board can make a decision with both; it cannot audit a blended fear premium.

Do not assume migration immediately removes these people. During parity testing and cutover, their knowledge becomes more important. Afterward, the best outcome may be to retain them as domain owners while removing pager duty tied to obsolete infrastructure. The saving can appear in avoided contractor renewals, reduced on-call coverage, or vacancies you no longer need to refill, rather than layoffs.

A workload bill needs evidence at every row

A defensible workload bill traces each allocated amount to a source and labels its behavior. Start with twelve months of invoices and usage reports so a year-end batch, seasonal peak, or annual support fee does not vanish from view. Reconcile the total back to the general ledger before assigning a cent to applications.

Use a table with this shape:

cost_id,annual_cash,billing_driver,contract_floor,renewal_date,workload_share,removal_trigger,evidence
SW001,REDACTED,product_peak_msu,REDACTED,YYYY-MM-DD,measured,lower_tier_at_renewal,SCRT_report
HW004,REDACTED,fixed_lease,full_term,YYYY-MM-DD,shared,lease_end,signed_contract
LAB007,REDACTED,dedicated_effort,none,YYYY-MM-DD,time_study,role_reassigned,staffing_plan

Keep amounts redacted in working examples, but require real amounts in the controlled model. The useful column is removal_trigger. It forces the analyst to name the event that changes cash: a lower software tier accepted at renewal, a license terminated, a lease ended, a machine decommissioned, a support tower resized, or a vacancy left unfilled. "Application migrated" is rarely enough.

Then classify every row into one of four behaviors: avoidable with this workload, avoidable only when a group of workloads leaves, fixed until a dated event, or retained after migration. Run the model twice. The workload view shows economic consumption; the cash view shows invoices and payroll that actually change. Both are legitimate, but only the second funds a business case.

A simple calculation makes the distinction visible:

run-rate allocation = annual cash x workload share
year-1 cash saving = annual cash x removable share x active fraction of year
net year-1 effect = year-1 cash saving - migration cash cost - overlap cost
steady-state saving = terminated and resized annual obligations

Do not put estimated risk reduction into the cash-saving line. Track reduced outage exposure, audit effort, recovery complexity, and hiring concentration as decision factors with an owner and evidence. If you later monetize them, show the probability and impact assumptions separately.

Before approval, make finance, platform operations, procurement, and the application owner sign the rows they understand. Procurement catches renewal traps. Operations catches shared dependencies. The application owner catches missing batch and interfaces. Finance prevents an allocation from masquerading as removed spend.

Migration removes contracts only after shared dependencies leave

Handle the million-line estate
The agentic platform reads systems over a million lines without splitting away cross-language dependencies.

A migration can remove application licenses, capacity demand, storage growth, batch windows, specialist coverage, and hardware commitments. It removes each one on a different date. The safe forecast is a dependency ladder, not one percentage applied to the total estate.

First, map the workload boundary. Include online transactions, scheduled jobs, file transfers, print, database procedures, security rules, operational scripts, reconciliation, and downstream consumers. A web front end moved to a new service while the system of record remains in Db2 has not removed the database or its operational burden. A batch rewrite that still submits JCL for three closing jobs has not removed scheduler coverage.

Second, identify the stranding threshold for each shared cost. A software product may disappear only when its last dependent workload leaves an LPAR. Tape operations may remain for regulatory retention. The recovery machine may be sized for the largest remaining service. Network circuits may support other systems. This is where a portfolio sequence can create more value than choosing applications solely by apparent cost.

Third, put termination work into the plan. Archive data under an approved retention rule, remove identities and scheduler entries, stop feeds, prove recovery for the replacement, update operating procedures, submit contract notices, and dispose of hardware correctly. A workload that no longer receives traffic can still cost money and fail an audit.

CodeHero rewrites the whole legacy codebase into Go, Rust, and TypeScript and verifies behavior with a parity harness against recorded production traffic. That can compress the delivery work to under 30 days, but your contract notices, retention duties, and shared-platform exit dates still control when the annual cost falls.

Several costs survive on the new platform

Reach the dated exit sooner
Every rewrite project is delivered in under 30 days, leaving contract dates visible as separate gates.

Migration changes the cost structure; it does not abolish production engineering. The replacement needs compute, databases, storage, observability, backups, security controls, support, incident response, disaster recovery, and people who understand the business rules. A model that sets these to zero is advocacy, not analysis.

Cloud bills need the same discipline as mainframe bills. Separate baseline capacity, burst demand, managed database charges, network transfer, backup retention, nonproduction environments, and support. Do not compare a fully burdened mainframe figure with a bare virtual-machine estimate. Include parallel running during verification and the temporary storage used for data movement.

Some duties become cheaper because common labor markets and tooling replace specialized platform knowledge. Others merely move. RACF policy may become identity and access policy. SMF monitoring may become logs, metrics, and traces. Db2 backup and recovery may become Postgres backup and recovery. Name the receiving owner for each control before removing the old owner.

Performance headroom also survives. Mainframes handle mixed transaction and batch loads with mature workload controls. A replacement must meet observed latency, throughput, close-window, and recovery behavior under real demand. Size it from production traces and test results rather than source-line counts. Architecture modernization can reduce waste, but the budget should recognize the capacity you have proved.

Define the baseline before requesting migration proposals. Freeze the estate boundary, the twelve-month cost period, and the treatment of shared services. List any planned processor refresh, contract renewal, data-center move, or staffing change that would happen without migration. Otherwise the project receives credit for savings already scheduled, or gets blamed for a cost increase that the current platform would also incur.

Present at least three time views. The committed view includes obligations that cannot yet change. The exit-year view reflects partial-year terminations and parallel running. The steady-state view begins only after old production, recovery, retention, and support duties have ended. Label the starting date of each view. An annual figure without a calendar quietly assumes that every saving begins on day one.

Use gates for the exit rather than a hopeful cutover date. Traffic cutover is one gate. Financial exit may also require a completed production cycle, reconciled outputs, accepted recovery test, expired rollback period, archived records, removed access, signed service acceptance, and acknowledged contract notice. Give every gate an owner and evidence. If one gate slips, the cost model should move the connected saving automatically.

The parity period needs a budget of its own. Both platforms may process the same recorded or live business cases while teams compare outputs, investigate differences, and prove operational behavior. Include duplicate compute and storage, data extracts, test execution, defect work, and the people who approve equivalence. Do not hide this temporary spend inside a contingency percentage; tie it to the verification plan and expected cycles.

Set a rule for residual exceptions. A replacement that handles nearly all traffic but sends rare cases back to the mainframe has preserved the dependency that blocks decommissioning. Count the frequency, identify the business rule, and decide whether to implement it, retire it with approval, or operate a bounded manual process. Leaving an undefined fallback path open can keep an entire license and support chain alive.

After cutover, compare invoices with the forecast for several billing cycles. Confirm that reported peaks, product tiers, support quantities, storage, circuits, and contractor charges changed when expected. Close purchase orders and remove automatic renewals rather than assuming an unused service will stop billing itself. Record forecast variance by cost row so the next migration uses observed removal behavior instead of another generic savings percentage.

The final approval should state who owns stranded cost. If one application moves but a shared product remains, the saving cannot appear in that application's case and then again in a later estate case. Keep a central ledger of claimed, realized, and still-stranded amounts. This prevents double counting and gives portfolio planners a rational reason to group the next workloads around the dependencies that are expensive to leave behind.

The decision rests on removable cash and a dated exit

The investment case should present three totals: today's annual cash run rate, the steady-state replacement run rate, and the transition cash required to move between them. Next to those totals, show a dated schedule of obligations that terminate. This makes delayed savings visible and prevents year-one overlap from being hidden in an annualized number.

Stress the assumptions that can reverse the decision. Test a software peak that does not fall, a lease that cannot end early, a product that remains for another application, higher replacement capacity, longer data retention, and an extra period of parallel operation. If the case works only when every shared cost vanishes on cutover day, it does not work.

The go or no-go decision is not purely financial. An estate can justify migration because change lead time, recovery risk, or talent concentration blocks the business, even when the narrow first-year cash saving is modest. Say that openly. Do not bury strategic reasons inside invented cost avoidance.

Start the approval pack with the contract register, SCRT-backed peak view, capability map, workload boundary, and removal-trigger table. Those five artifacts give technical and finance leaders something they can challenge. Once every claimed saving has an owner, a trigger, and a date, the annual number stops being a debate about MIPS folklore and becomes a plan the company can execute.

FAQ

How much does a mainframe cost per year?

There is no credible universal figure. Add the actual software contracts, hardware or lease payments, maintenance, storage, facilities, recovery, support, and labor, then separate allocated consumption from cash that can truly be removed.

Can I calculate mainframe cost from MIPS alone?

No. MIPS can support capacity comparisons, but it does not capture contract floors, product-specific pricing, storage, hardware terms, or staffing. Use it as a reconciliation signal, not as the whole bill.

What is the difference between MIPS and MSUs?

MIPS estimates instruction-processing capacity, while MSUs are service-unit capacity measures used in mainframe management and some software pricing. Neither has a universal monetary rate, and contract terms decide how measured capacity becomes a charge.

What is the rolling four-hour average?

The R4HA smooths service consumption over a four-hour window and often matters in eligible sub-capacity software reporting. Removing CPU work does not guarantee it falls because the peak can move to another window or workload.

Does moving one application reduce software licensing immediately?

Often it does not. The product may remain installed for other workloads, the reported peak may stay above the same tier, or a committed minimum may apply until renewal.

Which mainframe costs disappear after migration?

Only obligations whose removal triggers have been met disappear: terminated licenses, resized capacity tiers, ended leases, retired maintenance, reduced support, or roles genuinely reassigned or left unfilled. Shared costs stay until their last dependency leaves.

Should mainframe specialists count as migration savings?

Count only a staffing cash change that management actually plans and can date. Specialists often carry business knowledge into parity testing and the replacement, so a whole salary rarely disappears at cutover.

How should we allocate shared mainframe costs?

Use the driver that causes each cost, such as peak contribution, storage occupied, environments used, or dedicated effort. Keep the workload allocation separate from the cash-saving forecast so shared overhead does not look removable.

What replacement costs are usually missed?

Teams often omit nonproduction environments, observability, backup retention, network transfer, incident coverage, disaster recovery, parallel running, and data-migration storage. The new platform still needs production engineering.

What evidence should a mainframe migration business case include?

Include a contract register, twelve months of invoices, SCRT-backed capacity reports, a capability map, a complete workload boundary, and a dated removal trigger for every saving. Finance and engineering should be able to trace every total to those records.