PdM Platform The Toolbox Services Academy Library About Contact Open PdM →Open the Toolbox →
Operate Track · Tool 02 Guide

PM Optimisation: reviewing a programme against what actually happened

A maintenance programme is written before the equipment runs an hour, and then runs for years while the plant changes around it. This guide covers the loop that closes that gap under NORSOK Z-008:2024 — why §11.4 and §11.5 are different lists, what three CMMS exports will and will not tell you, how an interval extension is earned, and where the honest answer is that the data cannot say.

NORSOK Z-008:2024 Clause 9 · Clause 11 ISO 14224:2016 Free tool
In short

A maintenance programme is written once and then runs for years while the plant changes around it. PM optimisation is the loop that closes that gap — comparing what the plan says against what actually happened, and changing the plan where the evidence supports it.

NORSOK Z-008:2024 splits this into two clauses that this tool keeps strictly apart. §11.4 lists what should make you investigate. §11.5 lists what justifies changing the programme. They are not the same list, and a repeated failure is a reason to find out why, not a reason to shorten an interval.

You upload three exports — the plan register, PM completion history, failure history — and get a verdict per plan line with the clause, the sample size and the attribution behind it. It is free, it runs in your browser, and nothing is uploaded anywhere.

The Bluestream Toolbox PM Optimisation tool showing a 26-line PM plan reviewed against four years of completion and failure history. The verdict summary reads 1 shorten, 11 extend, 12 reconcile, 1 refer to barrier management and 1 refused. The first row is expanded to show the evidence behind the shorten recommendation, including the clause quoted, a counter-argument, the consequence class, the peer comparison with its raw and multiplicity-corrected p values, and the evidence grade.
Tool 02 in the Toolbox, reading a 26-line PM plan against four years of history. The first line is shortened because that pump’s bearings fail well above the rest of its peer group — 11 failures against the 2.2 expected of it. The card behind the verdict carries the §11.5 wording it rests on, the counter-argument against shortening, the peer comparison with both the raw and the multiplicity-corrected p, and the attribution tier — so the recommendation can be argued with rather than taken on trust.

Why programmes drift

Almost every preventive maintenance programme starts life as a reasonable set of assumptions: the vendor’s recommended intervals, a criticality study, and whatever the last plant did. Those assumptions were made before the equipment ran a single hour.

Then the plant runs. Some tasks turn out to be finding nothing, year after year. Some assets fail far more often than their identical siblings. The floor quietly settles into a rhythm that is not the one written on the plan. None of this shows up in a report, because a maintenance programme has no natural feedback loop — a PM that finds nothing looks exactly like a PM that was worth doing.

Z-008 puts the obligation plainly. §11.5: “A maintenance program shall be reviewed and updated at regular intervals. A continuous improvement process shall be in place in order to enhance safety levels and, at the same time, reduce equipment downtime and OPEX.” That is a shall. The difficulty was never whether to do it; it is that doing it by hand across a few thousand plan lines is a job nobody has time for.

The two clauses that decide everything

This is the part most PM optimisation exercises get wrong, and it is worth being pedantic about because it changes what you are allowed to conclude.

§11.4

Triggers for investigation

Eight of them, from HSE-related equipment failure to reduced energy efficiency. The clause scopes them to “finding the root cause for the deviation from the KPIs” and concludes that the root cause should be identified and necessary actions taken.

Output: an investigation.
§11.5

Initiations for a programme update

Seven bullets this time, and the first splits into two opposite limbs — eight initiations in all. Only two are visible in a CMMS export: a higher observed failure frequency, and a lower one or no observed damage at PM. The other six are things a human declares.

Output: a change to the plan.

The distinction matters because the natural instinct — “this pump keeps failing, shorten the interval” — jumps straight from a §11.4 trigger to a §11.5 action, skipping the step where you find out why it keeps failing. If the cause is misalignment at installation, a shorter inspection interval buys nothing and costs money.

So in this tool, nothing on the trigger register moves an interval — it produces an investigation and stops there. Every verdict card carries a button that hands that line straight to Root Cause Analysis, which is where a repeated failure belongs before anyone argues about its interval.

The PM Optimisation trigger register, listing the eight NORSOK Z-008 clause 11.4 triggers for a detailed investigation in the clause's own order. Each shows a state: clear, unavailable, or fired with a count. Notes beside each explain what it needs, including that the production-loss threshold has no default because the word unacceptable belongs to the operator.
The §11.4 register, in the clause’s own order. All eight are shown even where the data cannot support them — an auditor reading the clause counts eight bullets and will look for eight. Two have no source in a CMMS export at all and say so. Where a trigger needs a number the standard does not supply, the tool asks for yours rather than inventing one: “unacceptable” production loss is your threshold, not ours.

What you upload

Three exports. Only the first is required, and the tool tells you which analyses each one unlocks before you run it rather than leaving you to discover the gaps afterwards.

ExportWhat it isWhat it unlocks
PM plan register
required
One row per planned task: tag, interval, ideally a task id and a maximum allowed interval.Interval conformance, and the refusal register.
PM completion historyCompleted preventive orders with dates, and ideally the condition found.Reconcile plan against floor; the “no damage at PM” extension limb.
Failure / corrective historyCorrective orders or notifications with a failure date, ideally coded.Shorten, drop, and the §11.4 triggers.

§11.2 names the fields worth having, reprinting Table 3 from NS-EN ISO 14224:2016 — failure date, mode, mechanism, cause, impact, operating condition and detection method on one side; maintenance category, condition before and after, man-hours, spares, start and finish, active maintenance time and downtime on the other. Real exports carry perhaps half of that. The clause itself allows for this: the required reporting “will vary between systems and equipment and focus shall be on safety critical and production critical items”.

So a gap is not a conformance failure. It is a list of analyses that stay switched off, and the tool prints that list, plus a specification you can hand to whoever pulls the next export.

Nothing leaves your browser. The spreadsheets are parsed in the page. There is no upload, no server-side analysis and no model in the loop — every number in the output can be re-derived by hand from the exported CSV, which matters when someone in the review meeting decides to check one.

The verdicts

Each plan line gets exactly one, with the clause it rests on printed beside it.

VerdictRests onWhen
KeepNo §11.5 initiation is met. Shown with its evidence anyway — a silent keep is an unreviewed line.
Reconcile plan vs floor§11.5The plan says six months and the floor does nine. Decide which one you mean before optimising either.
Extend interval§11.5Lower failure frequency than the peer group, or no damage found at PM — and every gate below is passed.
Shorten interval§11.5Higher failure frequency than the peer group, on failures actually linked to this task.
Change strategy§9.1, §11.5A calendar task where the concept says condition-based, or a hidden failure with no function test.
Replace the unit§11.5Shortening would drive the interval below the shortest sensible rung. The same clause bullet allows it.
Drop to planned corrective§9.2.1The clause permits it where no PM is required or cost-effective — may, not shall, and the card says so.
Add task§9.2.2, §9.3, §9.4.2A concept line with nothing covering it, an unsafe failure mode with no task, or a barrier with no function test.
Investigate§11.4From the trigger register only. Never touches an interval.
Insufficient evidenceNames the exact field that would change the answer. Expect this to be the commonest verdict on a first upload.

Reconcile is the one that always works

It needs only a tag, a task and some dates — no failure coding, no condition field, no criticality study. And it is often the most valuable thing in the run, because a plan that says 180 days against a floor that runs at 270 is not an optimisation problem yet. It is a disagreement about what the programme actually is, and no interval arithmetic on top of it means anything until it is settled.

Earning an extension

§11.5’s second observable initiation reads: “lower failure frequency or no observed damage at PM can point towards extension of intervals or omitting certain tasks.” Two limbs, and the tool requires one of them plus a set of gates.

Then the gates: a consequence class to select against, a maximum allowed interval to stop at, enough observed cycles, not a safety-driven item, not a hidden-failure function test, and no bulk close-out in the data. The asymmetry is deliberate and the panel says so: it is much harder to earn an extension than a shortening, because absence of evidence over a short window is not evidence of reliability.

A clean function test is not good news about the interval. For a hidden failure, a clean test is the expected outcome — that is what hidden means. A run of them tells you the item has not failed, not that you can test it less often. The defensible bound is a tolerable multiple-failure probability from barrier analysis, which this tool does not hold and will not guess at. It refuses.

What it refuses, and why

The refusals are a first-class output, not error handling. They are listed on their own tab with counts, and each names the clause that forces it — or says plainly that it is a Bluestream rule rather than the standard’s.

There is also no OREDA benchmark. The comparisons this tool makes are against the peer group inside your own fleet and against the Bluestream concept library — which for a question about your plant is better evidence than a population average anyway.

The statistics, plainly

Z-008 supplies no formula for any of this. Every number below is a Bluestream choice, and the method note in the tool labels each one as such.

A tag is compared against the rest of its own peer group, leaving itself out of the baseline — otherwise a bad actor inflates the very average it is being judged against. The comparison is an exact conditional test rather than a rate against an estimated average, because on groups of a dozen the average is itself uncertain and pretending otherwise overstates the result.

Then the correction that most exercises skip. Testing twelve tags at the usual five per cent and reporting the hits is twelve tests, not one. Simulate a fleet where every tag genuinely shares one failure rate — so every flag is a false positive by construction — and that rule still raises at least one bad actor in roughly a third of runs; on a group of thirty tags, in more than half. Correcting for the size of the family removes most of that. Exactly how much depends on how the correction is arranged — which turns out to matter more than it sounds.

One test per tag, not two

There is one question per tag — does this tag differ from its peers? — so there is one test, and it is two-sided. Whether the tag sits above or below the group is read off afterwards from the smaller tail. It costs nothing extra, because a tag cannot be in both tails at once.

Asking the two questions separately instead — is it higher, and is it lower — and acting on whichever fires is a two-sided test at double the level, whether or not anyone says so. This tool did exactly that until August 2026: it corrected the two directions as two separate families and then reported a single family size, which described neither. On the simulated fleet above, that arrangement produced at least one false recommendation in 4.4–8.5 % of runs against 2.1–4.1 % after the fix, depending on group size and window length.

That is not free, and it would be dishonest to present it as free. The cost lands on marginal findings: a tag failing at three to four times its group’s rate is now detected about seven percentage points less often. The gap narrows to roughly two points at six times the group rate, and to nothing measurable at eight — a severe bad actor is caught exactly as readily as before. Halving the rate of unnecessary interventions is worth that trade, because a shorten verdict is not free either: it buys more intrusive work, more infant mortality, and a maintenance budget spent on a tag that was never the problem.

The card prints the raw p and the corrected one, with the size of the family it was corrected against. The corrected value on its own cannot be checked: it would not tell you whether a finding was strong in its own right or only survived because the family happened to be small. The one-sided tail for the reported direction, if you want it, is exactly half the raw figure.

Where zero failures were observed, no rate is reported as 0.00. An upper bound is printed instead, because zero failures in a short window and zero failures in a long one are not the same claim.

Measuring the programme

§11.3 asks for KPIs and offers Table 4 as examples — the clause’s own word, and the panel repeats it rather than presenting them as required outputs. The tool reproduces all seven rows with the standard’s own purpose and comment text beside each.

The PM Optimisation KPI panel showing NORSOK Z-008 Table 4's seven example key performance indicators, each with the standard's own purpose and comment text. Preventive workload reads 301 orders and 666 hours with a note that man-hours are present on 297 of them; the overdue rows show counts of open orders past their deadline; the failure-fraction row refuses to give a PFDavg.
Table 4, as the standard prints it. The workload row carries its own caveat — 666 hours across 301 preventive orders, with man-hours recorded on 297 of them — so the reader knows what the total is missing. Overdue is counted as a snapshot of open orders at the as-of date, never over completed history, because the latter counts every order that was ever late and produces a number nobody recognises.

§11.3 also asks for a combination of lagging and leading indicators. Table 4’s examples are predominantly backwards-looking and the panel says so; choosing a leading one is yours to do.

Common mistakes

Treating a bulk close-out as execution

When four hundred orders are technically completed on three dates, the observed intervals stop describing the floor and start describing an administrative tidy-up. It flattens the gaps and raises apparent compliance at the same time, so a compliance threshold does not catch it — it is satisfied by it. The tool detects the pattern and refuses the verdicts that are built on completion gaps for that run.

Guessing the maintenance category

If PM01 and ZM2 are not mapped to preventive and corrective, a corrective order can end up counted as evidence that a preventive task is working. The tool asks you to map the values it actually found and blocks the PM/CM analyses until the mapping covers almost everything.

Shortening the wrong task

A tag that fails often is a tag, not a task. If the bearings are failing, shortening the seal inspection adds cost and infant mortality and removes no failure. Where the data supports task-level attribution the tool uses it, and where it does not, it says the verdict is at asset level and claims nothing finer.

Extending on an empty column

The commonest way to get a wrong answer out of a PM review is to read “no failures recorded” as “no failures”. Recording gaps are the norm in maintenance data. That is why limb A requires a comparison and limb B requires the condition field.

Getting the file from your CMMS

The three exports map onto standard transactions. Pull them for the same window and the same tag scope.

SAP PM

Three transactions

Where each export comes from

IP24 — maintenance plans and items, for the plan register IW39 — order list, filtered to your preventive order types IW29 — notification list, which is where the damage and cause codes live

IW29 is the one worth arguing for. Orders carry what was done; notifications carry what was wrong, and the failure coding that decides whether a verdict can be attributed to a task at all comes from there.

IBM Maximo

Two applications

Where each export comes from

PM — the preventive maintenance application, for the plan register WOTRACK — work order tracking, split by work type into the two histories

Both histories come out of the same application, so export it twice with the work type filter set differently rather than trying to separate preventive from corrective afterwards.

Microsoft Dynamics 365

Asset Management

Path

Modules → Asset managementInquiries

Maintenance plans give you the register; work orders give you both histories, separated by their maintenance type. Take the maintenance type as a column rather than filtering it away — the tool asks you to map the values it finds, and it cannot map a value that was filtered out before export.

Include the order status and a due or overdue date if you can. Without them the tool cannot tell which orders are still open, and the four overdue KPIs are refused rather than computed over closed history.

References

Next steps

PM optimisation is where the Operate track closes back onto Develop. A shorten verdict with no known cause belongs in Root Cause Analysis; a change of strategy belongs back in RCM; a task that should never have been on the plan belongs in the concept that generated it.