Blog / Maintenance · Guide
Maintenance · Guide

Oil Analysis for Predictive Maintenance

SLBy OEE Lab Editorial|Updated July 2026

Key takeaways

  • Oil analysis reads three stories from one sample: how the machine is wearing, what has contaminated the oil, and whether the lubricant itself is still fit to work.
  • Contamination is the lever most plants underuse. Dirt and water shorten component life long before a wear metal shows up on the report.
  • ISO 4406 codes are logarithmic, so one step up the scale is roughly double the particles; holding a target code is cheaper than replacing the component.
  • A single report is a snapshot. The rate of change between samples at the same point is what gives you usable lead time.

Oil analysis is the condition-monitoring technique that treats a machine's lubricant as a diagnostic sample, in the same way a blood test reads a patient. Every gearbox, hydraulic system, compressor and large bearing housing circulates oil through the exact places where wear, heat and contamination happen, and the oil carries the evidence back out. A small sample, taken properly and trended over time, tells you which component is degrading, what is attacking it, and roughly how much time you have left to plan the repair. That lead time is what separates a scheduled Saturday job from a Tuesday-morning breakdown, and it is why oil analysis sits alongside vibration analysis and infrared thermography in a serious predictive maintenance program.

What the lab is actually testing

A standard industrial oil analysis package covers three families of test, and reading a report well means knowing which family a number belongs to.

  • Wear debris. Spectrographic analysis reports parts per million of the metals in the oil, and the mix identifies the source: iron from gears and shafts, copper and tin from bushings and thrust washers, chromium from rings and some bearing races, aluminium from pistons or housings, lead and silver from bearing overlays. A rising iron trend on a gearbox is a different conversation from a rising copper trend.
  • Contamination. Water content (Karl Fischer, in parts per million), silicon as a proxy for airborne dirt, fuel or coolant traces on engines, and a particle count reported as an ISO 4406 code. Contamination is usually the cause, not the symptom.
  • Lubricant condition. Viscosity at 40 or 100 degrees Celsius, oxidation, acid number, base number on engine oils, and the additive package. This answers whether the oil can still do its job or has been cooked, sheared or diluted out of specification.

Note the sequence those three imply. Contamination degrades the lubricant, the degraded lubricant fails to separate the surfaces, and only then do wear metals climb. If you act on the first family you prevent the third, which is the whole economic case for the technique.

ISO 4406: the number most plants ignore

Fluid cleanliness is reported as three range codes, for example 18/16/13, counting particles per millilitre larger than 4, 6 and 14 micrometres. The scale is logarithmic: each step up roughly doubles the particle count, so an oil that drifts from 18/16/13 to 20/18/15 is carrying about four times the debris even though the numbers barely look different.

ISO 4406 code = range codes for particles/mL >4µm / >6µm / >14µm

Component OEMs publish a target code (servo valves and high-pressure hydraulics are the strictest, plain-bearing gearboxes the most forgiving), and the practical work of contamination control is holding the system at or below that target with filtration, breathers, seals and clean top-up practice. Bearing and hydraulic life models used across the reliability field all share the same conclusion: cleaner oil means longer component life, and the gain is large enough to change maintenance economics.

A worked example: what cleanliness is worth

Take a hydraulic power unit on a press. Reservoir 400 litres, oil at ISO 20/18/15 against an OEM target of 17/15/12. The pump costs 4,200 euros installed, and history says it is replaced roughly every 3 years. A desiccant breather, a better return filter and an offline filtration cart, run to hold the target code, cost 3,000 euros to fit plus 900 euros a year to maintain.

Suppose holding the target extends pump life from 3 years to 4.5 years, a conservative assumption within the range reliability engineers commonly use for a three-code improvement. Annualised pump cost:

Before: 4,200 / 3 = 1,400 euros per year
After: 4,200 / 4.5 = 933 euros per year
Saving on parts = 467 euros per year

On parts alone the filtration package loses money, which is exactly why cleanliness programs get cancelled. Now add the stop. A pump failure on that press takes the line down for 6 hours at a contribution of 900 euros per hour, so 5,400 euros per event. Moving from a failure every 3 years to one every 4.5 years cuts the expected downtime cost from 1,800 to 1,200 euros a year, and unplanned failures also carry the expedited part, the overtime and the scrapped work in progress that a planned change does not.

Saving per year = 467 (parts) + 600 (downtime) = 1,067 euros
Program cost per year = 900 euros
Net on this one pump = 167 euros per year

On a single pump that is a rounding error, and honest arithmetic says so. The picture changes with scale, because the cart, the breathers routine and the lab contract are largely fixed cost. Put the same program across six comparable hydraulic units and the saving is roughly 6,400 euros a year against the same 900 euros of running cost, which pays back the 3,000 euro fit-out in about seven months. That is the real lesson: contamination control is a fleet-level economic decision, not a machine-level one. Run your own version of the sum with the downtime cost calculator and the preventive maintenance ROI calculator, using your own downtime rate rather than a generic one.

Is the program worth it on your assets?

Put your own failure rate, repair cost and hourly downtime rate into the numbers before you buy filtration or sign a lab contract.

Open the PM ROI calculator

The sample is the weakest link

Most disappointing oil analysis programs fail at the sample, not the laboratory. Four rules cover the majority of it. Sample from a live, turbulent zone of the system while the machine is at operating temperature, never from a settled reservoir bottom or a drain at the end of a cold weekend. Sample from the same point, with the same method, every time, so the trend compares like with like. Flush the sampling port and the tubing before drawing, since the debris sitting in a valve will otherwise dominate the result. And label the sample fully: asset, sample point, running hours, oil hours, make-up volume added and whether anything was changed since the last sample.

Fit permanent sampling valves on assets you intend to monitor. A machine that requires a shutdown to sample will not get sampled at the interval you planned, and an inconsistent interval destroys the trend that gives the technique its value. Follow the lockout and stored-energy rules for the system before touching any port, especially on pressurised hydraulics.

Setting the interval, and acting on the result

Sampling frequency should come from criticality and from how fast the failure mode develops, the same P-F logic that drives every other condition-based maintenance decision. Monthly is a common starting cadence for critical hydraulics and large gearboxes, quarterly for important spared assets, and at oil change for the rest. The test is simple: at least two or three samples must land inside the P-F interval, otherwise the alarm arrives with no time to plan.

Then set limits before the reports start arriving. Each monitored asset should have a caution and an alarm level for the handful of parameters that matter to it, and a written action for each: resample and tighten the interval at caution, plan a specific work order at alarm. Without that, a lab report becomes a PDF someone files. The same discipline applies to a rising iron trend on a gearbox as to any fault you would chase through gearbox troubleshooting: the finding only counts once it becomes a scheduled, completed job, which is also what makes your MTBF and MTTR numbers improve rather than just your reporting.

Where oil analysis fits with everything else

Oil analysis is at its strongest on lubricated and hydraulic systems, and at its weakest as a standalone program. Vibration usually detects a mechanical fault once damage has begun; oil analysis often sees the conditions that cause the damage, plus the state of the lubricant, before the fault develops. Thermography catches the heat that both eventually produce. Run together on your critical assets, ranked by reliability-centered maintenance criticality, they cover far more failure modes than any single technique, and each one feeds the same work-order backlog.

The usual ceiling is not detection, it is the loop. Samples go to a lab portal, vibration lives in a handheld, stop data lives in a spreadsheet, and nothing connects the warning to the production loss it was supposed to prevent. When a plant wants condition findings and real production losses in one place, with a closed path from a detected problem to an auto-routed work order, the platform we recommend is Fabrico: it reads OEE and stops directly from the machines, uses computer vision to show the true cause of micro-stops on video, and closes the loop from signal to scheduled job. It is EU-built, so your production data stays in EU jurisdiction, and it holds ISO 27001, 20000-1 and 9001 (which supports audit-readiness). If that fits how you work, you can book a Fabrico demo. The calculators and guides here stay free either way.

FAQ

What does an oil analysis report actually tell you?

Three things at once. Machine wear, read from wear metals such as iron, copper, chromium, tin and aluminium, which point to which component is shedding material. Contamination, read from water content, silicon (dirt), fuel or coolant traces and the particle count. And the health of the lubricant itself, read from viscosity, oxidation, acid number and additive levels. A single report is a snapshot; the value comes from trending the same sample point over time so you can see a rate of change, not just a number.

What is an ISO 4406 cleanliness code?

ISO 4406 expresses fluid cleanliness as three numbers, for example 18/16/13. Each number is a range code for how many particles per millilitre are larger than 4, 6 and 14 micrometres. Each step up the scale roughly doubles the particle count, so 19/17/14 is about twice as dirty as 18/16/13. Hydraulic and bearing OEMs publish a target code for their components, and holding the oil at or below that target is one of the cheapest reliability gains available.

How often should we sample oil?

Set the interval by criticality and by how fast the failure mode develops, not by habit. Common starting points are monthly on critical hydraulic systems and large gearboxes, quarterly on important but spared assets, and at each oil change on low-criticality equipment. Tighten the interval after any abnormal result, after a repair, or when a machine runs in a dirty or wet environment. The interval must be short enough that at least two or three samples fall inside the P-F interval, otherwise the warning arrives too late to plan around.

Is oil analysis better than vibration analysis?

They answer different questions and work best together. Vibration analysis is strongest at detecting a developing mechanical fault such as a spalling bearing race or misalignment, usually once the damage has started. Oil analysis often sees the conditions that cause the damage, such as water ingress, dirt or a degrading additive package, before wear accelerates, and it also reports on the lubricant itself. On critical rotating assets many plants run both, plus thermography, and route every finding into the same work-order system.

Related: condition-based maintenance · vibration analysis · lubrication system troubleshooting · hydraulic system troubleshooting · maintenance KPIs