Home/Field resources/Chasing the Intermittent Fault
Diagnostic field guide · free PDF

Chasing the Intermittent Fault

The fault that clears before anyone reaches the panel. How to capture evidence, sort the usual causes, and set a trap so the next occurrence tells you something.

Get the library by email →   Browse the full library

One email, the whole 24-guide library. No drip campaign, no spam.

From the guide

How it reads

Excerpt

Who this is for and what it does: a maintenance tech, controls engineer, or plant reliability person who keeps getting handed a fault that clears itself before anyone reaches the panel. This walks you through capturing evidence, narrowing the cause, and setting a trap so the next occurrence actually tells you something.

How to use this Read the triage order once. When the fault hits, work from evidence, not the last thing someone swapped. If it's already cleared, jump to "Set a trap" and get instrumentation in place first. The frustrating thing about an intermittent is that the failure and the diagnosis rarely coincide. By the time you reach the panel with a laptop, the terminal is tight again, the condensation has burned off, the noise source has cycled off, and everything meters fine. So the first job is never "fix it." The first job is to make the fault leave a record.

Step one: capture evidence before you touch anything An intermittent won't show itself on demand, so before you start swapping parts, get these working: First-out / first-fault logging. Most controllers and safety systems record which condition tripped first in a cascade. One device faulting looks identical to ten until first-out tells you which moved first. Turn it on if it isn't already; if your platform is an older or basic one that has no first-out at all, note that and lean on the external trap described at the end of this guide. Event and alarm logging with real timestamps. Confirm the clock is right and, if you have more than one device, that they're synchronized. A programmable logic controller (PLC) log that says 14:32:07 and a drive log that says 14:29:51 for the same event will send you chasing a phantom. Fix time sync before you trust any correlation. Trending on the suspect signals. Log the analog values, comms status bits, supply voltages, and digital inputs around the fault. Sample fast enough to catch the event you're chasing: a onceper-second trend misses a 200-millisecond dropout entirely, and even a 200-ms sample can straddle and miss it. To reliably catch a glitch that short, sample well under 100 ms, or better, use a

Inside the guide

What it covers

  • Step one: capture evidence before you touch anything
  • The usual families of cause
  • Symptom you see
  • Likely causes
  • What to check / what to log
  • Fault correlates with

Dealing with this on your plant right now?

Set the trap and it’s still slipping through? Send me what the trap caught and I’ll take a look.

Get in touch Call (707) 690-7054