Why PID Loops Oscillate: A Practical Guide to Finding the Real Cause and Fixing It
Why PID Loops Oscillate: A Practical Guide to Finding the Real Cause and Fixing It
Quick answer: why PID loops oscillate
A proportional-integral-derivative (PID) loop oscillates when the complete feedback path has enough gain and phase lag to sustain repeated movement, or when a nonlinearity, interacting loop, sampling effect or external disturbance forces a cycle. The PID may be creating the oscillation, amplifying it or correctly responding to a real periodic disturbance. Classify the source before retuning.
PID should not be confused with P&ID, which means piping and instrumentation diagram. A P&ID shows equipment, piping and instrumentation relationships. A PID controller uses proportional, integral and sometimes derivative action to regulate a process variable.
Use this seven-step workflow:
- Recognize: Confirm that the movement is recurring rather than random noise or a normal transient.
- Measure: Calculate period, frequency, amplitude and whether successive peaks decay, remain steady or grow.
- Classify: Decide whether the evidence points toward tuning, process dynamics, saturation, valve behavior, measurement, interaction or external forcing.
- Test: Perform only authorized, bounded checks that can distinguish competing explanations.
- Correct: Address the demonstrated cause rather than applying a generic detuning change.
- Verify: Repeat the same measurements under comparable operating conditions.
- Document: Record evidence, approvals, original settings, final settings, limits, modes and restoration status.
| Possible cause | Typical trend signature | Best next check |
|---|---|---|
| Excessive loop gain | Smooth, repeating PV and controller output; amplitude changes substantially when gain is reduced | Verify controller form, action and scaling before a cautious gain change |
| Excessive integral action | Controller output reverses only after accumulated error changes; damping improves when Ti is increased | Confirm whether the setting is integral time or repeats per minute |
| Valve stiction | Controller output ramps while actual travel is stationary, followed by a travel jump | Trend actual valve travel, positioner signals and supply pressure |
| Saturation and windup | Applied output is pinned while nominal output continues beyond the limit; recovery is delayed | Compare nominal and applied output and inspect anti-windup tracking |
| Derivative acting on noise | High-frequency output chatter while PV is much smoother at the process timescale | Inspect raw PV, quantization, derivative location and filtering |
| Dead time and multiple lags | Output changes, but PV responds late enough to produce substantial phase lag | Estimate delay and lag from an approved test or existing disturbance |
| External periodic disturbance | Cycling persists in manual while actual actuator travel remains constant | Trace upstream, utility, pressure, feed, composition or analyzer forcing |
| Loop interaction | Several loops share a period, with one signal consistently changing first | Align timestamps and inspect shared process paths and utilities |
The key rule is simple: do not retune until the oscillation source has been classified.
Safety before diagnosis or testing
Passive diagnosis from existing synchronized trends should normally come before active testing.
What PID loop oscillation looks like
Terminology: oscillation, cycling, hunting and instability
An oscillation is repeated movement around a value or trajectory. Operators may also call it cycling or hunting. These words describe what the trend looks like; they do not identify the cause.
Instability is more specific. In a linear analysis, instability means deviations grow rather than settle. A loop can nevertheless show a sustained cycle without being linearly unstable. Saturation, stiction, deadband, backlash, relay behavior or quantization can create a bounded nonlinear limit cycle.
Calling every repeating trend “bad tuning” skips the most important diagnostic step.
Oscillation versus a normal control response
After a setpoint or load change, a functioning loop may overshoot and cross its target one or more times. That is not automatically a sustained oscillation.
Three basic response patterns help establish urgency:
- Decaying response: Each same-side peak becomes smaller. The loop is returning toward equilibrium.
- Sustained response: Peak amplitude remains approximately constant. Linear gain and phase conditions or a nonlinear mechanism may be sustaining the cycle.
- Growing response: Successive peaks become larger. The loop may be unstable and can reach output limits, alarms or trips.
- Non-oscillatory response: The PV approaches its target without crossing it repeatedly. It may still be too slow or limited, but it is not cycling.
A growing trend deserves prompt attention, but a bounded trend should not be dismissed. A sustained limit cycle can repeatedly move a valve, consume utilities, disturb downstream units and hide smaller process changes.
Types of oscillation and the measurements that describe them
A useful description should quantify the cycle instead of relying on “fast,” “large” or “almost stable.”
Measure several cycles where possible. A single peak interval can be distorted by noise, a setpoint change or historian sampling. Record the baseline used for amplitude and whether peaks were measured from raw or filtered data.
A self-regulating process reaches a new steady value after a fixed input change. An integrating process continues to ramp while the net input imbalance remains. Their trends can look different under the same controller behavior, so process type matters when interpreting waveform shape.
Forced versus self-sustained oscillation
A self-sustained oscillation originates within the closed feedback path. Excessive gain, integral action combined with delay, valve stiction or a quantized controller-actuator path can keep the loop moving without an external periodic input.
A forced oscillation enters from elsewhere. Examples include an upstream cycle, utility pressure variation, periodic feed composition, an analyzer cycle or another controller. In this case, the PID may be doing exactly what it should: responding to a real disturbance.
Manual mode can help distinguish the categories, but it is not a complete verdict:
- If cycling persists in manual and actual actuator travel is constant, suspect external forcing.
- If cycling persists in manual and travel continues moving, inspect the positioner, air supply, actuator and valve mechanics.
- If cycling stops in manual, the feedback loop was necessary for that cycle, but this does not distinguish poor tuning from stiction or another feedback-dependent nonlinearity.
Waveform shape and manual behavior are clues. Neither should be used as proof without checking actual travel and relevant disturbances.
PID control in brief
A PID controller combines:
- Proportional action, based on present error.
- Integral action, based on accumulated error.
- Derivative action, based on the rate of change of error or another selected signal.
Common signal names are:
- SP: Setpoint.
- PV: Process variable.
- CO or OP: Controller output.
- Applied CO: Output after limits, selectors, rate limits and other downstream logic.
- Valve travel: Measured physical position. It is not the same signal as CO.
Controller forms, units and action
Different PID equations can use identical parameter names while producing different behavior. Never transfer tuning numbers without identifying the controller form.
| Form | Equation or representation | Parameter meaning | Important conversion |
|---|---|---|---|
| Parallel | u = Kp e + Ki ∫e dt + Kd de/dt + b | Kp: output/error; Ki: output/(error·time); Kd: output·time/error | Ki is an integral gain, not an integral time |
| Ideal or ISA | u = Kc[e + (1/Ti)∫e dt + Td de/dt] + b | Kc: controller gain; Ti: integral time; Td: derivative time | Kp=Kc; Ki=Kc/Ti; Kd=Kc·Td |
| Series or interacting | C(s)=Kc(1+1/(Ti s))(1+Td s) | Integral and derivative factors interact | Expansion adds proportional contribution, so numbers are not directly interchangeable with ideal form |
For ideal form, a smaller Ti produces stronger integral action. If the implementation uses repeats per minute, Ri = 1/Ti when Ti is in minutes, so a larger repeats-per-minute value produces stronger action. Confusing Ti with Ki or repeats per minute can create factor-of-60 errors when seconds and minutes are mixed.
Proportional band is another source of confusion. PB(%) = 100/|Kc| applies only when signals are consistently normalized to their spans. It should not be applied blindly to engineering-unit signals.
The expected direction of control must also be documented. One convention may define error as SP−PV; another may use PV−SP with a corresponding direct/reverse setting. Verify the expected physical result: if PV rises, should CO rise or fall? Wrong action changes negative feedback into positive feedback and can cause rapid divergence.
Derivative on error reacts to a setpoint step and can create derivative kick. Derivative on PV avoids that specific kick. Two-degree-of-freedom implementations can weight setpoint terms separately, but vendor definitions vary. Derivative also magnifies rapid measurement changes approximately in proportion to ΔPV/Δt. Filtering can reduce chatter while adding phase lag; excessive filtering can therefore worsen stability. There is no universal derivative-filter value.
Related topic: derivative kick and derivative on measurement.
Where oscillation can enter the loop
A controller is only one component in the feedback path. The complete loop includes the sensor, filtering, controller calculations, limits, network or scan delays, actuator, positioner, valve and process. Disturbances and interacting controllers can enter at several points.
The practical question is not merely “Is the PID oscillating?” It is: Where does the repeating movement first appear, and what mechanism sustains it?
Signals to trend before changing tuning
Trend signals on one synchronized time base
A useful diagnostic trend should include more than SP, PV and the displayed output. Collect the following on a synchronized time base:
- SP, including remote or cascade setpoint.
- Raw PV and the filtered PV used by the controller.
- Nominal CO before limiting.
- Applied CO after limits, selectors or overrides.
- Actual valve or actuator travel.
- Controller mode and cascade status.
- Output, rate and integral limits.
- Positioner output and supply pressure where available.
- Relevant feed, utility, composition, pressure and analyzer signals.
- SP, PV and CO from suspected interacting loops.
- Alarm, override or selector states relevant to the event.
If the controller display provides only one output, determine whether it is nominal or applied. A nominal request of 140% may appear as an applied 100% after limiting. Likewise, a 55% CO signal does not prove that a valve traveled to 55%.
Related topic: how to read process trends.
Sampling for diagnosis
Sampling must capture the fastest feature relevant to the suspected mechanism. Slow historian data may show the process cycle while completely missing valve jumps or output chatter.
As an illustrative planning target—not a universal requirement—collect about 10–20 samples across the shortest feature of interest and include multiple complete cycles. A one-second valve jump requires a different collection rate from a 20-minute vessel-temperature cycle.
Record:
- Sample interval and timestamp resolution.
- Controller scan and historian collection rates.
- Clock alignment between systems.
- Missing values and interpolation.
- Compression or exception-reporting settings.
- Filter settings.
- Quantization.
- Task jitter, communication delay or execution-order changes where relevant.
Do not differentiate a heavily quantized or exception-compressed PV and then interpret the resulting spikes as process physics.
Waveform and phase clues
There is no universal rule that PV and CO must be “in phase” or “out of phase” for a particular fault. Their relationship depends on process gain sign, controller action, process type, frequency, dead time, filtering and whether the plotted output represents actual actuator movement.
For a positive-gain first-order-plus-dead-time process,
Gp(s) = K e^(−θs)/(τs+1)
the phase from CO to PV at angular frequency ω is:
phase(PV/CO) = −ωθ − atan(ωτ)
Dead time and process lag both add phase delay. At the ultimate frequency of the stated model, PV is 180° behind CO, but that result cannot be converted into a universal trend-reading rule for every implementation.
| Observed pattern | Consistent with | Interpretation boundary | Next check |
|---|---|---|---|
| Smooth, roughly sinusoidal PV and CO | Linear gain/phase-margin problem | Not proof of poor tuning | Compare actual travel, constraints and common-frequency signals |
| Square-like PV with triangular or sawtooth CO | Stiction in some self-regulating loops | Recognized clue, not proof | Look for flat travel while CO ramps, then a jump |
| Ramp or triangular PV | Integrating process under a cycling input | A process-dependent clue rather than a universal pattern | Confirm process type and material/energy balance |
| Flat CO tops or bottoms | Saturation or clipping | Display scaling can conceal downstream limits | Compare nominal CO with applied CO |
| CO ramps while travel is flat, then travel jumps | Stick-slip | Stronger evidence than PV/CO shape | Check positioner, air, linkage and valve mechanics |
| Travel fails to respond after reversal | Backlash or reversal deadband | One-direction tests can miss it | Use an authorized bidirectional test |
| Fixed PV–CO phase offset | Gain sign, action, process dynamics, dead time or filtering | No universal phase rule | Check implementation and estimate dynamic delay |
| Same period across several loops | Interaction or common forcing | Earliest signal is a candidate source, not proof | Align clocks and trace shared paths |
| Cycling persists in manual with constant travel | External disturbance | Confirm manual mode and actual travel | Trace feed, utility, pressure, composition or analyzer forcing |
| Cycling persists in manual with moving travel | Positioner, air or mechanical behavior | May be autonomous actuator cycling | Inspect the final-control element |
Cross-correlation, PV–CO plots, ellipse or parallelogram methods and higher-order statistics can screen routine data, but they can produce false positives and negatives. Confirm any automated diagnosis with travel feedback or a controlled test.
Diagnostic decision tree
Use the following sequence rather than starting with a tuning change:
- Confirm recurrence. Is the pattern repeated, or is it a single disturbance and recovery?
- Quantify it. Record period, frequency, amplitude and whether peaks decay, remain sustained or grow.
- Align the evidence. Trend SP, PV, raw PV, nominal CO, applied CO, valve travel, mode, limits and likely disturbances.
- Check constraints and data quality. Look for saturation, clipping, rate limits, selectors, quantization, filtering, compression and scan-related timing.
- Verify implementation. Confirm PID form, error equation, action, gain scaling, integral units, time base and derivative location.
- Compare command with movement. Does actual travel follow applied CO smoothly and promptly?
- Classify the process. Determine whether it is primarily self-regulating or integrating in the operating region.
- Compare other loops. Search for common periods, shared utilities and upstream signals.
- Use an authorized discriminating test if needed. Manual mode or a bounded change may help separate feedback-dependent cycling from external forcing.
- Correct the demonstrated cause.
- Verify under representative conditions.
- Document evidence, approvals, final configuration and restoration.
A planned downloadable version of this decision tree is proposed for field use: download link: placeholder.
Common causes of PID loop oscillation
Excessive proportional gain
Proportional action changes CO in direct proportion to the current error. Increasing controller gain can improve responsiveness, but it also magnifies loop movement. At a frequency where the complete feedback path has nearly 180° of phase lag, sufficient gain turns corrective action into reinforcement of the preceding movement.
A smooth, approximately sinusoidal PV and CO cycle is consistent with excessive loop gain, especially if amplitude or damping changes substantially after Kc is reduced. It is not proof. Valve motion, output constraints and periodic disturbances can produce similar PV trends.
A gain reduction that makes a trend calmer does not prove the original problem was tuning. Detuning can conceal a mechanical or interaction problem while reducing control performance.
Excessive integral action
Integral action accumulates error and drives steady-state offset toward zero. It also adds phase lag. If it is too aggressive relative to process delay and lag, CO can keep moving after the process has begun an unseen response.
In ideal form:
integral contribution = Kc(1/Ti) ∫e dt
A smaller Ti means stronger integral action. In parallel form, Ki is an integral gain, and larger Ki means stronger action. If the controller uses repeats per minute:
Ri = 1/Ti, with Ti in minutes.
A loop dominated by excessive integral may show delayed output reversals because the integrator must unwind accumulated error. Increasing Ti can improve damping and change the period, but it also affects load-disturbance recovery. The change must therefore be evaluated rather than judged from one quiet trend.
Related topic: PID tuning methods and when each fails.
Derivative action, noise and filtering
Derivative anticipates movement by responding to a rate of change. That can add damping in suitable applications, but differentiation magnifies rapid measurement changes, quantization and communication steps.
A noisy PV may produce high-frequency CO chatter even when the physical process changes slowly. Derivative on error can also react sharply to a setpoint step. Derivative on PV avoids that specific setpoint kick but is not universally required.
Filtering derivative or PV can reduce chatter. However, filtering adds delay and phase loss. Adding a stronger filter without evaluating stability can trade visible noise for a slower oscillation. Vendor definitions of derivative filter factor N differ; there is no universal value.
Do not assume derivative always damps an oscillation. Poor implementation, noisy measurement or excessive filtering can aggravate it.
Dead time and multiple process lags
Dead time is the interval between an input change and the first observable effect on the measured output. During that interval, the controller receives no evidence that its last action is working.
A first-order-plus-dead-time model is:
Gp(s) = K e^(−θs)/(τs+1)
where:
Kis process gain.θis dead time.τis the first-order time constant.
This model is a simplification, but it illustrates the mechanism. The delay term contributes phase lag proportional to frequency, while the first-order lag adds atan(ωτ). Multiple lags, filters, valve dynamics and computation delays can add further phase loss.
Dead time often comes from transport, analyzer cycles, thermal propagation, mixing, communication or batch sequencing. It cannot be removed by increasing gain. Aggressive gain or integral action generally becomes less tolerable as delay becomes more important relative to the process time scale.
Before estimating dead time from a trend, check timestamp alignment and filters. A historian timestamp or slow sensor filter can appear as process dead time.
Related topic: dead time in process control.
Operating-point nonlinearity
A loop may be stable at one flow, level, pressure, product grade or valve position and oscillatory at another. Process gain, time constant, dead time and valve gain can all change across the operating envelope.
A tuning set chosen at one point should not automatically be claimed as suitable everywhere. Identify where stability and performance change, then determine whether operating-region tuning, gain scheduling, valve characterization or process modification is justified.
Wrong action, scaling or units
A direct/reverse-action error can create positive feedback: rising PV produces a CO movement that drives PV further upward. Scaling errors can make moderate-looking tuning numbers extremely aggressive. Common mistakes include using normalized-span relationships on engineering-unit signals, confusing gain with proportional band and mixing seconds, minutes and repeats per minute.
Do not change action simply because CO and PV appear to move in the same direction. Expected direction depends on the process and valve arrangement.
Saturation and integral windup
Saturation occurs when the requested output exceeds what the applied output or actuator can deliver. Output limits, selector stations, rate limits, override control and physical actuator limits can all create a difference between nominal and applied CO.
Integral windup occurs when the integrator continues accumulating error while the requested action cannot be delivered. When the process finally becomes reachable, the stored integral contribution can delay output re-entry and produce a large recovery excursion.
Compare these signals:
- Nominal CO before limiting.
- Applied CO after limits and selectors.
- Actual travel.
- Integral state or contribution where available.
Anti-windup methods include:
- Clamping or conditional integration: Stop or condition integration while saturation would drive the output farther into the limit.
- Back-calculation: Feed the difference between nominal and applied output back to the integrator. A tracking time constant
Ttcontrols correction speed. - External reset or tracking: Make the integrator track an actual actuator or selected signal, useful around overrides, selectors and cascade limits.
A representative back-calculation expression is:
I-dot = Ki e + (u_applied − u_nominal)/Tt
Signs depend on the implementation. Feedforward contributions must be included consistently in tracking.
No anti-windup method creates missing actuator capacity. Clamping and back-calculation also recover differently, and back-calculation adds another setting.
Related topic: integral windup and anti-windup.
Valve stiction, deadband, backlash and hysteresis
The controller can calculate the correct output while the final-control element fails to deliver it. The strongest practical evidence comes from comparing applied CO with actual travel.
| Term | Practical definition | Distinguishing feature |
|---|---|---|
| Stiction | Resistance to the start of motion: static friction holds the valve until breakaway, then travel slips or jumps | CO can ramp while travel remains stationary, followed by a jump |
| Deadband | Input range over which no observable output change occurs | Describes observed input/output behavior; backlash can be one cause |
| Backlash | Lost motion after the direction of input reverses | Travel response is delayed specifically around reversal |
| Hysteresis | Output depends on previous excursions and direction | The same input can correspond to different outputs depending on history |
| Dead zone | Input region that produces no output response | Not necessarily associated with reversal; distinct from deadband |
In a stick-slip cycle, integral action changes CO while the valve remains stuck. Once the command exceeds breakaway resistance, the valve jumps. The error reverses, the command changes direction, and the sequence repeats.
A square-like PV with triangular or sawtooth CO can be a stiction clue in some self-regulating loops. It is not proof. Flat travel while CO changes and then jumps is much stronger evidence.
A useful way to state the evidence:
Valve stiction is a mechanical nonlinearity, so ordinary retuning does not remove the underlying fault. With integral action present, a stick-slip limit cycle can persist over a broad range of tuning. Changing gain or integral time often changes the cycle period and may change its amplitude, but the result depends on process dynamics, process type, operating point, stiction model, and controller structure. A calmer trend after detuning is not proof that stiction has been repaired.
No single rule describes how tuning changes affect a stiction cycle: amplitude and period responses depend on the model and the loop.
A one-direction bump may miss backlash. An authorized bidirectional test is more informative because reversal exposes lost motion. Positioner output, air supply, actuator pressure, linkage and packing condition may also be relevant.
Related topic: valve stiction explained.
Measurement, sampling, filtering and quantization
A loop can oscillate because the controller is acting on an inaccurate or poorly timed representation of the process.
Potential mechanisms include:
- Sensor noise or intermittent connections.
- Quantized signals that move in coarse steps.
- Historian compression that hides fast changes.
- Aliasing when the sample rate is too low.
- Communication delay or jitter.
- Controller tasks executing in an unexpected order.
- Filters adding more phase lag than expected.
- A control update tied to a slow analyzer or network scan.
A small limit cycle spanning one or a few digital counts can come from quantization. A cycle tied closely to scan or update timing suggests computation or sampling behavior, but it must be checked against physical periodic disturbances at the same timing.
Adding a filter is not a neutral repair. It may suppress noise while adding phase lag. Reassess stability after any filter change.
Interaction, cascade loops and split-range systems
A loop can cycle because another controller changes the same process path. Several loops may share a manipulated stream, pressure header, heat source, vessel inventory or upstream disturbance.
When several loops show the same period:
- Align timestamps accurately.
- Compare waveform distortion as well as period.
- Identify which signal changes earliest and most cleanly.
- Trace physical and control-system connections.
- Treat the earliest signal as a candidate source, not proof.
In cascade control, the primary controller sets the setpoint of a faster secondary loop. If the secondary loop is not sufficiently faster, is saturated or has poor actuator control, the primary may repeatedly demand corrections that the inner loop cannot deliver.
In split-range control, one controller output drives two or more final elements across different ranges. Cycling can concentrate around the split point if there is a gap, overlap, discontinuity, incorrect characterization or selector problem.
Selectors and override controllers can also cause nominal and applied outputs to differ. An integrator that does not track the selected signal may wind up even though the displayed controller output appears reasonable.
Related topic: cascade control.
Related topic: split-range control basics.
External periodic disturbances
The PID may be responding correctly to a disturbance that repeats. Sources can include upstream cycling, utility-pressure variation, composition changes, analyzer updates or another unit.
A strong diagnostic observation is a cycle that persists in manual while actual travel remains constant. That points away from the controller calculation and toward external forcing. If travel keeps moving in manual, investigate the positioner, actuator, air supply and valve instead.
Ultimate gain and ultimate period
The Ziegler–Nichols closed-loop method uses:
- Ultimate gain, Ku: Proportional-only gain at which the idealized loop reaches a sustained cycle.
- Ultimate period, Tu: Period of that cycle.
For the ideal PID form Kc[1 + 1/(Ti s) + Td s], the traditional settings are:
- P:
Kc = 0.5Ku - PI:
Kc = 0.45Ku,Ti = Tu/1.2 - PID:
Kc = 0.6Ku,Ti = Tu/2,Td = Tu/8
Parallel equivalents are:
- PI:
Kp = 0.45Ku,Ki = 0.54Ku/Tu - PID:
Kp = 0.6Ku,Ki = 1.2Ku/Tu,Kd = 0.075KuTu
Do not label Ki as integral time. Do not transfer these values into series or vendor-specific forms without conversion.
The method is historically associated with a quarter-amplitude-decay objective and can be aggressive for many applications. It does not adequately account for noise, hysteresis, constraints, valve wear, interaction or safety. It should not be called optimal, and no single tuning rule is universal.
Relay results can be misleading with stiction, backlash, deadband, saturation, rate limits, integrating or open-loop-unstable plants, asymmetry, noise chatter, selectors and interacting loops.
Textbooks or qualified training can help engineers understand the method’s assumptions, but no book, course or product can make an unsafe plant test safe without site-specific review and authorization.
Worked examples using scripted synthetic data
All four examples below use scripted synthetic data, not plant, field or laboratory data. Model parameters were chosen for teaching. Generated signals are illustrative, and reported metrics are calculated from those signals. They are not universal tuning recommendations.
Example 1: measuring ultimate period, decay ratio and clipping
SCRIPTED SYNTHETIC DATA — NOT PLANT DATA
A normalized loop uses the convention e = SP−PV, positive process and controller gain, and signals expressed as percent of span. A synthetic sustained record has:
- Sample interval: 5 seconds, or approximately 0.0833 minute.
- Record duration: 72 minutes.
- Baseline PV: 50%.
- Candidate ultimate gain:
Ku = 4.0 %CO/%PV.
Eight same-side PV peaks occur at approximately:
12.0, 19.8333, 28.0, 35.9167, 43.9167, 52.0, 60.0833 and 68.0833 min
The peak values are:
55.0239, 54.9969, 55.0100, 54.9934, 55.0116, 55.0199, 55.0042 and 55.0065%
The seven measured periods are:
7.8333, 8.1667, 7.9167, 8.0, 8.0833, 8.0833 and 8.0 min
Using the first and last peak:
Tu = (68.0833−12.0)/7 = 8.011904761904761 min
For practical reporting, Tu ≈ 8.0 min. Peak-time uncertainty is at least one sample interval.
The corresponding values are:
fu = 0.12481426448737 cycles/minωu = 0.78423115275347 rad/min- Full-record PV range:
10.060964631539314percentage points - Half-range amplitude:
5.030482315769657points
The full-record range includes noise, so it slightly exceeds the nominal five-point peak amplitude.
| Dataset | Controller form | Calculated settings |
|---|---|---|
| Ex1 nominal: Ku=4.0, Tu=8.0 min | P | Kc=2.0 |
| Ex1 nominal | PI ideal | Kc=1.8; Ti=6.666666666666667 min |
| Ex1 nominal | PI parallel | Kp=1.8; Ki=0.27 %CO/(%PV·min) |
| Ex1 nominal | PID ideal | Kc=2.4; Ti=4.0 min; Td=1.0 min |
| Ex1 nominal | PID parallel | Kp=2.4; Ki=0.6 %CO/(%PV·min); Kd=2.4 %CO·min/%PV |
| Ex1 measured: Ku=4.0, Tu=8.011904761904761 min | PI ideal | Kc=1.8; Ti=6.676587301587301 min |
| Ex1 measured | PI parallel | Kp=1.8; Ki=0.26959881129271923 %CO/(%PV·min) |
| Ex1 measured | PID ideal | Kc=2.4; Ti=4.0059523809523805 min; Td=1.0014880952380951 min |
| Ex1 measured | PID parallel | Kp=2.4; Ki=0.599108469539376 %CO/(%PV·min); Kd=2.403571428571428 %CO·min/%PV |
| Ex2 analytic: Ku=4.251212494222509, Tu=7.441522726018012 min | P | Kc=2.1256062471112545 |
| Ex2 analytic | PI ideal | Kc=1.913045622400129; Ti=6.201268938348344 min |
| Ex2 analytic | PI parallel | Kp=1.913045622400129; Ki=0.308492607145281 %CO/(%PV·min) |
| Ex2 analytic | PID ideal | Kc=2.550727496533505; Ti=3.720761363009006 min; Td=0.9301903407522515 min |
| Ex2 analytic | PID parallel | Kp=2.550727496533505; Ki=0.6855391269895132 %CO/(%PV·min); Kd=2.3726620791666386 %CO·min/%PV |
These calculations illustrate the importance of stating controller form beside every number.
The separate synthetic decay-ratio dataset is:
PV(t)=50+8e^(−αt)[cos(ωt)+(α/ω)sin(ωt)]
where α=ln(4)/8 and ω=2π/8. Same-side peaks at t=0, 8, 16 and 24 min are 58, 52, 50.5 and 50.125%. Their amplitudes from the 50% baseline are 8, 2, 0.5 and 0.125 points.
Therefore:
DR = |52−50| / |58−50| = 2/8 = 0.25
That is quarter-amplitude decay.
Now apply artificial PV limits of 48–52%. The detected period remains 8 minutes, but the half-range becomes 2.0 points and peak-to-peak range becomes 4.0 points. This clipped waveform is invalid for ultimate-cycle inference because the amplitude is set by the limits rather than by an unconstrained linear cycle.
Example 2: how dead time changes stability margin
SCRIPTED SYNTHETIC DATA — NOT PLANT DATA
Consider the illustrative process model:
Gp(s) = 2e^(−2s)/(10s+1)
Thus:
- Process gain
K=2. - Time constant
τ=10 min. - Dead time
θ=2 min.
At the ultimate frequency, the process phase is −π radians:
2ω + atan(10ω) = π
A simple bisection check gives:
- At
ω=0.8,2ω+atan(10ω)=1.6+1.4464=3.0464, below π. - At
ω=0.9,2ω+atan(10ω)=1.8+1.4601=3.2601, above π.
Refining the interval gives:
ωu = 0.8443413449792343 rad/min
At that frequency:
- Phase:
−3.1415926535897927 rad ωuτ = 8.443sqrt(1+(ωuτ)^2) ≈ 8.502|Gp(jωu)| = 0.23522700908482508Ku = 1/|Gp(jωu)| = 4.251212494222509 %CO/%PVTu = 2π/ωu = 7.441522726018012 min
These results are exact for the stated mathematical model within the numerical solution shown. They do not include transmitter dynamics, filters, valve behavior, computation delay or process nonlinearity.
A P-only synthetic simulation uses:
e=SP−PV- Time step
dt=0.005 min - A 400-sample first-in-first-out dead-time buffer
- Exact zero-order-hold lag update
- No saturation
- Duration 120 minutes
- Unit SP step
The responses are:
| P-only setting | Kc | Median same-side amplitude ratio | Synthetic result |
|---|---|---|---|
| 0.5Ku | 2.1256 | 0.1046 across 13 peaks | Strong decay |
| 0.9Ku | 3.8261 | 0.7563 across 16 peaks | Decay |
| Ku | 4.2512 | 1.0093 across 15 peaks | Nearly sustained |
| 1.1Ku | 4.6763 | 1.2962 across 15 peaks | Growing |
At Ku, the ratio differs slightly from exactly 1.0 because of discretization drift relative to the continuous analytic model.
At 1.1Ku, the no-saturation model reaches a maximum PV of 55.86, and the PV is still swinging at 6.56 at 120 minutes. The mathematical output is unbounded. A real plant would encounter physical limits and might alarm, trip or be damaged before reproducing this theoretical behavior.
This example explains why ultimate-cycle testing is not an invitation to increase a real controller to Ku. Even a simple model grows rapidly when gain exceeds the boundary.
Example 3: saturation and four anti-windup responses
SCRIPTED SYNTHETIC DATA — NOT PLANT DATA
Consider:
Gp(s)=e^(−s)/(5s+1)
The illustrative ideal PI controller has:
Kc=2.0Ti=4.0 min- Parallel equivalent
Ki=0.5 %CO/(%PV·min) - Output limits
0–100 %CO - Time step
dt=0.02 min
The setpoint schedule is:
- 20% until 5 minutes.
- 120% from 5 to 25 minutes.
- 50% from 25 to 80 minutes.
The 120% setpoint is deliberately unreachable. With process gain 1 and maximum CO of 100%, the PV can only approach approximately 100%.
Four illustrative methods are compared:
- No anti-windup.
- Conditional integration.
- Back-calculation with
Tt=1.0 min. - Instantaneous external-reset tracking using
I(k+1)=u_applied−Kc·e.
The external-reset expression is illustrative and not a vendor algorithm.
| Method | Total saturation duration | Nominal CO re-entry after t=25 | Maximum PV above new 50% SP | ±2% settling after return | Maximum absolute integral state |
|---|---|---|---|---|---|
| None | 30.46 min | 10.46 min | 49.82 points | 25.92 min | 456.70 %CO |
| Conditional integration | 19.74 min | 2.18 min | 48.53 points | 11.78 min | 56.58 %CO |
| Back-calculation, Tt=1.0 min | 22.0 min | 2.0 min | 48.53 points | 10.88 min | 70.58 %CO |
| External reset | 10.32 min | 1.04 min | 48.51 points | 8.44 min | 100.0 %CO |
The “maximum PV above new SP” values are large because the PV is still approximately 98–99% when the SP returns to 50%. This metric describes the process state left by the unreachable high setpoint; it is not a small conventional overshoot.
Without anti-windup, nominal CO reaches approximately 448.7% at 20 minutes and 360.3% at 25 minutes while applied CO remains at 100%. The integral state reaches 456.7% at 25 minutes. Applied CO does not leave its upper limit until approximately 40 minutes.
The ranking is specific to this schedule, model, Tt, limits and algorithm semantics. It does not establish one anti-windup method as universally superior. Most importantly, anti-windup cannot make the unreachable 120% target reachable.
Example 4: a phenomenological stiction model
SCRIPTED SYNTHETIC DATA — NOT PLANT DATA
The illustrative process is:
Gp(s)=e^(−s)/(5s+1)
An ideal PI controller has:
Kc=1.2- Two Ti cases:
1.5 minand3.0 min - Sample time: 1 second, or
1/60 min - SP: 50% initially, then 55% at 5 minutes
- CO limits: 0–100%
The two-parameter stiction model is phenomenological. It is not a friction-force or positioner model.
Its rules are:
- While stuck, if
|c−v|<S, valve travelvdoes not change. - If
|c−v|≥S, breakaway occurs and travel jumps byJin the command direction. - During continuing motion,
v=c−d(S−J), wheredis the direction. - Command reversal returns the valve to the stuck state.
Illustrative parameters are:
- Stiction threshold
S=2%span. - Jump
J=1%span.
The calculated results are:
| Ideal PI setting | Mean late peak period | Late PV half-range | Moving samples | Largest stationary-run CO range |
|---|---|---|---|---|
| Kc=1.2, Ti=1.5 min | 20.4 min | 0.769 points | 2317 | Approximately 3.0 points |
| Kc=1.2, Ti=3.0 min | 35.27 min | 0.636 points | 896 | Approximately 3.0 points |
In this model, doubling Ti increases period by 14.87 min and reduces PV half-range by 0.133 points. This is model-specific. The continuing stick-jump behavior shows that stiction was not repaired.
The synthetic manual-mode case sets constant CO to 50%. Travel remains at 50%, PV settles at 50%, and no autonomous cycle occurs. This simulated result shows only that feedback was necessary for this model’s cycle. It does not distinguish stiction from poor tuning by itself, and the simulation contains no external disturbance or autonomous positioner fault.
An illustrative bidirectional step sequence begins with travel at 50%:
- Command
50→51%: no movement because the change is belowS=2%. - Command
51→53%: breakaway; travel progresses50→51→52%. - Command
53→52%: travel remains at 52%. - Command
52→50%: reverse breakaway; travel moves52→51%and re-sticks.
This direction-dependent pattern resembles deadband or backlash behavior in the synthetic model. Real tests require authorization and may produce different travel paths.
The lesson is not that Ti should be doubled. It is that tuning changes can alter a stiction cycle without removing its mechanical cause.
Corrective-action matrix
| Evidence | Likely category | Verification | Corrective direction |
|---|---|---|---|
| Smooth cycle changes strongly when Kc is reduced | Excessive loop gain | Verify form, action and scaling | Reduce Kc cautiously and reassess robustness |
| Response improves when Ti is increased | Integral too aggressive | Confirm ideal-form Ti versus parallel Ki | Increase Ti cautiously; check disturbance rejection and constraints |
| CO chatters with noisy or quantized PV | Derivative/noise | Inspect raw PV, derivative location and filter | Fix measurement first; then review derivative and filtering |
| Nominal CO exceeds applied CO | Windup | Check limits, selectors, tracking and integral state | Evaluate clamping, back-calculation or external reset |
| CO ramps while travel stays flat, then jumps | Stiction | Inspect travel, pressure, air and mechanics | Repair final-control element |
| Travel is lost after reversal | Backlash/deadband | Use authorized bidirectional test | Inspect linkage, actuator and positioner |
| Cycle concentrates around split point | Split-range discontinuity | Check overlap, gap, characterization, selector and handoff | Correct split-range design or configuration |
| Many loops share a period | Interaction or common forcing | Compare timestamps, spectra and physical paths | Address shared utility, upstream source or interaction |
| Manual mode plus constant travel | External disturbance | Trace feed, pressure, composition, utility and analyzer signals | Correct disturbance source or add justified compensation |
| Manual mode plus moving travel | Positioner, air or mechanical fault | Inspect positioner, supply, actuator and valve | Repair or service final-control element |
| Stability changes with load | Operating-point nonlinearity | Identify dynamics and constraints by region | Review gain scheduling or region-specific strategy |
| Cycle tied to scan/update period | Sampling or computation | Inspect task period, jitter, execution order, communications and filtering | Correct timing, sampling or implementation |
Related topic: control loop performance monitoring.
Safe testing and management of change
Testing should begin with a written question. For example: “Does actual travel remain stationary while applied CO changes?” A narrowly designed test is safer and more informative than changing several tuning parameters and observing whether the trend looks calmer.
A test/change record can use these fields:
| Field | Required entry |
|---|---|
| Loop and equipment | Tag, service, final-control element and affected units |
| Diagnostic question | Competing hypotheses the test will distinguish |
| Baseline | Operating point, product or grade, mode, limits, disturbances and current metrics |
| Original configuration | PID form, Kc/Kp, Ti/Ki, Td/Kd, action, scaling, filters, limits and anti-windup |
| Risk review | Process limits, stored energy, alarms, trips, downstream effects and safeguards |
| Authorization | Named roles permitted to start, pause, abort and restore |
| Test bounds | Maximum amplitude, duration, cycles and recovery time |
| Stop criteria | Numeric or clearly observable conditions |
| Data plan | Signals, sample rate, clocks and historian settings |
| Result | Evidence, calculations and unresolved questions |
| Restoration | Final settings, modes, limits, tracking, anti-windup and backup status |
| Follow-up | Maintenance, engineering review, temporary-change expiry and closeout |
Online tests and tuning changes must be screened under the site’s management-of-change (MOC) and temporary-change procedures. Legal requirements are jurisdiction- and process-specific.
For processes covered by the United States Environmental Protection Agency Risk Management Program, 40 CFR 68.75 requires written procedures for changes other than replacements in kind and consideration of technical basis, safety and health impact, operating-procedure modifications, duration and authorization. OSHA Process Safety Management, 29 CFR 1910.119, applies only to covered highly hazardous chemical processes and includes management of change.
Center for Chemical Process Safety guidance notes that control-system programming or sequence changes may require MOC rather than treatment as replacement in kind, depending on facility procedure. Temporary tuning changes need expiry, restoration and closeout. These United States examples do not mean every tuning change worldwide is legally subject to the same rule. Confirm jurisdiction, regulatory coverage and corporate policy locally.
A PID test can increase demand on alarms, interlocks or safety instrumented functions. Identify shared sensors, trip setpoints and process safety time. Do not alter trip logic, voting, setpoints or bypasses as an incidental part of PID work; those are separate safety-system activities.
Training materials and textbooks may help a team plan a test, but they do not replace competent review, local procedure or operating authority.
Related topic: management of change for control changes.
A planned test/change-record template and MOC/test-plan checklist are proposed: download link: placeholder.
Verification and documentation
A correction is not complete when the trend first looks smoother. Verify the mechanism and the operating result.
Use this checklist:
- Repeat the same period, amplitude, peak and recovery measurements used in the baseline.
- Use comparable operating points and disturbance conditions.
- Verify SP, PV, nominal CO, applied CO and actual travel remain synchronized.
- Confirm the actuator follows applied CO without unexplained stationary periods or jumps.
- Test recovery from relevant constraints without deliberately violating safe limits.
- Check performance after normal load or setpoint disturbances.
- Review representative operating regions rather than only one steady condition.
- Confirm controller form, action, scaling and time units.
- Verify final modes, cascade status, selectors, output limits, rate limits, tracking and anti-windup.
- Confirm filters and sample settings.
- Restore temporary instrumentation, trends and test logic as required.
- Record original and final parameters.
- Record evidence supporting the diagnosed cause.
- Close maintenance, MOC or temporary-change actions.
- Note remaining limitations and conditions not tested.
Do not permanently detune a loop merely to avoid valve maintenance. Do not present a synthetic trace as field evidence. If the correction is valid only within one operating region, document that boundary.
Frequently asked questions
Why do PID loops oscillate?
A PID loop oscillates when gain and phase lag sustain repeated movement, or when a nonlinearity, saturation, interaction, sampling effect or external disturbance forces a cycle. The PID may create the movement, amplify another source or correctly respond to a real periodic disturbance. Trend nominal and applied output, actual travel and disturbances before deciding to retune.
How can I tell valve stiction from poor tuning?
Compare applied CO with actual valve travel. A smooth PV and sawtooth CO can suggest stiction, but it is not proof. Stronger evidence is CO ramping while travel remains flat, followed by a sudden travel jump. Use positioner and supply-pressure information and, when authorized, a bidirectional valve test to confirm the mechanical behavior.
What does it mean if the oscillation continues in manual?
If the PV cycles in manual while actual actuator travel remains constant, suspect an external periodic disturbance. If travel continues moving, inspect the positioner, air supply, actuator and valve mechanics. Confirm that the loop is genuinely in manual. Persistence does not automatically exclude the valve, because a positioner or actuator can cycle independently.
Can too much integral action cause oscillation?
Yes. Integral action adds phase lag and continues accumulating error. In ideal form, a smaller Ti means stronger integral action; in repeats-per-minute form, a larger value is stronger. Excessive integral can delay output reversal and worsen windup during constraints. Verify controller form and units before increasing Ti or reducing parallel-form Ki.
Can derivative action cause or worsen oscillation?
Derivative can amplify noise, quantization and rapid sample-to-sample changes, producing CO chatter. Derivative on error can also react to setpoint steps. Filtering may reduce chatter but adds phase lag, so a poorly chosen implementation can aggravate stability. Inspect raw PV, derivative location, sample rate and filter definition before assuming derivative provides damping.
Why does dead time make a loop difficult to tune?
During dead time, the controller sees no measured evidence that its previous action is affecting the process. It may therefore add more proportional or integral correction before the delayed response arrives. Dead time adds phase lag without advance warning in the PV, reducing the gain and integral aggressiveness that the complete loop can tolerate.
What is integral windup?
Windup is integral accumulation while the requested controller action cannot be delivered because of output limits, selectors, rate limits or actuator capacity. Nominal CO can continue beyond the applied limit, causing delayed recovery when the constraint clears. Compare nominal and applied CO, then evaluate clamping, back-calculation or external-reset tracking for the implementation.
Are ultimate-gain and ultimate-period tests safe?
No active forced-cycle test is inherently safe. An ultimate-gain test deliberately approaches the stability boundary and can grow if Ku is exceeded. Saturation also invalidates linear ultimate-cycle inference. Testing requires a reviewed, bounded plan with limits, stop criteria, authority, monitoring and restoration steps. Manual mode does not immediately remove delayed or stored process effects.
Can a PID oscillation be fixed without retuning?
Yes. Retuning is unnecessary when the root cause is wrong action, scaling, time units, measurement noise, sampling, valve mechanics, limits, interaction or external forcing. Repairing a sensor, restoring actuator obedience, correcting split-range handoff or removing a periodic disturbance can solve the problem while preserving otherwise appropriate controller settings.
How do I verify that the real cause was fixed?
Repeat the same period, amplitude, recovery and constraint metrics under comparable conditions. Verify that actual travel follows applied CO and that recovery from normal limits no longer shows delayed integral release. Check representative operating points, final modes, tracking and filters. Document the original configuration, evidence, approvals, final settings and any untested conditions.
Planned downloads and next steps
The following supporting resources are planned. They are not presented as currently available files:
- Diagnostic decision tree — download link: placeholder
- Trend data-collection sheet — download link: placeholder
- Oscillation period and amplitude worksheet — download link: placeholder
- Test/change-record template — download link: placeholder
- MOC/test-plan checklist — download link: placeholder
- Synthetic CSV datasets and Python script — download link: placeholder
- Corrective-action matrix PDF — download link: placeholder
Start with passive trend collection. Measure the cycle, compare nominal CO with applied CO and actual travel, then identify the earliest credible source.
Primary CTA: Use the planned decision tree and data-collection sheet to classify the oscillation before anyone changes tuning.
Soft workbook CTA: Record each diagnostic step, test boundary, result and final configuration in the planned troubleshooting workbook.
The practical sequence remains:
Recognize → Measure → Classify → Test → Correct → Verify → Document.
References and disclosure
The article is based on the supplied research pack and the following registered sources:
- ISA-TR5.9-2023 Realizing and Achieving Best PID (ISA): https://www.isa.org/intech-home/2023/june-2023/features/isa-tr5-9-2023-realizing-and-achieving-best-pid
- Fundamentals of PID Control (ISA): https://www.isa.org/intech-home/2023/june-2023/features/fundamentals-pid-control
- A Practical Guide to PID Controller Implementation (Sundström, Bauer, Guzmán, Hägglund, Soltesz): https://arxiv.org/pdf/2604.15918
- Anti-Windup Control Using PID Controller Block (MathWorks): https://www.mathworks.com/help/simulink/slref/anti-windup-control-using-a-pid-controller.html
- PID Control (Caltech Feedback Systems wiki): https://www.cds.caltech.edu/~murray/FBS/PID_Control.html
- Different PID Equations (Control.com): https://control.com/textbook/closed-loop-control/different-pid-equations/
- Quantitative PID Tuning Procedures (Control.com): https://control.com/textbook/process-dynamics-and-pid-controller-tuning/quantitative-pid-tuning-procedures/
- Ziegler–Nichols Method (Michigan Tech): https://pages.mtu.edu/~tbco/cm416/zn.html
- Ziegler–Nichols Controller Tuning Example (Colorado School of Mines): https://people.mines.edu/jjechura/wp-content/uploads/sites/120/2019/02/CHEN403_14_ZieglerNicholsExample.pdf
- Modern PID Control (Bhattacharyya/UC Berkeley): https://msc.berkeley.edu/assets/files/PID/modernPID1-overview.pdf
- Industrial Process Models (Iqbal/LibreTexts): https://eng.libretexts.org/Bookshelves/Industrial_and_Systems_Engineering/Introduction_to_Control_Systems_(Iqbal)/01%3A_Mathematical_Models_of_Physical_Systems/1.05%3A_Industrial_Process_Models
- Stiction: Definition and Discussions (Choudhury, Shah, Thornhill; Springer): https://link.springer.com/content/pdf/10.1007/978-3-540-79224-6_11.pdf
- Benchmarking Control Loops with Oscillations and Stiction (Horch; Springer): https://link.springer.com/content/pdf/10.1007/978-1-84628-624-7_7.pdf
- Managing the Performance of Control Loops with Valve Stiction (Patwardhan et al.): https://www.researchgate.net/profile/Rohit-Patwardhan/publication/284446000_Managing_the_performance_of_control_loops_with_valve_stiction_An_Industrial_Perspective/
- Analysis and Compensation of Control Valve Stiction-Induced Limit Cycles (Fang, Wang, Tan, Shang; IEEE): https://ieeexplore.ieee.org/document/7577013
- Closed-loop Control Troubleshooting (Smuts, ISA): https://www.isa.org/intech-home/2018/march-april/departments/closed-loop-control-troubleshooting
- Control Valves and Cycling Control Loops (Monsen, Valve World): https://valve-world.net/control-valves-and-cycling-control-loops/
- CCPS Key Principles of Process Safety: Management of Change: https://www.aiche.org/sites/default/files/html/ccps/3028311/V1/files/downloads/MOC%20final%20Jan%202024.pdf
- 40 CFR 68.75: https://www.ecfr.gov/current/title-40/chapter-I/subchapter-C/part-68/subpart-D/section-68.75
- 29 CFR 1910.119: https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XVII/part-1910/subpart-H/section-1910.119
The original 1942 Ziegler–Nichols paper, Optimum Settings for Automatic Controllers, was not directly inspected by the editorial team. Its bibliographic details and coefficients are therefore treated as partially verified through technical reproductions. Paid standards are not reproduced.
All numerical examples and worked-example figures are based on scripted synthetic data, not plant, field or laboratory results.
Disclosure: This page may display third-party advertising. No affiliate links or product endorsements are included in this article. Textbooks and training resources mentioned are educational aids and do not guarantee that a control-loop problem will be fixed.
This article is educational and is not a substitute for site operating procedures, hazard review, management of change, competent engineering judgment or local regulatory requirements.
Comments
Post a Comment