Fan Speed Issues

Dev Account
Dev Account
  • Updated

Fan Speed Issues

Summary

Fan speed issues cover fans running too fast, fluctuating, not spinning, reporting wrong telemetry/GPU fan ERR, or making abnormal noise. Causes include BMC/BIOS control, failed fan modules, chassis/fan-board paths, seating/contact issues, thermal-control mismatch, false sensor telemetry, and normal high-airflow behavior.

Frequency

  • 194 tickets mention fan speed, fan noise, fan telemetry, or fan-control faults.

Common Causes

  • BMC/BIOS fan-control or telemetry faults. Fans ran high, oscillated, reacted backwards to temperature, appeared mislabeled, reported 0 RPM while physically spinning, or reacted to false thermal sensors such as ExpanderA 235°C (#22547, #31541, #41986, #42596, #44859, …and 60+ more).
  • Single failed or noisy fan module. Cases narrowed to one chassis, CPU-adjacent, PSU, GPU, or liquid-cooler radiator fan with grinding, humming, non-spin, repeated warnings, or fault LEDs (#21703, #32480, #38267, #41804, #44958, …and 50+ more).
  • Chassis, fan-board, harness, or seating path faults. Fan swaps/reseats, GPU reseats, cable checks, chassis/fan-board repair, or depot work were needed when the control path did not clear (#31541, #32471, #34438, #35680, #44357, …and more).
  • Thermal-control mismatch after service/configuration change. Fan behavior changed after RMA, firmware, platform updates, OS state, or thermal policy changes while the system otherwise ran (#24162, #24164, #30085, #40557, #40811).
  • Expected, load-dependent, unconfirmed, or vendor-driver scoped fan behavior. Some servers were loud by design, used separate CPU/GPU fan zones, scaled normally with load, closed before root cause confirmation, or involved NVIDIA driver-stack fan-control permissions rather than Exxact hardware; one 4x RTX PRO 6000 Blackwell Max-Q TS4 server was expected around 11,000 RPM idle and nearly double under load (#11548, #17481, #23091, #37709, #44674, #45551).

Diagnostic Steps

  • Classify the symptom. Separate constant high RPM, oscillation/reversed response, one noisy/non-spinning fan, PSU fan faults, GPU fan ERR, telemetry-only alarms, and loud-but-normal airflow (#21703, #22547, #31541, #44958).
  • Check management evidence. Review BMC/IPMI readings, fan mode/profile, BIOS/BMC versions, SEL logs, sensor-to-temperature consistency, physical labels, false CPU/expander temperatures, whether IPMI Power Saving mode is available, and whether the symptom is an NVIDIA/Xorg driver permission regression rather than chassis fan hardware (#24164, #39977, #42596, #44674, #44859, #45551).
  • Isolate hardware paths. Swap/reseat suspect fans, GPUs, cables, and modules; inspect fan boards/chassis harnesses, dust, obstruction, connector seating, and PSU fan/fault LEDs before assuming firmware or chassis failure (#32471, #34438, #38267, #43894, #44357).
  • Reproduce under controlled conditions. Compare idle/load and temperature response when inverted fan curves, unstable RPM, or post-repair recurrence is suspected (#31541, #36410, #40557).

Solutions

  • Replace failed fan/module. Use for confirmed noise, non-spin, fan alerts, liquid-cooler radiator fan failure, or PSU-internal fan failure treated as a PSU module issue (#21703, #32480, #38267, #41804, #44958, …and 50+ more).
  • Update, reset, or tune BMC/BIOS fan control. Firmware/settings remediation, BMC reset, fan-mode changes, or telemetry correction can resolve false readings, unstable curves, and control faults, though reset may be temporary when a sensor defect persists (#22547, #31541, #37005, #41268, #44859).
  • Repair chassis-side control hardware. Use fan-board/chassis repair, chassis replacement, or depot repair when fan swaps and reseats do not clear the fault path (#31541, #32471, #34438, #35680, #40699).
  • Validate before return. Burn-in and thermal checks confirm corrected fan behavior after reproduction, repair, or chassis-side work (#22547, #31541, #32471, #36410).
  • Clarify expected behavior or monitor after reseat. Explain normal server acoustics, separate fan banks, and IPMI Power Saving limits; monitor intermittent GPU fan ERR after GPU reseat/swap when symptoms clear, route NVIDIA driver-stack fan-control permission regressions to NVIDIA/vendor channels, and avoid unvalidated power-management changes such as ASPM because GPUs may fall off the bus (#11548, #42596, #44357, #44674, #45551).

Edge Cases

  • Repeat post-RMA fan behavior can recur after prior repair (#24162, #24164, #40699).
  • Reversed control logic can make fans speed up as temperatures drop (#31541, #36410).
  • Fan symptoms can co-occur with overheating, GPU instability, software corruption, boot failure, or no-POST; one loud-fan/high-temperature case recovered after OS reinstall, and one no-POST case had fans recover after DIMM reseat/CMOS while platform RMA was still needed (#30085, #40557, #41684, #43116, #43675).
  • IPMI/label ambiguity, 0-RPM readings with physically spinning fans, false CPU temperature, false ExpanderA 235°C, or NVIDIA driver fan-control permission changes can indicate telemetry/firmware/software behavior rather than a failed fan (#39977, #41986, #42596, #44859, #45551).
  • Simple fan/acoustic cases can still stall on part ID, shipment timing, follow-up gaps, return logistics, or clarifying whether only the fan or full workstation returns (#21703, #32480, #41804, #43535, #43894, #44453).

Related Issues

Was this article helpful?

0 out of 0 found this helpful

Have more questions? Submit a request

Comments

0 comments

Please sign in to leave a comment.