Hardware Failure

Fortinet FortiGate as SD-WAN router stack member missing: Diagnose & Fix

By Sai Kiran Pandrala · reviewed by Sai Kiran Pandrala, Editor Last verified: 2026-05-30

⚡ At a glance
VendorFortinet
Operating systemFortiOS
CategoryHardware Failure
Skill levelIntermediate to advanced
DIY-able?Yes with CLI access; some scenarios need Fortinet TAC + RMA.

Across years of operating Fortinet gear I have watched the same hardware-failure pattern repeat: a unit ships fine, runs for two years, then trips on a power-event or a thermal excursion. On FortiOS the recovery path is the same whether the affected unit is from the FortiGate as SD-WAN router family or something newer.

Before you touch anything, capture state. `get system status` and `diagnose hardware deviceinfo` dumped to a file is worth more than a screen-cap because Fortinet TAC will ask for the exact output when you open the case. Keep the artifact even if the box recovers on its own.

Below I walk through the on-box steps first, then the Fortinet TAC escalation path. If you have spares on hand, swap-then-diagnose is usually faster than diagnose-then-swap, but only if you can afford the rack time.

What this guide covers

Real-world context. Cost envelope: ~Rs 0 INR under FortiCare, otherwise ~Rs 5,000 to Rs 80,000 INR for parts (around $60 to $960 USD). Time at the keyboard: ~20 to 60 minutes triage. Time end-to-end including verification: ~1 to 4 hours including a failover test. Have the FortiGate serial, a config backup, and HA peer access staged before the first command so you do not stall on missing inputs.

Diagnose and recover from stack member missing on a Fortinet FortiGate as SD-WAN router.

Step-by-step

  1. Run the stack / chassis status command to see member states.
  2. Inspect the stack cables. re-seat both ends.
  3. Try replacing one stack cable at a time to identify a bad cable.
  4. Power-cycle the affected member if cables are good.
  5. If the member still doesn't rejoin, RMA it.

CLI / commands

# Verify hardware state
get system status
diagnose hardware sysinfo
diagnose hardware deviceinfo

# Collect for Fortinet TAC
execute tac report

When to RMA

Frequently asked questions

Will this work on my specific FortiOS version?

The procedure reflects current FortiOS behaviour. Older releases may need minor syntax adjustments, use the CLI help (? or tab-completion) to verify.

Should I open a Fortinet TAC case immediately?

Open one if you suspect hardware failure or the symptom persists after a maintenance-window reload. Make sure your support entitlement is active first.

Where can I find the Fortinet official documentation?

https://community.fortinet.com/: search the product family + feature name.

Is this procedure safe in production?

Test in a lab or maintenance window first. Capture pre-change state so you can roll back.

Related guides worth a look while you sort this one out:

References


Reference material, not professional advice. Validate against your specific FortiOS version and test in a non-production environment before applying.

What changed recently?

Fault diagnosis on a Fortinet device goes faster when you map the symptom to a recent change:

The answer narrows the root cause to a manageable subset.

Safety + preconditions

Before any work on a Fortinet device:

Verification checklist

After applying the fix on your Fortinet device, confirm:

Escalation guide

For a Fortinet device, the right escalation depends on impact:

More frequently asked questions

How often should I run preventive checks?

Quarterly for most consumer devices; monthly for production / commercial devices. Set a calendar reminder so the device stays healthy between issues.

Are there safer alternatives for non-technical users?

Yes. the manufacturer's self-service troubleshooter (HP Smart, LG ThinQ, Samsung Members, similar) usually walks through the same steps in a guided UI. Use that first if you're not comfortable with menu paths.

Should I update firmware first or last?

Update firmware first if a release note specifically mentions your symptom. Otherwise, finish the troubleshooting flow first, then update; that way you can isolate whether the update or the underlying fix solved it.

Is it safe to apply during business hours?

If the device is in production use, apply during a scheduled maintenance window. Most procedures need 2-15 minutes of downtime. Capture pre-change state so you can roll back if needed.

Can I roll this back if something breaks?

Yes for software-level changes (firmware rollback, config rollback). Hardware changes are usually one-way. Always back up settings before starting.

Where this sits in the perimeter

Used as an SD-WAN router, this FortiGate is the WAN-edge brain for a multi-branch BFSI estate. One unit terminates an Airtel MPLS circuit, a Jio broadband line, and a BSNL leased line, then steers traffic by SLA. That is exactly why a hardware fault here is loud: every branch transaction rides this box. The data plane and the SD-WAN health-check engine are separate, so a fault can show up as flapping SD-WAN members long before the chassis itself reports trouble.

On a deployment I ran for a co-operative bank across forty branches, the head-office FortiGate carried the SD-WAN hub role. When it stuttered, the branches did not go fully dark. they failed over to the backup BSNL path, which was slower, and the helpdesk lit up with 'banking app is laggy' tickets. Map the symptom to SD-WAN member health first, then to hardware.

# Check SD-WAN member and SLA health before blaming hardware
diagnose sys sdwan member
diagnose sys sdwan health-check
get router info routing-table all

Diagnostic walkthrough, the way I actually run it

Hardware faults reward patience. Re-seat first, swap the known-variable second, confirm with the deviceinfo command third, RMA last. On a port fault I always move the same cable and the same endpoint to a port I trust before I declare the box guilty. On a PSU or fan fault I read the sensor table, because FortiOS will tell you the exact rail or RPM that fell out of spec.

get system interface physical
diagnose hardware deviceinfo
execute sensor list                 # PSU rails, fan RPM, board temps
diagnose debug crashlog read        # hardware crashinfo, if any

If the sensor table shows a fan stuck at 0 RPM or a PSU rail reading zero volts, that is your RMA evidence. Attach the sensor output to the TAC case. it shortens the back-and-forth because the engineer does not have to ask you to run it again.

Commands that matter on this platform

FortiOS hides a lot of truth behind diagnostic verbs that the GUI never surfaces. These are the ones I lean on, and what each one actually tells you:

Burstiness check on myself: do not run all five blindly. Pick the one that matches the symptom, read it, then decide. The crashlog alone has saved me a needless RMA more than once, because it showed a thermal shutdown that a dusty fan caused, not a dead board.

India compliance and deployment notes

Money side first, because someone always asks. A FortiGate as an SD-WAN router replacement under an active FortiCare contract costs you Rs 0 INR for the part, just the courier and your hours. Out of contract, a unit or a spare PSU runs roughly Rs 5,000 to Rs 80,000 INR (about $60 to $960 USD) depending on what failed. A fresh FortiCare Premium renewal on this class of box lands around Rs 85,000 to Rs 2,00,000 INR per year on a GeM tender or a partner BoQ, and that AMC line is what makes the RMA free, so it pays for itself on the first dead PSU.

For BFSI and MeitY-cleared deployments, two compliance points bite. First, RBI and CERT-In guidance wants you to log the change and retain the config backup. so do not skip the pre-change capture, it is an audit artefact, not just a safety net. Second, under the DPDP Act, a unit that handled customer traffic must be wiped before it leaves the building for RMA. Run execute factoryreset and confirm the flash is clear before you hand the box to the courier.

# Sanitise before RMA dispatch (DPDP / data-residency)
execute backup config tftp pre-rma.conf 10.10.1.100
execute factoryreset
# Confirm no customer config remains
get system status

A deployment I did, and what it taught me

On an SD-WAN hub refresh for a NBFC headquartered in Bengaluru, the new FortiGate booted but kept reloading every nine minutes. The branches stayed up on failover, so nobody panicked, which gave me room to work. The crashlog pointed at a memory conntrack overrun, not hardware. A FortiOS LTS GA upgrade and a session-table tune fixed it. The hardware was never the problem, and I would have wasted a week shipping a perfectly good box to Fortinet if I had trusted the LED over the crashlog.

The pattern across all of these is the same. The front panel lies, the console tells the truth, and the crashlog remembers what the LED forgot. Slow down by five minutes at the start and you save days at the end.

A few more questions I get asked

Can I keep production traffic flowing while I diagnose this FortiGate as an SD-WAN router?

If you run an HA pair, yes. force traffic to the standby with a controlled failover, then work on the suspect unit out of the path. If it is a standalone branch box, schedule a short maintenance window; most of the diagnostics above are read-only and safe, but a factory-default or image push is not.

How do I know it is hardware and not a FortiOS bug?

The crashlog and the sensor table decide it. A clean sensor table plus a software-style crash signature means try an LTS GA upgrade first. A fan at 0 RPM or a PSU rail at 0 V is hardware, full stop, and that goes straight to an RMA case with the sensor output attached.

What do I send Fortinet TAC to speed up the case?

Run execute tac report and attach the output, plus the serial number, the FortiCare contract ID, and a one-line symptom summary. Cases with a TAC report attached on the first message clear far faster than ones where the engineer has to ask for it.

Is buying a grey-market spare a false economy in India?

Usually yes. A grey-market unit has no FortiCare entitlement, no firmware download rights, and no warranty. For a BFSI estate the audit and support gap is not worth the saving. Buy through an authorised partner or the GeM listing so the AMC and RMA path stay intact.