Deployment Automation

Nvidia (Mellanox) SN2010: How to deploy with the vendor's controller / manager

By Sai Kiran Pandrala · reviewed by Sai Kiran Pandrala, Editor Last verified: 2026-05-30

⚡ At a glance
VendorNvidia (Mellanox)
Operating systemCumulus Linux / NVOS / SONiC
CategoryDeployment Automation
Skill levelIntermediate to advanced
DIY-able?Yes with CLI access; some scenarios need Nvidia Enterprise Support + RMA.

Bulk operations on Nvidia (Mellanox) get expensive fast if you do them serially. Cumulus Linux / NVOS / SONiC tolerates parallel pushes well, but only if you respect the rate limits and the activation order on the SN2010 family. Get either wrong and you create a self-inflicted outage.

I always wrap automation runs in pre- and post-check captures. cl-support (Cumulus) / show techsupport (SONiC) before and after gives you a diff that Nvidia Enterprise Support can act on if anything looks off later.

What follows is a battle-tested pattern. Adapt the concurrency to your environment, there is no universal right answer, only ranges that work.

What this guide covers

How to deploy with the vendor's controller / manager for Nvidia (Mellanox) SN2010 (Cumulus Linux / NVOS / SONiC).

Step-by-step

  1. Choose the automation surface: vendor controller, API, or CLI scripting.
  2. Verify reachability + credentials from your automation host.
  3. Test the change on a single device + maintenance window.
  4. Roll out in waves of 10-20 devices to limit blast radius.
  5. Pre-collect baseline, push the change, post-collect; diff.
  6. Roll back any device whose post-check fails.

Sample CLI invocation

# Manual baseline
nv show system
nv show platform inventory
nv show interface

# Push change (via vendor CLI)
nv config (NVUE)
nv set interface swp1 ip address 10.0.0.1/24
nv config apply
nv config save

# Verify
nv show interface

Best practices

Frequently asked questions

Will this work on my specific Cumulus Linux / NVOS / SONiC version?

The procedure reflects current Cumulus Linux / NVOS / SONiC behaviour. Older releases may need minor syntax adjustments. use the CLI help (? or tab-completion) to verify.

Should I open a Nvidia Enterprise Support case immediately?

Open one if you suspect hardware failure or the symptom persists after a maintenance-window reload. Make sure your support entitlement is active first.

Where can I find the Nvidia (Mellanox) official documentation?

https://docs.nvidia.com/networking/, search the product family + feature name.

Is this procedure safe in production?

Test in a lab or maintenance window first. Capture pre-change state so you can roll back.

Related guides worth a look while you sort this one out:

References


Reference material, not professional advice. Validate against your specific Cumulus Linux / NVOS / SONiC version and test in a non-production environment before applying.

What changed recently?

Fault diagnosis on a Nvidia device goes faster when you map the symptom to a recent change:

The answer narrows the root cause to a manageable subset.

Before you start

A few things to confirm so the Nvidia device fix goes cleanly:

How to confirm it's actually fixed

On a Nvidia device, the test is rarely "reboot and see". Use this list:

Escalation guide

For a Nvidia device, the right escalation depends on impact:

More frequently asked questions

Can I roll this back if something breaks?

Yes for software-level changes (firmware rollback, config rollback). Hardware changes are usually one-way. Always back up settings before starting.

Why is this happening on a brand-new unit?

Out-of-box defects do occur. If you've owned the device under 30 days and the symptom persists after a factory reset, escalate to the seller for replacement under DOA terms before opening a manufacturer support case.

Does this affect other devices on my network?

Generally no. The procedure is local to this device. Network-side changes (firmware updates that affect TLS, SMB, or routing) are flagged explicitly in the steps.

Is it safe to apply during business hours?

If the device is in production use, apply during a scheduled maintenance window. Most procedures need 2-15 minutes of downtime. Capture pre-change state so you can roll back if needed.

What if my model isn't exactly the same revision?

Cross-check the model code on the rating plate against the manufacturer support page. Major firmware generations sometimes shift the menu path; the option is usually under a similarly-named section.