How to decode Xid 13/31/63/79/119/120 errors
| Hardware family | NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) |
|---|---|
| Category | Computer Hardware |
| Guide type | Procedure |
| Skill level | Intermediate to advanced |
| Time | 15 - 60 minutes including verification |
When How to decode Xid 13/31/63/79/119/120 errors bites you on NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200), the first instinct is to open an RMA ticket. Most of the time you do not have to. The steps below are the ones a senior hardware tech would walk you through at a repair bench.
What how to decode xid 13/31/63/79/119/120 errors actually involves on NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)
This task on NVIDIA Datacenter H100 H200 B200 is one of the more searched operational topics across vendor forums and Tom's Hardware in the last 12 months. The procedure below is the path that works on a current NVIDIA Datacenter H100 H200 B200 setup with default config.
The rest of this page is the structured fix path. Start with diagnose, then remediation, then the automation options so you do not have to do this by hand the next time it surfaces. Verify and safety sections at the end are the discipline that keeps the fix from regressing in production.
Diagnose first, fix second
Second pass: run the vendor pre-boot diagnostic before Windows loads, because Windows will hide half the truth on a NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200). Dell: tap F12 at the splash, pick Diagnostics, let SupportAssist ePSA run the full memory and storage suite (write down the validation code if a test fails). HP: tap F2 for UEFI Diagnostics, run System Fast Test and then the Memory Extensive plus Storage Extensive (it will spit a failure ID like 6XL1KP-8RV85F-MFPV6J-60SB03). Lenovo: tap F10 or boot Lenovo Vantage Hardware Scan. Apple Silicon: hold power, choose Options, then Diagnostics; Intel Macs: power-on with D held. MSI laptops: open MSI Center, System Diagnosis tab. ASUS desktops have MyASUS Diagnostics under Customer Support, which mirrors the ePSA-style suite for ROG and ProArt boards. On AMD AM5 boards, the DRAM Q-LED flashing at boot followed by no display is almost always EXPO instability after the AGESA 1.2.0.3C update; drop EXPO, boot bare, and confirm Memory Context Restore is on before retraining. Capture the raw result codes in a notebook with timestamps - the second time the same SKU fails the same test, you have a fleet pattern, not an incident, and the RMA narrative writes itself.
Start by capturing the exact failure signal in writing before you change a single thing on your NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) rig. On a desktop board that is the 2-digit Q-Code hex on the postcode display (top-right corner on Asus ROG, lower edge on MSI MEG, near the 24-pin on Gigabyte Aorus, by the chipset on ASRock Dr.Debug). On a laptop it is the amber power-LED blink pattern (Dell does 2+1 for CPU, HP does 3+5 for memory, Lenovo flashes the Caps-Lock LED). On Windows it is the BSOD stop code at the bottom of the blue screen and the bugcheck hex in parentheses. Photograph it. Do not paraphrase. On Dell add the SupportAssist Pre-Boot Assessment validation code that ePSA prints at the end of any failed test - it is the only string Dell ProSupport will accept on the RMA portal at support.dell.com when you open the case. On HP capture the Failure ID in the 6XL1KP-8RV85F-MFPV6J-60SB03 format from HP UEFI Diagnostics; the Care Pack agent decodes it directly into a part recommendation. On Lenovo include the Hardware Scan QM3WMRP-style result code from Lenovo Vantage, which the Premier Support workflow uses to skip the Tier 1 triage queue.
Sixth: pin down the thermal envelope on the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) under real load. Launch HWiNFO64 in Sensors-only mode, hit the clock icon to log to CSV, then run a known workload: Cinebench R23 30-minute loop for sustained CPU, Unigine Superposition 4K Optimized for sustained GPU, FurMark only briefly and with very cautious use (it pushes PL2 / TBP past spec and can melt under-rated 12V-2x6 connectors in minutes). Watch CPU Package, Tctl/Tdie, VR VOUT, VR T-Junction, GPU Hot Spot, and GDDR6X memory junction. Confirm the AIO pump is plugged into CPU_FAN or AIO_PUMP at 100 percent, not CPU_OPT, or BIOS will throw CPU FAN ERROR while the pump silently sits at 0 RPM. Run OCCT CPU+Cache (Large data set, AVX2) for 30 minutes to provoke IMC errors that Cinebench will not, then OCCT Power for combined CPU plus GPU draw to test PSU transient response. Use HWiNFO64 8.x with the latest sensor patches because older builds mis-read AMD VSOC on AGESA 1.2.0.3C; if Tctl exceeds 95C at stock the cooler mount pressure is wrong, repaste with PTM7950 phase-change pad (refrigerated 1 hour pre-application, cut to die size) and fit a Thermalright contact frame on LGA1700.
Solution-focused remediation path
For NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) systems where storage is suspect, open CrystalDiskInfo 9.x and read SMART honestly. Any Reallocated Sector Count climbing past the threshold, or Pending Sector above zero, means back up tonight and start the RMA. For Gen5 NVMe drives sitting at 86C under sustained writes, fit the heatsink that came in the motherboard box and add direct airflow; throttling explains most "slow" reports. Samsung 990 Pro showing the 0E health drop needs the firmware update via Samsung Magician immediately, not next week. DRAM-less SSD stalls during big writes are by design: accept the budget or swap to a DRAM-cached drive. Run CrystalDiskMark 8.x at the 64 GiB size and compare against the manufacturer spec sheet - sustained writes under 300 MB/s on a Gen4 drive rated 5000+ MB/s indicate cache exhaustion or thermal throttling, not a dead controller. Decision point: in-warranty NVMe with documented SMART degradation goes to the SSD vendor RMA portal (Samsung Members, WD support, Crucial RMA), out-of-warranty drives with bulk data go to a recovery shop (Ontrack, DriveSavers) only if the data is worth four-figure recovery fees; otherwise restore from the Macrium Reflect image and swap to a fresh DRAM-cached drive with at least 5 year warranty (Samsung 990 Pro, WD SN850X, Crucial T705).
When the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) fault tracks to display, hangs, or TDR events in Reliability Monitor, treat the GPU stack as suspect. Boot Safe Mode, run DDU 18.x to fully strip NVIDIA, AMD, and Intel display drivers, then reinstall via NVIDIA App, AMD Adrenalin, or Intel Graphics Software (clean install ticked). Verify Resizable BAR is actually active in GPU-Z: that needs Above 4G Decoding on, Re-Size BAR Support on, CSM off. On RTX 4090 and 5090 reseat the 12V-2x6 firmly until you hear the audible click, keep bend radius at or above 35 mm of straight cable before the curve, prefer a native ATX 3.1 PSU cable over the included adapter, and never daisy-chain 12V-2x6. Decision point: if TDR persists after a clean DDU reinstall on the previous driver branch (not the latest), photograph the connector seated and the GPU PCB next to the I/O bracket and open NVIDIA RMA or AMD RMA via the support portal; on prebuilts the path is OEM RMA first (Dell ProSupport, HP Care Pack) because cracking the chassis voids coverage. Part-number convention: Dell DPN starts with 0 (for example 0WTRP4), HP part numbers start with letters (L29483-001), Lenovo uses FRU PN (5M11A12345), ASUS uses a P/N like 90YV0HP0-M0NA00; record the correct format in the RMA narrative or the ticket bounces back.
When the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) box reboots under load, posts black, or coil-whines under any draw, suspect the PSU before blaming the GPU. Put a multimeter on the 12V rail under load and confirm it holds at or above 11.4V (idle should sit at 11.95 to 12.10V, transient sag below 11.4V on Cinebench R23 ramps means the PSU is on the edge). Listen carefully: coil whine is a high pitch from chokes, fan motor noise is lower and steady, an audible click from the PSU is replace-immediately territory. Substitute a known-good unit if you have one on the shelf. RTX 5090 connector melting cases almost universally trace back to non-ATX-3.1 supplies, adapters, or daisy-chained 12V-2x6: use an ATX 3.1 PSU rated 1000W minimum with a native 12V-2x6, and reseat the connector after every chassis move with the audible click and at least 35 mm of straight cable before any bend. Decision point: if multimeter readings on the EPS 12V dip below 11.4V under OCCT Power, RMA the PSU via the vendor portal (Corsair, Seasonic, be quiet!, EVGA each have direct portals) before touching the GPU; on prebuilts the Dell ProSupport or HP Care Pack RMA replaces the PSU under warranty without a separate part order. Photograph the 12V-2x6 connector seated, latch click visible, before relocating the system - that photo is what the RMA team asks for first on any melted connector claim.
Automate this fix so you do not do it twice
Automate vendor diagnostic and SMART pull via vendor CLI
On the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200), regular SMART snapshots catch reallocated sectors, pending sectors, and NVMe Media and Data Integrity Errors well before the drive disappears mid-boot. Pair smartctl long self-tests with the OEM diagnostic CLI (Dell SupportAssist, HP Image Assistant, Lenovo Vantage) so both controller-side and OS-side issues land in one folder. The vendor installers all support silent install via /SILENT or /VERYSILENT flags - dcu-cli.exe installs unattended with /SILENT /NORESTART, HP Image Assistant ships as a self-extracting EXE with /S, and Lenovo Thin Installer accepts /VERYSILENT for the bootstrap before the actual /CM scan. Run the scheduled task under Windows PowerShell 5.1 for broadest compatibility; if you have standardized on PowerShell 7.x, the script-block syntax below works without change. Pipe the JSON output through ConvertFrom-Json for downstream parsing into the fleet dashboard.
$smartctl = "C:\Program Files\smartmontools\bin\smartctl.exe"
$out = "C:\Logs\NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)-smart-$(Get-Date -Format yyyyMMdd).txt"
& $smartctl --info --health -a /dev/nvme0 | Out-File $out
& $smartctl -t long /dev/nvme0 | Out-File $out -Append
# Dell unattended scan (silent, log to file)
& "C:\Program Files (x86)\Dell\CommandUpdate\dcu-cli.exe" /scan -outputLog="C:\Logs\dcu-NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200).log"
# HP Image Assistant unattended
& "C:\HPIA\HPImageAssistant.exe" /Operation:Analyze /Silent /ReportFolder:"C:\Logs\HPIA-NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)"
# Lenovo Thin Installer silent bootstrap then scan
& "C:\Lenovo\ThinInstaller\ThinInstaller.exe" /VERYSILENT
& "C:\Lenovo\ThinInstaller\ThinInstaller.exe" /CM -search A -action SCAN -noiconMonitor and alert via HWiNFO64 logging + Performance Counters
For the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200), the most useful long-running telemetry is HWiNFO64 8.x sensor logging to CSV (CPU package temp, VRM temp, GPU hotspot, GPU memory junction, SSD composite) sampled every 2 seconds, plus Windows Performance Counters for GPU engine and memory usage. Argus Monitor adds SMART-over-time; a homelab Grafana is optional but pays off past a handful of machines. Register the Get-Counter sampler via Task Scheduler XML (schtasks /create /XML) so the task definition is identical across the fleet and survives image redeploys. The Get-Counter pattern below runs identically on Windows PowerShell 5.1 and PowerShell 7.x; if you push the CSV to a central collector, wecutil event forwarding on the source nodes carries the WHEA correlation events to the same dashboard so thermal events and machine checks line up on one timeline.
# HWiNFO64 INI (excerpt) - place next to HWiNFO64.exe
# SensorsOnly=1
# OpenSensors=1
# MinimizeMainWnd=1
# MinimizeSensors=0
# Logging.Enabled=1
# Logging.File=C:\Logs\NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)-hwinfo.csv
# Logging.Interval=2000 # PowerShell: sample GPU engine + memory counters every 5s for 1h
Get-Counter -Counter "\GPU Engine(*engtype_3D)\Utilization Percentage",` "\GPU Process Memory(*)\Local Usage" ` -SampleInterval 5 -MaxSamples 720 | Export-Counter -Path "C:\Logs\NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)-gpu.blg" -Force
# Register via schtasks XML for reproducibility across the fleet
# schtasks /create /TN "NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)-gpu-sample" /XML C:\Tasks\gpu-sample.xml /RU SYSTEMCodify the BIOS fix as a saved profile and backup USB
Once a stable BIOS revision is identified for the NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200), save it as a named profile in the UEFI (slot 1 through 8, with date and AGESA or microcode tag in the name) and prepare a recovery USB. ASUS BIOS Flashback needs a specific filename produced by the BIOSRenamer utility, and Gigabyte Q-Flash Plus expects GIGABYTE.bin on a FAT32 USB in the white-rimmed port. PowerShell makes the rename reproducible across rebuilds. The snippet below targets Windows PowerShell 5.1 syntax so it runs on stock Windows 10 / 11 without PowerShell 7 installed; if you standardize on pwsh 7.x for the fleet, the same Copy-Item and Get-ChildItem calls work identically. Stage the recovery USB next to a printed label (system serial, BIOS rev, AGESA, date) and store in a labeled drawer; the second time a board bricks at 2 a.m. you do not want to be rebuilding the stick from scratch.
$src = "C:\BIOS\NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)\X670E-HERO-ASUS-2401.CAP"
$dst = "E:\X670E.CAP" # name from BIOSRenamer
Copy-Item $src $dst -Force
# Gigabyte Q-Flash Plus expects GIGABYTE.bin at root
Copy-Item "C:\BIOS\NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200)\B650-AORUS-F36.bin" "E:\GIGABYTE.bin" -Force
Get-ChildItem E:\ | Format-Table Name,Length,LastWriteTime
# Label profile in UEFI as: 2026-05-31_AGESA_1.2.0.3C_stable
Common pitfalls and what to watch for
Read-only validation before any write is the single step most NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) repairs skip, and it is the step that lets you roll back when a fix backfires. Photograph every existing UEFI page (Auto versus manual values matter), capture the current Q-Code or MSI EZ Debug LED state on a phone video, export SMART data through CrystalDiskInfo to PDF, and photograph cable routing including the 12V-2x6 seating angle before any disconnect. On NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) AM5 and X3D platforms record VSOC voltage in HWiNFO64 before toggling EXPO, because a VSOC above 1.30V on early AGESA is the documented degradation path. On RTX 4090 and 5090 builds photograph the 12VHPWR or 12V-2x6 connector fully seated with the latch click visible before relocating the system.
The mirror-image mistake is confusing a user-error symptom with a hardware fault on NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200). No display on an RTX 50 card is often a DisplayPort 2.1 cable rated only for UHBR 10 rather than a dead panel. A WHEA Uncorrectable Machine Check might be a degraded Intel 13th or 14th gen die that still needs the 0x12B microcode rather than bad DIMMs. Plugged in, not charging on a 140W gaming laptop is frequently a 65W USB-C PD brick negotiating an underpowered profile, not battery EOL.
Verify the fix worked
- Reproduce the original symptom path on NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200). If it still surfaces on any unit in the fleet, you have not fixed it.
- Watch for 24 to 48 hours via Windows Reliability Monitor + Event Viewer (Windows Logs > System filtered to Error) + HWiNFO64 sensor log. Cached health masks slow-burn thermal drift and memory bit-rot.
- Smoke-test under realistic load: Cinebench R23 30-min for CPU, Unigine Superposition for GPU, CrystalDiskMark for storage, MemTest86 1 pass for RAM.
- Capture the new state in a runbook so the next person on call does not rediscover this. Note BIOS version + microcode revision + driver branch + Q-Code seen + verbatim error string + fix applied. Push to a shared wiki.
- If the fix involved a BIOS change, save the working BIOS to a USB labeled with the system serial, and screenshot every BIOS page for archival.
Safety, rollback, blast radius
- Test on a non-production rig or back up via Macrium or Clonezilla before any write that touches NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200).
- Anti-static wrist strap clipped to bare chassis metal. Non-carpeted surface. Photograph cable routing before any disconnect.
- Label every screw + screw location (egg-carton trick). Never force connectors. Bend radius >=35mm on 12V-2x6 cables.
- Know your rollback path. BIOS flash is reversible via BIOS Flashback if you saved the previous file; component swap is not if you damaged a socket pin.
- For rack-mounted servers, line up a maintenance window with stakeholder notification before iDRAC / iLO / XCC firmware update.
FAQ
References
- Vendor support docs for NVIDIA Datacenter (H100, H200, B200 Blackwell, GH200, GB200) (Dell SupportAssist, HP UEFI Diagnostics, Lenovo Vantage, ASUS MyAsus, Apple Self Service Repair)
- Reddit hardware subs (r/buildapc, r/Amd, r/intel, r/nvidia, r/sffpc, r/homelab, r/MiniPCs, brand-specific subs)
- Tom's Hardware, GamersNexus, TechPowerUp, Notebookcheck, ServeTheHome
- Vendor status pages, BIOS/firmware release notes, and driver changelogs
Related fixes
Related guides worth a look while you sort this one out: