RTX 5090 Suprim Liquid - Reproducible VIDEO_TDR_FAILURE (0x116) after 3-month stable period

latal

New member
Joined
Apr 29, 2017
Messages
7
Hi everyone,
I'm posting here because I'm using an MSI GeForce RTX 5090 Suprim Liquid and I'm trying to determine whether this points to an isolated GPU issue or whether other MSI RTX 5090 owners have observed similar crashes and WinDbg signatures.

System
- GPU: MSI GeForce RTX 5090 Suprim Liquid
- Motherboard: ASUS ROG Crosshair X870E Hero
- CPU: Ryzen 9 9900X3D
- RAM: 96 GB DDR5 G.Skill 6000 CL30
- PSU: ASUS ROG Strix Platinum 1200W (ATX 3.1)
- Windows 11 24H2

Symptoms
Random black screens resulting in a VIDEO_TDR_FAILURE (0x116) and system reboot.
The crashes occur in different situations:
- Unreal Engine 5
- Autodesk Maya
- Fortnite
- Crimson Desert

Sometimes during normal desktop use, often after previously running GPU-accelerated applications.
The crashes are not exclusively related to heavy GPU load.
No artifacts are ever visible before the crash.
GPU temperatures remain around 55°C under load.

Timeline
- The issue first appeared in early 2026.
- It then disappeared completely for approximately three months without any hardware changes.
- Since mid-May 2026, the crashes have returned and now occur regularly.
- During both periods, the crashes produced the same WinDbg signature.

WinDbg
Every single minidump shows the same result:
- VIDEO_TDR_FAILURE (0x116)
- IMAGE_NAME: nvlddmkm.sys
- Offset: nvlddmkm+0x1958210
- Arg3 (NTSTATUS): 0xC000009A
- FAILURE_BUCKET_ID: 0x116_IMAGE_nvlddmkm.sys
- FAILURE_ID_HASH: {c89bfe8c-ed39-f658-ef27-f2898997fdbd}

Sometimes Windows also logs:
- Event ID 14 (nvlddmkm)
- Event ID 153 (nvlddmkm)
- "GpuRcReset TDR occurred on GPUID:100"
The WinDbg signature is identical across all crashes.

Already tested
- Clean Windows reinstall
- DDU before every driver installation
- Multiple NVIDIA Game Ready drivers
- Multiple NVIDIA Studio drivers
- Different motherboard BIOS versions
- CMOS reset
- EXPO enabled/disabled
- PCIe Auto and Gen4
- HAGS enabled/disabled
- MPO disabled
- G-Sync enabled/disabled
- Single monitor
- All monitors forced to the same refresh rate
- iGPU disabled
- Hyper-V/VBS investigated
- GPU underclock (-200 MHz core)
- Power Limit reduced to 70%
- Prefer Maximum Performance
- No overlays
- MemTest: PASS
- UserDiag: PASS
The crash remains identical.

Hardware checks
- OCCT 3D (multiple modes): PASS
- OCCT VRAM: PASS
- OCCT Power: PASS
- GPU temperatures remain below 55°C during stress tests
- No artifacts
- No WHEA errors

My question:
At this point I'm trying to determine whether this points to an isolated hardware failure or whether multiple RTX 5090 owners are reproducing the same failure signature.

Has anyone using an MSI RTX 5090 (Suprim, Suprim Liquid, Vanguard, Gaming Trio, etc.) observed the same WinDbg signature, especially:
- nvlddmkm+0x1958210
- Arg3 = 0xC000009A
- VIDEO_TDR_FAILURE (0x116)
or similar crashes?

I'm not trying to conclude that this is necessarily a driver issue or a hardware defect. I'm simply trying to determine whether other MSI RTX 5090 owners are reproducing the same failure signature before sending my card for RMA.

If you've analyzed your crash dumps with WinDbg and see similar results, I'd really appreciate comparing outputs, even if your hardware configuration is different.
 
Hi,

I may be having the same or similiar issue. My system:

- GPU: Asus TUF 5090 (non-OC)
- Motherboard: MSI MAG x870 Tomahawk Wifi
- CPU: Ryzen 9 9800X3D
- RAM: 64GB of TEAMGROUP T-FORCE DELTA RGB (2x32GB) DDR5-6000 CL30
- PSU: Corsair RM1200x Shift ATX3.1 with Corsair 90° 12V-2x6 power cable
- Windows 11 25H2
- NVidia driver: Geforce Game Ready 610.47

My problems first started in early April. There have two phases although symptoms have always been the same through both - screen goes black, fans slow down, and 5 to 20 seconds later the machine reboots. No artefacts and temperatures are fine. 3DMark stress tests always pass with flying colours and the crashes seem to happen in transition states (e.g. in the Arc Raiders menu, or only 1 or 2 minutes into Battlefield 6 or Avatar: Frontiers of Pandora gameplay). As a result, I don't believe it's uniquely GPU load issue. Heck, I've had 2 of the crashes in Age of Empires 2, which makes most GPUs yawn and go for a nap.

Phase 1: My Event Viewer was getting absolutely littered with a storm of nvlddmkm Event 153 and 14 errors. After about a month of futzing with BIOS settings (e.g. PCI4, other tweaks) I finally reseated my GPU and GPU power cable. I thought I completely cured the issue as I had 60 days and over 30 hours of heavy gaming after this.

Phase 2: Yesterday, it started again. Same symptoms but the that dozens of event 153/14s are gone, but there is still the same minidump with the 0x116 TDR with parameter 3 locked straight to 0xC000009A for STATUS_INSUFFICIENT_RESOURCES.

I honestly don't know where to go from here. It doesn't seem like a hardware issue given the many hours of heavy gaming I just went through before these new crashes appeared.
 
Last edited:
Hi,

I may be having the same or similiar issue. My system:

- GPU: Asus TUF 5090 (non-OC)
- Motherboard: MSI MAG x870 Tomahawk Wifi
- CPU: Ryzen 9 9800X3D
- RAM: 64GB of TEAMGROUP T-FORCE DELTA RGB (2x32GB) DDR5-6000 CL30
- PSU: Corsair RM1200x Shift ATX3.1 with Corsair 90° 12V-2x6 power cable
- Windows 11 25H2
- NVidia driver: Geforce Game Ready 610.47

My problems first started in early April. There have two phases although symptoms have always been the same through both - screen goes black, fans slow down, and 5 to 20 seconds later the machine reboots. No artefacts and temperatures are fine. 3DMark stress tests always pass with flying colours and the crashes seem to happen in transition states (e.g. in the Arc Raiders menu, or only 1 or 2 minutes into Battlefield 6 or Avatar: Frontiers of Pandora gameplay). As a result, I don't believe it's uniquely GPU load issue. Heck, I've had 2 of the crashes in Age of Empires 2, which makes most GPUs yawn and go for a nap.

Phase 1: My Event Viewer was getting absolutely littered with a storm of nvlddmkm Event 153 and 14 errors. After about a month of futzing with BIOS settings (e.g. PCI4, other tweaks) I finally reseated my GPU and GPU power cable. I thought I completely cured the issue as I had 60 days and over 30 hours of heavy gaming after this.

Phase 2: Yesterday, it started again. Same symptoms but the that dozens of event 153/14s are gone, but there is still the same minidump with the 0x116 TDR with parameter 3 locked straight to 0xC000009A for STATUS_INSUFFICIENT_RESOURCES.

I honestly don't know where to go from here. It doesn't seem like a hardware issue given the many hours of heavy gaming I just went through before these new crashes appeared.
Hi,

Thanks for replying. Your case sounds surprisingly close to mine, especially the two phases and the fact that the issue disappeared for a long period before coming back.

One thing that caught my attention is that we both have:
- VIDEO_TDR_FAILURE (0x116)
- Parameter 3 = 0xC000009A (STATUS_INSUFFICIENT_RESOURCES)
Crashes that often occur during transition states rather than under sustained GPU load.

Today I also managed to reproduce another crash while running UserDiag Extreme. Before the 0x116 bugcheck, Event Viewer logged:
- Event ID 153: "Resetting TDR occurred on GPUID:100"
- Event ID 153: "BusReset TDR occurred on GPUID:100"

Like you, I'm finding it difficult to believe this is a straightforward hardware failure. The issue disappeared completely for around three months without any hardware changes, then returned with exactly the same symptoms and WinDbg signature. Of course that doesn't rule out a hardware fault, but it does make me wonder whether there's another factor involved.

I was wondering if you could check one thing for comparison.

If you've had a chance to open one of your minidumps (located in C:\Windows\Minidump) in WinDbg, could you run !analyze -v and check whether it reports the same NVIDIA driver information?

Mine consistently reports:
- IMAGE_NAME: nvlddmkm.sys
- SYMBOL_NAME: nvlddmkm+1958210
- FAILURE_BUCKET_ID: 0x116_IMAGE_nvlddmkm.sys
- Arg3 = 0xC000009A (STATUS_INSUFFICIENT_RESOURCES)
If yours matches as well, that would be a very interesting data point.

Thanks!
 
Last edited:
Hi,

Regarding your crash while running UserDiag Extreme, I checked my event viewer to see if I've ever seen anything similiar. Before I re-seated my GPU, when I used to get dozens of Event 153s before the final BugCheck, each one said the same thing in the event data: "GpuRcReset TDR occurred on GPUID:100". I'm not sure how generic this error is but it seems like we're swimming in the same waters.

I checked my latest minidump with WinDbg as you requested (with !analyze -v) and it reports the exact same NVIDIA driver information to the letter as yours:

- IMAGE_NAME: nvlddmkm.sys
- SYMBOL_NAME: nvlddmkm+1958210
- FAILURE_BUCKET_ID: 0x116_IMAGE_nvlddmkm.sys
- Arg3 = 0xC000009A (STATUS_INSUFFICIENT_RESOURCES)

That being said, I'm not sure what a next step would be here. I've tried much of the same config changes as you, including PCIe Auto and Gen4, Prefer Maximum Performance, various BIOS versions, DDU-first driver installs, etc.

One reddit commenter (see: ) suggested the following config change combination:

- PCIE1: Set to Gen4
- FCH Spread Spectrum: Set to Enabled
- Power Supply Idle Control: Set to Typical Current Idle
- NB/SOC Voltage: Set to Override Mode at 1.25V

But it feels like I'm just stabbing in the dark (plus, I've already tried the first 2 with no positive outcome). I suppose I could start swapping out the motherboard or PSU but you and I have different motherboards and PSUs but seem to be having the same issue. What we do share is the AMD AGESA 1.3.0.x and Nvidia driver stacks... so I'm inclined to think the problem is somewhere in there.
 
Hi,

Regarding your crash while running UserDiag Extreme, I checked my event viewer to see if I've ever seen anything similiar. Before I re-seated my GPU, when I used to get dozens of Event 153s before the final BugCheck, each one said the same thing in the event data: "GpuRcReset TDR occurred on GPUID:100". I'm not sure how generic this error is but it seems like we're swimming in the same waters.

I checked my latest minidump with WinDbg as you requested (with !analyze -v) and it reports the exact same NVIDIA driver information to the letter as yours:

- IMAGE_NAME: nvlddmkm.sys
- SYMBOL_NAME: nvlddmkm+1958210
- FAILURE_BUCKET_ID: 0x116_IMAGE_nvlddmkm.sys
- Arg3 = 0xC000009A (STATUS_INSUFFICIENT_RESOURCES)

That being said, I'm not sure what a next step would be here. I've tried much of the same config changes as you, including PCIe Auto and Gen4, Prefer Maximum Performance, various BIOS versions, DDU-first driver installs, etc.

One reddit commenter (see: ) suggested the following config change combination:

- PCIE1: Set to Gen4
- FCH Spread Spectrum: Set to Enabled
- Power Supply Idle Control: Set to Typical Current Idle
- NB/SOC Voltage: Set to Override Mode at 1.25V

But it feels like I'm just stabbing in the dark (plus, I've already tried the first 2 with no positive outcome). I suppose I could start swapping out the motherboard or PSU but you and I have different motherboards and PSUs but seem to be having the same issue. What we do share is the AMD AGESA 1.3.0.x and Nvidia driver stacks... so I'm inclined to think the problem is somewhere in there.
One thing just came to mind.
Since your ASUS TUF 5090 also has a dual GPU BIOS (Performance / Quiet), which BIOS are you currently using?

My MSI Suprim Liquid also has a Silent and Gaming BIOS. I had been using the Silent BIOS since I bought the card, but I switched to the Gaming BIOS today. So far, I haven't had any crashes, but it's still much too early to draw any conclusions since I only switched a few hours ago.

Since changing to the Gaming BIOS, another UserDiag Extreme run completed successfully, whereas it had previously caused my system to crash. I'm going to keep testing over the next few days to see if the behaviour really changed or if it's just a coincidence.

Have you ever tried running your card with the Performance BIOS instead of the Quiet BIOS?
 
One thing just came to mind.
Since your ASUS TUF 5090 also has a dual GPU BIOS (Performance / Quiet), which BIOS are you currently using?

My MSI Suprim Liquid also has a Silent and Gaming BIOS. I had been using the Silent BIOS since I bought the card, but I switched to the Gaming BIOS today. So far, I haven't had any crashes, but it's still much too early to draw any conclusions since I only switched a few hours ago.

Since changing to the Gaming BIOS, another UserDiag Extreme run completed successfully, whereas it had previously caused my system to crash. I'm going to keep testing over the next few days to see if the behaviour really changed or if it's just a coincidence.

Have you ever tried running your card with the Performance BIOS instead of the Quiet BIOS?
Interesting. Mine came with the switch already on Performance mode and I've never touched it.
 
Interesting. Mine came with the switch already on Performance mode and I've never touched it.
Interesting, thanks for checking. That probably weakens the idea that this is simply caused by Quiet/Silent BIOS profiles in general.

In my case, switching from the Silent BIOS to the Gaming BIOS seems to have changed the behaviour so far, but it may be MSI-specific or simply a coincidence. I've only been testing it for a few hours, so I'll keep testing over the next few days before drawing any conclusions.

The fact that we still have the exact same nvlddmkm+1958210 / 0xC000009A signature on two different RTX 5090 models is still the most interesting part to me.

Another user on the NVIDIA forum also replied to my thread. His issue isn't identical (he gets LiveKernelEvent 141 instead of a 0x116 bugcheck), but he went through an RMA as well. His original RTX 5090 was replaced with another RTX 5090, yet the replacement showed the exact same behaviour. As a control test, he then reinstalled his old RTX 3090 in the exact same system, and all of the crashes disappeared completely.

That doesn't necessarily mean we're all facing the same root cause, but it's another data point suggesting this may not simply be isolated defective hardware.
 
I just ran a UserDiag Extreme for the first time without issue (well, besides my room temperature going up a couple degrees). Both CPU and GPU were pinned and drawing maximum power. No errors/critical events in the Windows event log.

Agreed that our identical error signature is interesting. I acknowledge the non-scientific nature of this forthcoming statement, however, to me this feels like the most advanced hardware available running at the absolute limits of the thresholds of power and component communication on the AM5 and Windows platform. Occasionally, rapid power transitions and handshaking protocols get nudged too far and we get our crash. That being said, if this is the case, there must be some mix of configuration that can mitigate this.

Can you share the link to your post on the NVIDIA forum?
 
I just ran a UserDiag Extreme for the first time without issue (well, besides my room temperature going up a couple degrees). Both CPU and GPU were pinned and drawing maximum power. No errors/critical events in the Windows event log.

Agreed that our identical error signature is interesting. I acknowledge the non-scientific nature of this forthcoming statement, however, to me this feels like the most advanced hardware available running at the absolute limits of the thresholds of power and component communication on the AM5 and Windows platform. Occasionally, rapid power transitions and handshaking protocols get nudged too far and we get our crash. That being said, if this is the case, there must be some mix of configuration that can mitigate this.

Can you share the link to your post on the NVIDIA forum?
Sure, here's the NVIDIA thread:

Regarding the temperatures during UserDiag Extreme, that's expected. UserDiag itself states that the Extreme test is not representative of normal PC usage. It's designed to stress the system as much as possible for stability testing, so both the CPU and GPU can run at or near their maximum power and temperature limits during the test.

I agree with your reasoning. I don't think we have enough evidence yet to say it's definitely a power transition or PCIe handshaking issue, but the timing of my crashes certainly points in that direction. The most reproducible one happened about 30–60 seconds after an OCCT 3D Adaptive Switch test had already completed successfully, rather than during the heavy load itself.

For now, I'm going to keep using the Gaming VBIOS on my MSI card and avoid changing anything else. If the system remains stable for several days, that'll be an interesting data point. If it crashes again with the exact same signature, then I'll know the VBIOS change wasn't the explanation.

I'll keep you updated if I find anything reproducible.
 
Sure, here's the NVIDIA thread:

Regarding the temperatures during UserDiag Extreme, that's expected. UserDiag itself states that the Extreme test is not representative of normal PC usage. It's designed to stress the system as much as possible for stability testing, so both the CPU and GPU can run at or near their maximum power and temperature limits during the test.

I agree with your reasoning. I don't think we have enough evidence yet to say it's definitely a power transition or PCIe handshaking issue, but the timing of my crashes certainly points in that direction. The most reproducible one happened about 30–60 seconds after an OCCT 3D Adaptive Switch test had already completed successfully, rather than during the heavy load itself.

For now, I'm going to keep using the Gaming VBIOS on my MSI card and avoid changing anything else. If the system remains stable for several days, that'll be an interesting data point. If it crashes again with the exact same signature, then I'll know the VBIOS change wasn't the explanation.

I'll keep you updated if I find anything reproducible.
Yes, I was aware that UserDiag would push temps quite far. Everything stayed within design spec.

Fingers crossed on your VBIOS switch on your GPU. I update my motherboard BIOS (I was only a bit behind) and NVIDIA graphics driver and then had that successful UserDiag Extreme test and then an additional 2.5 hours of Avatar: Frontiers of Pandora (an absolute monster of a game engine) without issue so I'm going to see how stable things stay. I think my next round of tests will first be changing my GPU VBIOS to "Quiet" and then perhaps the 4 settings that reddit commenter suggested that I listed earlier in the thread. I'll let you know if I learn anything new.
 
Yes, I was aware that UserDiag would push temps quite far. Everything stayed within design spec.

Fingers crossed on your VBIOS switch on your GPU. I update my motherboard BIOS (I was only a bit behind) and NVIDIA graphics driver and then had that successful UserDiag Extreme test and then an additional 2.5 hours of Avatar: Frontiers of Pandora (an absolute monster of a game engine) without issue so I'm going to see how stable things stay. I think my next round of tests will first be changing my GPU VBIOS to "Quiet" and then perhaps the 4 settings that reddit commenter suggested that I listed earlier in the thread. I'll let you know if I learn anything new.
Quick update on my side:
It's now been 9 days since I switched my MSI RTX 5090 Suprim Liquid from the Silent VBIOS to the Gaming VBIOS, and I haven't experienced any black screen or VIDEO_TDR_FAILURE so far.

I haven't been using Unreal Engine as heavily as before, but I've been using the PC normally and playing at least one Fortnite session every day. Right after the switch, I also passed multiple tests that had previously triggered crashes, including OCCT 3D Adaptive and UserDiag Extreme.

It's still too early to say the VBIOS change fixed the issue, especially since the problem was intermittent, but the system behavior has definitely been different since the switch.

Did you get a chance to test your ASUS card in Quiet VBIOS mode? I'm curious whether changing the GPU firmware profile had any impact on your side too.

Also, Lenovo recently published a GPU VBIOS update for some RTX 50-series laptops mentioning BSOD 0x133/0x116 fixes in the release notes, so it looks like GPU firmware-related stability fixes are at least a possibility on this generation.
 
Last edited:
Which BIOS version the GPU is using for Silent vbios and which version for Gaming vbios?
Look for each one of them in GPU-Z, Advanced tab, then from upper horizontal scroll select NVIDIA BIOS.
 
Quick update on my side:
It's now been 9 days since I switched my MSI RTX 5090 Suprim Liquid from the Silent VBIOS to the Gaming VBIOS, and I haven't experienced any black screen or VIDEO_TDR_FAILURE so far.

I haven't been using Unreal Engine as heavily as before, but I've been using the PC normally and playing at least one Fortnite session every day. Right after the switch, I also passed multiple tests that had previously triggered crashes, including OCCT 3D Adaptive and UserDiag Extreme.

It's still too early to say the VBIOS change fixed the issue, especially since the problem was intermittent, but the system behavior has definitely been different since the switch.

Did you get a chance to test your ASUS card in Quiet VBIOS mode? I'm curious whether changing the GPU firmware profile had any impact on your side too.

Also, Lenovo recently published a GPU VBIOS update for some RTX 50-series laptops mentioning BSOD 0x133/0x116 fixes in the release notes, so it looks like GPU firmware-related stability fixes are at least a possibility on this generation.
That's promising. Fingers crossed your stability continues. I am currently testing the latest BIOS for my motherboard that I installed 2 weeks ago and am now at about 10 hours of heavy gaming (Arc Raiders, Battlefield 6, Avatar) plus a successful UserDiag Extreme test. So, I'm not going to introduce any new variables for now. If I have another crash then I'm going to switch the VBIOS to quiet mode and test that. These testing cycles take a long time as there can sometimes be many hours in between my crashes. Will keep updating this thread with any new info.
 
Quick update on my side:
It's now been 9 days since I switched my MSI RTX 5090 Suprim Liquid from the Silent VBIOS to the Gaming VBIOS, and I haven't experienced any black screen or VIDEO_TDR_FAILURE so far.

I haven't been using Unreal Engine as heavily as before, but I've been using the PC normally and playing at least one Fortnite session every day. Right after the switch, I also passed multiple tests that had previously triggered crashes, including OCCT 3D Adaptive and UserDiag Extreme.

It's still too early to say the VBIOS change fixed the issue, especially since the problem was intermittent, but the system behavior has definitely been different since the switch.

Did you get a chance to test your ASUS card in Quiet VBIOS mode? I'm curious whether changing the GPU firmware profile had any impact on your side too.

Also, Lenovo recently published a GPU VBIOS update for some RTX 50-series laptops mentioning BSOD 0x133/0x116 fixes in the release notes, so it looks like GPU firmware-related stability fixes are at least a possibility on this generation.
Another crash just now. Same symptoms. Loaded into Avatar and after 30 seconds in the game world screen went black. This time, though, the computer did not reboot. I had to hold down the power button to shut it off. Event viewer only shows the unexpected shutdown. I'm at a bit of a loss at this point. I suppose I can try running my card in quiet VBIOS mode but that feels like a stab in the dark (especially since, physically, all this is doing is keeping temps lower by being more aggressive with fan curves).
 
Back
Top