Thanks for the detailed data — this is really useful.
Looking at your numbers: power isn't flat/stuck — it moves with load (idle ~24-31W → steady ~40-47W → bursts to 68-69W), and your graphics clock is holding near max boost (2483MHz) throughout.
Two things in your own log we'd like clarified before going further:
1. Spark-2 shows 95% SM utilization at "idle" while Spark-1 shows 0% — was there any residual traffic or NCCL sync activity on Spark-2 at that moment?
2. The mem column in dmon reads 0% under active load — that's unexpected for a decode workload and worth double-checking against your own monitoring daemon's readings.
Could you also run this on both nodes while under load, and paste the full output?
nvidia-smi -q -d PERFORMANCE
We're specifically looking for two sections in the output:
- * Clocks Event Reasons — a line called SW Power Cap that will say either Active or Not Active
- * Clocks Event Reasons Counters — a line called SW Power Capping showing a number in microseconds (run the command twice, a few seconds apart, to see if that number is increasing)
If SW Power Cap shows Not Active and the counter isn't climbing, that tells us the GPU isn't being software-capped, which points us toward the power draw simply being what this workload needs rather than a limiting issue.
Hi, sorry for the delay in getting back to you!
On 1 - Since with this model running across two nodes, 'Spark-2' is the worker so I suspect it's node2 polling for incoming data across the CX-7 link. I'm not sure if it uses keeping the SM high through the polling?
On 2 - the nvidia-smi dmon was effectively reporting "Memory-Usage: Not Supported".
This is live today:
$ nvidia-smi -q -d MEMORY
FB Memory Usage
Total : N/A
Used : N/A
Free : N/A
BAR1 Memory Usage
Total : N/A
Used : N/A
Free : N/A
$ nvidia-smi dmon -d 1 -c 3
sm mem
0 0
0 0
0 0
Regarding the nvidia-smi Performance command results whilst under load:
Power Cap Diagnostic — Dual Pass
GPU Util•
Value: 96%

Load active
Power Draw
Value: 43.43
WTemp
Value: 55°
CClock
Value: 2,502 MHz (P0 state)
Pass 1 — 15:15:06 UTC
SW Power Cap : Not Active
SW Power Capping : 2,156,970,640 us
Pass 2 — 15:15:14 UTC (8s later)
SW Power Cap : Not Active
SW Power Capping : 2,156,970,640 us ← IDENTICAL
Looks to be more memory bandwidth limited with a short test. I've then re-run the same thing over a 10 minute period taking the measurement at the 2nd, 5th and 9th minute.
Results — 10-Minute Sustained Load Power Cap Diagnostic
Load Profile
- Duration: 600 seconds (10 minutes exactly)
- Requests sent:1,186 (1 every ~500ms)
- Model: DeepSeek-V4-Flash on Spark-1:8000 (TP=2 cluster)
- Exit code: 0 — clean completion
GPU State Throughout
GPU Util•
During Load: 96%•
Post-Load: 96% (residual)
Power Draw•
During Load: 46.8–47.1 W•
Post-Load: 45.5 W
Temp•
During Load: 70–71°C•
Post-Load: 70°C
---
Three nvidia-smi -q -d PERFORMANCE Readings
Timestamp
• ~17s (early): 15:23:14
• ~23s (early): 15:23:20
• 9:20 min

: 15:32:17
• Δ: +9m03s
SW Power Cap
• ~17s (early): Not Active
• ~23s (early): Not Active
• 9:20 min

: Not Active
• Δ: No change
SW Power Capping
• ~17s (early): 2,156,970,640 μs
• ~23s (early): 2,156,970,640 μs
• 9:20 min

: 2,156,970,640 μs
• Δ: Zero change
Verdict
Result: SW Power Cap Not Active consistently
Counter climbing?
• Result: Frozen — no movement across 9+ minutes of 96% GPU loadSoftware power limiting?
• Result: None whatsoever
The SW Power Capping counter did not increment by a single microsecond across 1,186 inference requests over 10 minutes of sustained decode load at 96% GPU utilisation.
---
During the same 10-minute load these are the GPU Power Draw readings:
Captured Power Draw Readings
15:23:02 UTC
• Elapsed: ~5s in
• GPU Util: 96%
• Power: 46.86 W
• Temp: 71°C
• Source: Load confirmation check
15:32:17 UTC
• Elapsed: ~9m20s in
• GPU Util: 96%
• Power: 46.97 W
• Temp: 70°C
• Source: 9-min PERFORMANCE + GPU state
15:34:04 UTC
• Elapsed: post-load
• GPU Util: 96%
• Power: 45.51 W
• Temp: 70°C
• Source: Final baseline
Dual-Node GPU Power Capture — Fixed Results
Test: 10-minute sustained inference load on DeepSeek-V4-Flash (TP=2 across both nodes)
Load: 1,186 requests via Spark-1:8000 | Capture: 120 samples @ 5s intervals
Under-load samples: 111 per node (≥95% util)
Side-by-Side
Avg GPU Power
• Spark‑1 (Node 1): 47.8 W
• Spark‑2 (Node 2): 46.4 W
• Δ: +1.5 W
Power range
• Spark‑1 (Node 1): 42.9 – 49.2 W
• Spark‑2 (Node 2): 42.1 – 48.1 W
• Δ: —
Avg Util
• Spark‑1 (Node 1): 96%
• Spark‑2 (Node 2): 96%
• Δ: identical
Avg Temp
• Spark‑1 (Node 1): 73°C
• Spark‑2 (Node 2): 71°C
• Δ: +1°C
Peak temp
• Spark‑1 (Node 1): 76°C
• Spark‑2 (Node 2): 75°C
• Δ: —
GPU clock
• Spark‑1 (Node 1): 2,463–2,515 MHz
• Spark‑2 (Node 2): 2,463–2,528 MHz
• Δ: comparable
Idle power
• Spark‑1 (Node 1): 14.6 W
• Spark‑2 (Node 2): 14.8 W
• Δ: ~identical
Thermal Ramp
First 1/3
• Spark‑1: 47.0W @ 70°C
• Spark‑2: 45.3W @ 68°C
Middle 1/3
• Spark‑1: 48.5W @ 75°C
• Spark‑2: 47.0W @ 74°C
Last 1/3
• Spark‑1: 48.0W @ 73°C
• Spark‑2: 46.8W @ 73°C
Ramp
• Spark‑1: 42.9W @ 56°C → 47.8W @ 72°C
• Spark‑2: 42.6W @ 54°C → 47.1W @ 73°C
Summary
- ~47W is the natural demand of a 96%-utilised GB10 Blackwell running NVFP4 MoE inference
- Spark-1 uses ~1.5W more (+3.2%) than Spark-2 — likely from API server overhead (HTTP parsing, tokenisation, scheduling)
- No throttling — power curve is flat after warm-up, SW Power Capping counter is frozen
- Idle floor ~14.7W — consistent between nodes
- Thermal soak at 73–75°C after ~6 minutes
So I guess you’ve pointed me to the conclusion that I’m not hitting a cap but that VLLM running deekseek across 2 nodes isn’t pushing the GB10 to it’s limit and perhaps its the CX-7 bandwidth which limits the performance?
P