AR
arXiv CS.AI
7/23/2026

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
Short summary
This paper benchmarks confidential GPU inference on NVIDIA H100 under Intel TDX, comparing standard and confidential execution modes for Mistral-7B and Qwen3-30B-A3B. Confidential mode increases average time-to-first-token by 21.8% and 27.8% respectively, with global token throughput dropping 17.7% and 21.1%. Throughput gaps remain 11.5-20.2% under concurrency, but larger models saturate earlier in confidential mode, requiring adjusted capacity planning.
- •Confidential mode on H100/TDX increases TTFT by 22-28% and reduces throughput by 18-21% for Mistral-7B and Qwen3-30B
- •Throughput gaps of 11.5-20.2% persist under concurrency, with larger models saturating earlier in confidential mode
- •Capacity planning must account for both steady throughput penalty and earlier saturation for larger models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

