HPE P67869-001
NVIDIA L40S 48 GB PCIe Accelerator Card
The HPE P67869‑001 is a factory‑qualified NVIDIA L40S Accelerator card boasting 48 GB of GDDR6 memory and a PCIe4.0 ×16 interface. Based on the NVIDIA Ada Lovelace architecture, the L40S delivers leadership-tier performance across AI inferencing, graphics rendering, simulation, and real-time visualization. With 18,176 CUDA cores, 568 fourth‑generation Tensor Cores, and 142 RT Cores, it accelerates generative AI (LLM inference/training), ray-traced graphics, scientific computing, and professional visualization workflows. Memory bandwidth reaches ~864 GB/sec and FP8 throughput up to 1,466 TFLOPS (with sparsity), enabling extreme scaling for multi-modal, multi-workload environments. This Accelerator is ideal for enterprise servers including HPE ProLiant DL and Edge platforms optimized for AI and GPU-intensive tasks.
Features
The L40S card fits in servers with PCIe 4.0 ×16 slots, delivering up to 64 GB/s bidirectional interconnect to the host CPU. Its dual-slot passive form factor supports up to 350 W of power consumption and requires a compatible server chassis with adequate cooling and thermal headroom. Supported on HPE platforms including DL380 Gen11/Gen12, DL385 Gen11/Gen12, DL320 Gen12, DL340 Gen12, and Edge Server e930t/e920d Gen10+, it integrates with HPE iLO for hardware monitoring and management. The GPU combines AI compute and RTX graphics capabilities: 91.6 TFLOPS FP32, 366 TFLOPS TF32, 733 TFLOPS FP16/BF16, and 1,466 TFLOPS FP8 (with structural sparsity) performance. Matrix and tensor workloads, real-time rendering, or simulation applications benefit from DLSS 3 and Transformer Engine. Multi-instance GPU (MIG) is not supported on L40S. Secure boot and NEBS Level 3 readiness add enterprise reliability.
Advanced Options
The NVIDIAL40S excels in generative AI inference workloads, offering up to 5× higher real-time performance compared to its predecessor A40. Its memory and compute architecture deliver smooth, multi-frame real-time graphics or LLM inference with low latency. In models using sparsity-aware formats and sparse tensor acceleration, peak FP8 performance reaches 1,466 TFLOPS, greatly enhancing throughput per watt. Designed for enterprise data centers, the L40S meets stringent reliability and operational standards—featuring passive cooling, support for secure boot, root-of-trust authentication, and NEBS Level 3 readiness. Integrated support via HPEiLO enables health monitoring, firmware updates, and power/thermal alerts. When used in qualified HPE servers, the card benefits from validated thermal and power profiles and compatibility certifications.
Product Key Features
-
48 GB GDDR6 ECC memory with ~864 GB/sec bandwidth
-
PCIe Gen 4 ×16 interface (~64 GB/s bidirectional)
-
Ada Lovelace architecture: 18,176 CUDA, 568 Tensor, 142 RT cores
-
AI and graphics acceleration: up to 1,466 TFLOPS (FP8 with sparsity)
-
FP32 @ 91.6 TFLOPS; TF32 @ 366 TFLOPS; FP16/BF16 @ 733 TFLOPS
-
Compatible with HPE DL3x0/385 Gen11/Gen12 platforms
-
Passive thermal design; max power draw ~350 W; requires sufficient cooling
Reliability & Performance
Deployments in AI inference clusters or HPC environments benefit from the card’s high-bandwidth PCIe interface and massive compute density. It supports AI frameworks via CUDA-X optimized libraries, NVIDIA AI Enterprise software, and GPUs in HPE servers validated for generative AI workloads. Whether for rendering or multi-instance AI service hosting, the L40S delivers predictable performance at scale.
Key Features :
-
48 GB ECC GDDR6 memory with 864 GB/sec bandwidth
-
PCIe Gen4 ×16 interface (~64 GB/s host bandwidth)
-
Supports FP8, FP16/BF16, TF32, and FP32 workloads
-
Ideal for AI inference, graphics, rendering, simulation
-
Compatible with select HPE ProLiant Gen11/Gen12 and Edge servers
-
Passive-cooled design; max power ~350 W; requires adequate airflow





Write a Review