Summary

Run type
Validate
Certification
PASS for the GPU only
Claim
GPU only NVIDIA L4
Host
Custom build Google Google Compute Engine
CPU
Intel(R) Xeon(R) CPU @ 2.20GHz · 1 socket · 2 cores / 4 threads
CPU flags
90 feature flags (informational)
All flags, as reported:
3dnowprefetchabmadxaesapicaratarch_capabilitiesavxavx2avx512_vnniavx512bwavx512cdavx512dqavx512favx512vlbmi1bmi2clflushclflushoptclwbcmovconstant_tsccpuidcx16cx8deermsf16cfmafpufsgsbasefxsrhlehthypervisoribpbibrsibrs_enhancedinvpcidlahf_lmlmmcamcemd_clearmmxmovbempxmsrmtrrnonstop_tscnoplnxpaepatpcidpclmulqdqpdpe1gbpgepnipopcntpsepse36rdrandrdseedrdtscprep_goodrtmsepsmapsmepssssbdssesse2sse4_1sse4_2ssse3stibpsyscalltsctsc_adjusttsc_known_freqvmex2apicxgetbv1xsavexsavecxsaveoptxsavesxtopology
GPU
NVIDIA L4 (nvidia 610.57.04) what this run certifies
Memory
16 GB · 1 × RAM
Memory modules
SlotChannelSizeType SpeedRankManufacturerPart number
DIMM 0 Not Specified 16 GB RAM - - - -
1 of 1 slots populated
AlmaLinux
-
Kernel
6.12.0-211.46.1.el10_2.x86_64 tainted (4096) An out-of-tree module was loaded
Suite
0.1.0 (0.1.0-7.el10)
Submitted by
jonathan · 2026-08-19 17:58

Result counts

11 pass

Linked components

Validation tests

TestCategorySeverityStatusDetail
bench.gpu.clpeak gpu pass
bench.gpu.cuda-bandwidth gpu pass
validate.gpu.cuda-bandwidth gpu required pass host to device 10.87 GiB/s, device to host 12.13 GiB/s
validate.gpu.cuda-devicequery gpu required pass CUDA sees 1 device(s): NVIDIA L4 (runtime 13030, driver 13030)
validate.gpu.cuda-vectoradd gpu required pass 16777216 elements computed correctly on the device
validate.gpu.driver gpu informational pass 1 with a driver (nvidia)
validate.gpu.nvidia-smi gpu required pass 1 NVIDIA GPU(s), driver 610.57.04
validate.gpu.opencl-compute gpu required pass 4194304 elements computed correctly on NVIDIA L4
validate.gpu.opencl-devices gpu required pass OpenCL sees 1 hardware device(s): NVIDIA L4
validate.gpu.vulkan-compute gpu required pass 16777216 bytes filled and verified on NVIDIA L4 (discrete)
validate.gpu.vulkan-devices gpu required pass Vulkan sees 1 hardware device(s): NVIDIA L4

Benchmarks

BenchmarkMetricValue
GPU compute bench.gpu.clpeak OpenCL Compute · Single precision opencl_single_precision_compute primary 26,537 GFLOPS ↑ better
GPU compute bench.gpu.clpeak OpenCL Compute · Double precision opencl_double_precision_compute 472 GFLOPS ↑ better
GPU compute bench.gpu.clpeak OpenCL Compute · Integer opencl_integer_compute 11,919 GOPS ↑ better
GPU compute bench.gpu.clpeak OpenCL Bandwidth · Global memory opencl_global_memory_bandwidth 258 GB/s ↑ better
GPU compute bench.gpu.clpeak OpenCL Bandwidth · Local memory opencl_local_memory_bandwidth 12,519 GB/s ↑ better
GPU compute bench.gpu.clpeak OpenCL Bandwidth · Image memory opencl_image_memory_bandwidth 189 GB/s ↑ better
GPU compute bench.gpu.clpeak OpenCL Bandwidth · Host transfer opencl_transfer_bandwidth 12.9 GB/s ↑ better
GPU compute bench.gpu.clpeak OpenCL Latency · Kernel launch opencl_kernel_launch_latency 17.4 us ↓ better
GPU compute bench.gpu.clpeak CUDA Compute · Single precision cuda_single_precision_compute 25,326 GFLOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · Double precision cuda_double_precision_compute 473 GFLOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · Half precision cuda_half_precision_compute 29,941 GFLOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · Mixed precision cuda_mixed_precision_compute 20,985 GFLOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · bfloat16 cuda_bfloat16_compute 23,849 GFLOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · Integer cuda_integer_compute 11,865 GOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Compute · Integer, int8 dot product cuda_integer_compute_int8_dp 31,650 GOPS ↑ better
GPU compute bench.gpu.clpeak CUDA Bandwidth · Global memory cuda_global_memory_bandwidth 247 GB/s ↑ better
GPU compute bench.gpu.clpeak CUDA Bandwidth · Local memory cuda_local_memory_bandwidth 12,011 GB/s ↑ better
GPU compute bench.gpu.clpeak CUDA Bandwidth · Image memory cuda_image_memory_bandwidth 188 GB/s ↑ better
GPU compute bench.gpu.clpeak CUDA Bandwidth · Host transfer cuda_transfer_bandwidth 13 GB/s ↑ better
GPU compute bench.gpu.clpeak CUDA Latency · Kernel launch cuda_kernel_launch_latency 7.14 us ↓ better
GPU compute bench.gpu.clpeak Vulkan Compute · Single precision vulkan_single_precision_compute 24,568 GFLOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · Double precision vulkan_double_precision_compute 469 GFLOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · Half precision vulkan_half_precision_compute 29,354 GFLOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · Mixed precision vulkan_mixed_precision_compute 24,550 GFLOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · bfloat16 vulkan_bfloat16_compute 21,106 GFLOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · Integer vulkan_integer_compute 8,064 GOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Compute · Integer, int8 dot product vulkan_integer_compute_int8_dp 8,405 GOPS ↑ better
GPU compute bench.gpu.clpeak Vulkan Bandwidth · Global memory vulkan_global_memory_bandwidth 247 GB/s ↑ better
GPU compute bench.gpu.clpeak Vulkan Bandwidth · Local memory vulkan_local_memory_bandwidth 11,490 GB/s ↑ better
GPU compute bench.gpu.clpeak Vulkan Bandwidth · Image memory vulkan_image_memory_bandwidth 226 GB/s ↑ better
GPU compute bench.gpu.clpeak Vulkan Bandwidth · Host transfer vulkan_transfer_bandwidth 12.9 GB/s ↑ better
GPU compute bench.gpu.clpeak Vulkan Latency · Kernel launch vulkan_kernel_launch_latency 130 us ↓ better
GPU transfer bench.gpu.cuda-bandwidth device_to_device 107 GiB/s ↑ better
GPU transfer bench.gpu.cuda-bandwidth device_to_host 12 GiB/s ↑ better
GPU transfer bench.gpu.cuda-bandwidth host_to_device primary 11.2 GiB/s ↑ better