Board note

Inference, GPU, and kernels demand in UK-eligible postings

Published 9 Aug 2026 · Figures updated 25 Sept 2026

GPU and serving stacks named in 259 UK-eligible JDs.

Inference hiring is the part of the market that is not “prompt engineer”. CUDA is kernels. TensorRT and Triton are serving. vLLM is the open engine labs actually type into a JD.

StackRoles
cuda12
tensorrt7
triton5
vllm5

Each /tech/cuda page is an apply list. This note is the inference slice: why those four names are the serving tell, and why “prompt engineer” is not on this board.