Board note
Inference, GPU, and kernels demand in UK-eligible postings
GPU and serving stacks named in 259 UK-eligible JDs.
Inference hiring is the part of the market that is not “prompt engineer”. CUDA is kernels. TensorRT and Triton are serving. vLLM is the open engine labs actually type into a JD.
| Stack | Roles |
|---|---|
| cuda | 12 |
| tensorrt | 7 |
| triton | 5 |
| vllm | 5 |
Each /tech/cuda page is an apply list. This note is the inference slice: why those four names are the serving tell, and why “prompt engineer” is not on this board.
Live list: /tech/vllm · All notes · Monday Board