Senior Edge AI Engineer: LLM Inference (TensorRT)
NVIDIA
NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache management, within embedded and edge platforms.
You will collaborate across CUDA and robotics teams, optimize transformer components, and contribute to kernel development while staying ahead of LLM/VLM trends.
#J-18808-Ljbffr