NVIDIA is a leader in AI technology, seeking a Software Engineer specializing in Deep Learning Inference. The role involves designing, building, and optimizing GPU-accelerated software for advanced AI applications, as well as contributing to high-performance open-source frameworks.
Responsibilities:
- Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI
- Scale performance of DL models across different architectures and types of NVIDIA accelerators
- Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions
- Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions
Requirements:
- Pursuing or recently completed a MS or PhD Computer Engineering, Computer Science, EECS, AI or related field or equivalent experience
- Software development experience
- Excellent C/C++ programming and software design skills
- SW Agile skills are helpful and Python experience is a plus
- Prior experience with training, deploying or optimizing the inference of DL models in production is a plus
- Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus
- GPU programming experience (CUDA, OAI TRITON or CUTLASS) is a plus
- Experience with Multi GPU Communications (NCCL, NVSHMEM)