NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The role involves leading cross-stack engineering efforts, validating new capabilities, and producing credible evidence while collaborating with various teams to enhance the software stack.
Responsibilities:
- Lead toolchain, build, code-health, and performance-workflow adoption projects across large software components
- Evaluate and deploy supported GCC and LLVM/Clang toolchains, compiler options, linkers, sysroots, and cross-compilation configurations
- Integrate modern toolchains and workflows into complex build systems and continuous-integration environments
- Establish useful Clang diagnostic builds and targeted sanitizer coverage in partnership with component owners. Evaluate techniques such as link-time optimization, profile-guided optimization, AutoFDO, and BOLT on representative software
- Use profiling, PMU data, flamegraphs, and binary/source analysis to identify actionable performance and code-quality findings
- Build automation, wrappers, validation scripts, dashboards, and migration helpers where they improve adoption and repeatability. Measure changes in runtime, code size, build time, launch latency, throughput, quality, or engineering velocity
- Drive complex blockers to the appropriate compiler, runtime, library, infrastructure, or component owner
- Document validated approaches as reusable playbooks for other engineering teams. Communicate technical results and tradeoffs clearly to engineers and leadership
Requirements:
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience
- 8+ years of relevant systems software engineering experience
- Strong C and C++ development, debugging, and code-review skills
- Strong Linux systems knowledge and hands-on experience with complex native software stacks
- Practical experience with GCC or LLVM/Clang, linkers, compiler options, and build systems
- Experience working in large, multi-component codebases and CI environments
- Experience with performance profiling, root-cause analysis, and before-and-after validation
- Ability to lead projects with substantial technical and organizational ambiguity
- Excellent written and verbal communication and a record of effective cross-team collaboration
- Experience with Arm64 systems, CPU architecture, vectorization, or SVE
- Experience with LTO, PGO, AutoFDO, BOLT, binary optimization, or code-layout analysis
- Hands-on experience with Clang diagnostics, AddressSanitizer, or other code-health workflows
- Familiarity with PMU analysis, perf, flamegraphs, BRBE, SPE, ETM, or similar profiling technologies
- Experience with cross-compilation, sysroots, large monorepositories, Perforce-scale development, or distributed build systems
- Experience turning a successful migration or optimization into a maintained workflow used by multiple teams
- Python or other scripting experience for engineering automation and data analysis