NVIDIA logo

Software Engineering Intern, DLFW Comms - 2027

NVIDIA · China Shanghai

Location
China Shanghai
Funding
~$5.13T–$5.43T
Posted
Sep 20, 2026

NVIDIA is hiring a Software Engineering Intern, DLFW Comms - 2027 based in China Shanghai. Every apply link on Engg.space goes straight to the company's own careers page - no recruiter middleman, no generic job-board form.

Apply directly at NVIDIA

Role details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are looking for a motivated Deep Learning engineer to integrate advanced communication technologies into AI stacks like PyTorch, vLLM, SGLang, TRT-LLM, and veRL. You will be working with the team that developed communication libraries -- such as NCCL and NVSHMEM -- for scaling Deep Learning applications. Your customers will have diverse multi-GPU needs, ranging from training on scales up to 100K GPUs to inference at microsecond latency. Communication performance between GPUs directly affects AI applications. Your work in AI toolkits will simplify these challenges for the community. This is an excellent opportunity for someone with an AI background to push the state of the art in this field. Are you ready to contribute to innovative technologies and help realize NVIDIA's vision? What you'll be doing: Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production. Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models. Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms. Conduct in-depth research to achieve SOL GPU performance. Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads. Collaborate with a very dynamic team across multiple time zones. What we need to see: You are pursuing a M.S. or Ph.D. in CE/CS/EE with a strong background in communication, kernel authoring, and/or AI training/inference. Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe). Solid understanding of LLM models and parallelisms. Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively. Ways to stand out from the crowd: Development experience with frameworks such as PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL. Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP & PP. Experience with CUDA kernel optimization and profiling. Experience with large-scale training or production inference stack. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

More roles at NVIDIA

Search all engineering roles