We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Senior Network Engineer - GPU Cluster Networking

Advanced Micro Devices, Inc.
$173,600.00/Yr.-$260,400.00/Yr.
United States, California, San Jose
2100 Logic Drive (Show on map)
Sep 25, 2026


ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou'redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger- technologythat moves the world forward.Join us and, together, we'll advance your career.

THE ROLE:

Join AMD's IT Systems Engineering team as a Senior Network Engineer and help power the infrastructure behind next-generation AI and HPC workloads.

In this role, you will architect, automate, and operate high-performance backend networks supporting large-scale AMD Instinct GPU clusters. Working across network, AI, platform, and data center teams, you will ensure the network fabric delivers the bandwidth, latency, and reliability needed to enable world-class AI training, inference, and HPC performance.

THE PERSON:

You are a collaborative, results-driven engineer who thrives in fast-paced, complex environments. You bring a strong sense of ownership, a passion for solving challenging problems, and a commitment to operational excellence.

You communicate effectively across teams, build strong partnerships, and influence technical decisions through data and thoughtful execution. You are comfortable leading initiatives, mentoring others, and driving continuous improvement to deliver scalable, reliable solutions.

KEY RESPONSIBILITIES:

  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design and scale network fabrics supporting AI and HPC environments from individual racks to clusters of 10,000+ GPUs.
  • Own the end-to-end backend network architecture from GPU servers and NICs through the data center fabric.
  • Design and optimize high-speed Ethernet and RoCEv2 networks using modern routing, switching, and congestion-management technologies.
  • Develop scalable, resilient fabric architectures and drive capacity planning, topology modeling, and long-term growth strategies.
  • Optimize end-to-end communication performance across compute, network, and storage infrastructure.
  • Lead production incident response, root-cause analysis, and reliability improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, hardware refreshes, and fabric migrations.

PREFERRED EXPERIENCE:

  • Significant experience designing, deploying, and operating large-scale data center networks supporting AI, GPU, HPC, cloud, or other distributed computing environments.
  • Experience building and scaling backend network infrastructure for GPU clusters with 10,000+ GPUs or comparable hyperscale compute environments.
  • Strong knowledge of modern data center networking, including routing and switching, BGP, ECMP, QoS, network segmentation, and Ethernet fabric design.
  • Hands-on experience with RDMA and RoCEv2, including congestion management, traffic engineering, and lossless or near-lossless Ethernet environments.
  • Experience designing and operating scalable network architectures, including leaf-spine, Clos, EVPN, and VXLAN-based fabrics.
  • Understanding of GPU cluster topology and performance optimization across GPUs, NICs, CPUs, PCIe, NUMA, storage, and network infrastructure.
  • Experience with network observability, telemetry, monitoring, and performance analysis in large-scale production environments.
  • Experience with Juniper networking platforms, AMD Instinct accelerators, ROCm, RCCL, AMD Pensando technologies, or similar AI infrastructure solutions.

ACADEMIC CREDITALS:

  • Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.

LOCATION:

San Jose, CA OR Austin, TX

This role is not eligible for visa sponsorship.

#LI-BS1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.

This posting is for an existing vacancy.

Applied = 0

(web-9db6c7984-hfltr)