Senior Solutions Architect – Customer Success, Partnership

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior Solutions Architect leading NVIDIA AI infrastructure, networking, and customer success projects. Designing, optimizing, and deploying large-scale GPU-accelerated systems for customers and partners.

Responsibilities:

  • Lead hands-on analysis, optimization, and performance tuning of complex GPU-accelerated systems and AI workloads
  • Ensure high availability and efficiency across customer data centers
  • Serve as a senior technical authority on NVIDIA technologies
  • Contribute to architecture reviews and guide infrastructure decisions at scale
  • Establish and refine monitoring and optimization methodologies using analytics, telemetry, and automation
  • Detect bottlenecks and improve infrastructure resiliency
  • Participate in post-deployment reviews and incident retrospectives
  • Provide insights into NVIDIA’s infrastructure strategy and help shape the customer experience
  • Complete and lead complex technical projects from initial design through implementation and continuous improvement
  • Ensure alignment to SLAs and mitigation of technical risks
  • Identify AI infrastructure opportunities in cloud and enterprise environments
  • Drive technical initiatives showcasing NVIDIA’s leadership
  • Interact with customers, partners, and internal teams on large-scale Networking, System Design, and Automation projects

Requirements:

  • 10+ years of experience in large-scale data center service operations with a focus on infrastructure
  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields
  • Strong analytical, problem-solving, and decision-making skills
  • Strong communication, time management, and organizational skills
  • Preferred certifications in data center, server, or networking technologies
  • Proficiency in system-level aspects, including Operating Systems, Linux kernel drivers, GPUs, NICs, and hardware architecture
  • Expertise in cloud orchestration software and job schedulers, including Kubernetes, Docker Swarm, and Slurm
  • Familiarity with cloud-native technologies and integration with traditional infrastructure
  • Deep familiarity with AI infrastructure and workflows, including training/inference pipelines, MLOps/DevOps tools, containerization, and large-scale system deployments
  • Knowledge of data center infrastructure operations, including safety, security, environmental controls, and standard operating procedures
  • Ability to lead discussions, influence outcomes, and build positive relationships with internal and external collaborators

Benefits:

  • Willingness to travel up to 25% for customer engagements and team collaboration