Senior Solutions Architect, Cloud Partner Operations

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

NVIDIA Solutions Architect improving Day 2 operations for AI cloud partners. Advancing reliability, performance, economics, and operational maturity across large-scale AI infrastructure.

Responsibilities:

  • Solve hard Day 2 operations problems at scale alongside partner engineers.
  • Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices.
  • Prepare partners for new NVIDIA platforms, capacity, services, and use cases.
  • Drive adoption in live environments without degrading service.
  • Improve reliability, performance, utilization, recovery time, and cost per token.
  • Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response.
  • Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows.
  • Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering.
  • Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.

Requirements:

  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.
  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
  • Experience building, operating, or improving distributed infrastructure under real production load.
  • Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience.
  • Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
  • Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling.
  • Strong Linux knowledge.
  • Experience with Python, Bash, or similar scripting for automation.
  • Evidence-led troubleshooting across system boundaries.
  • Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority.
  • Strong communication, prioritization, and time-management skills across multiple partner engagements.
  • Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.

Benefits:

  • Competitive salaries
  • Generous benefits package
  • Equity