Senior Solutions Architect, Cloud Partner Operations
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
NVIDIA Solutions Architect improving Day 2 operations for AI cloud partners. Advancing reliability, performance, economics, and operational maturity across large-scale AI infrastructure.
Responsibilities:
- Solve hard Day 2 operations problems at scale alongside partner engineers.
- Find causes, prototype approaches, validate them under representative load, and leave partners with operable practices.
- Prepare partners for new NVIDIA platforms, capacity, services, and use cases.
- Drive adoption in live environments without degrading service.
- Improve reliability, performance, utilization, recovery time, and cost per token.
- Identify and help close maturity gaps across people, process, tooling, telemetry, security, and incident response.
- Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows.
- Identify cross-partner patterns and provide field evidence to account teams, support, product, and engineering.
- Improve NVIDIA Cloud Partner Day 2 operations and ecosystem capability.
Requirements:
- BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field, or equivalent experience.
- 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
- Experience building, operating, or improving distributed infrastructure under real production load.
- Deep expertise in at least one part of the Day 2 stack, with hands-on large-scale GPU, HPC, or cloud infrastructure experience.
- Experience with relevant technologies such as DCGM, BMC/Redfish, firmware and driver lifecycle, InfiniBand or high-speed Ethernet, NCCL, UFM, Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
- Working experience with Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation using Terraform, Ansible, Argo CD, or similar tooling.
- Strong Linux knowledge.
- Experience with Python, Bash, or similar scripting for automation.
- Evidence-led troubleshooting across system boundaries.
- Ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority.
- Strong communication, prioritization, and time-management skills across multiple partner engagements.
- Preferred/standout experience operating GPU clouds, HPC environments, or large-scale AI platforms under customer load; building 24/7 operations; NVIDIA rack-scale platforms; NVIDIA operations technologies; fleet health or unit economics improvements.
Benefits:
- Competitive salaries
- Generous benefits package
- Equity















