NextGenEnergyJobsRenewable Energy Jobs
CompaniesCitiesIndustries

NextGenEnergyJobs

The #1 platform for renewable energy careers. Join thousands of professionals who've found their dream jobs in renewable energy, sustainability, and renewable tech.

0+Newsletter subscribers
25K+Jobs posted
100+Companies

Sustainability Partners

Sustainability Software DirectoryRefurbished Tech Guide

Find Jobs

  • All Jobs
  • By Location
  • By State
  • International
  • By Industry
  • Top Companies
  • Job Titles

Job Types

  • Remote Jobs
  • Hybrid Jobs
  • Full-time
  • Part-time
  • Contract
  • Internships
  • Visa Sponsored

Experience

  • Entry Level
  • Mid Level
  • Senior Level
  • Executive
  • Remote Internships

Resources

  • Career Advice Hub
  • Top 10 Jobs
  • Solar Sales Salary
  • Become Solar Engineer
  • Salary Insights
  • CV Analyzer
  • Post a Job

Popular Job Locations

San Francisco
245 jobs
Boston
189 jobs
Denver
167 jobs
Austin
143 jobs
New York
298 jobs
Chicago
132 jobs
Seattle
201 jobs
Portland
98 jobs
Los Angeles
176 jobs
San Diego
87 jobs
Washington DC
203 jobs
Atlanta
112 jobs

Hot Remote Specializations

Project ManagerSolar SalesCustomer SuccessData EntryAll Data Entry
© 2026 NextGenEnergyJobs. All rights reserved.
Privacy PolicyTerms of ServiceAbout UsContact
  1. Home
  2. Jobs
  3. Senior AI Infrastructure Engineer
Gatik logo

Senior AI Infrastructure Engineer

Gatik
Mountain View, Canada
Full Time
Posted May 27, 2026
$180k - $240k
Power Generation
~74 people viewed this recently
Apply Now

Application opens on company website

Job Description

At Gatik, we connect people of extraordinary talent and experience to an opportunity to create a more resilient supply chain and contribute to our environment’s sustainability.

Key Responsibilities

We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient This role is onsite 5 days a week at our Mountain View, CA office! • Distributed Training & ML Systems Support Scale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train. Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters. Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training. Intelligent Resource Scheduling: Optimize hardware utilization and cost-efficiency through Kubernetes-native GPU scheduling (NVIDIA GPU Operator, KubeFlow). Inference Performance Engineering: Deploy and scale optimized model artifacts using TensorRT, ONNX Runtime, and Triton Inference Server, fine-tuning pipelines for both real-time and batch processing • Agentic Infrastructure & Automation Self-Healing AI Infrastructure: Architect and deploy Autonomous AI Agents (LangGraph, CrewAI, or AutoGen) to monitor GPU cluster health, enabling automated real-time triage of hardware failures and NCCL timeouts. Agentic DevOps & CI/CD: Develop agent-driven automation, such as Agentic PR Reviewers for infrastructure code and AI agents that proactively suggest model-specific Kubernetes resource optimizations. Agentic Data Curation: Support researchers in building "Data Machines" where AI agents autonomously curate, label, and verify high-priority edge cases from raw data. • Model Management & Lifecycle (MLOps) Automated Lifecycle Management: Design and maintain ML infrastructure leveraging MLFlow, Argo Workflows, and Kubernetes to automate the end-to-end model lifecycle. Experiment & Model Tracking: Integrate feature stores and experiment tracking systems to provide a robust system of record for every model iteration. Deployment Strategies: Implement robust serving mechanisms, including A/B testing, shadow deployments, and rollback mechanisms • Cloud-Native Foundations & Data Integration Infrastructure as Code: Drive the "Everything as Code" philosophy using Terraform and Helm. Data Pipelines: Collaborate with data teams to scale ETL pipelines using Apache Airflow, Kafka, and Spark for large-scale dataset management. ○ Integrated Data Factories: Collaborate with data engineering teams to scale high-bandwidth ETL pipelines using Apache Airflow, Kafka, and Spark, ensuring seamless data flow from raw sensor logs to optimized storage in S3, GCS, or Delta Lake • Monitoring & Observability System Metrics: Define and track key ML system metrics, including training convergence, latency, throughput, and drift detection. Infrastructure Health: Maintain deep visibility into platform health using Prometheus, Grafana, OpenTelemetry, and ELK Stack. Deep Stack Observability: Develop comprehensive monitoring using Prometheus, Grafana, and OpenTelemetry to track low-level infrastructure health alongside high-level ML metrics like training convergence and throughput. AI-Specific Metrics & Drift: Define and monitor critical ML system KPIs, including model latency, inference throughput, and feature drift detection

Requirements

• Experience: 5+ years in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments. • ML Expertise: Deep understanding of multi-GPU training strategies (FSDP, DeepSpeed, Ray Train) and high-performance networking (NCCL, InfiniBand). • Infrastructure Automation: Mastery of Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration. • AI Agent Frameworks: Proven experience building or supporting Agentic Workflows for infrastructure or data automation (e.g., using LLMs to drive DevOps tasks). • Platform Mastery: Expertise in MLFlow, Argo Workflows, and Kubernetes. • Containerization: Strong experience with Docker, Kubernetes, and Helm. • Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation. • Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plus • Advanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs. • Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models). • Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss. Salary Ranges - $180,000- $240,000

Ready to Apply?

Join Gatik and make an impact in renewable energy

Apply Now

Stay Updated on Sustainability Jobs

Get the latest renewable energy jobs and career tips delivered to your inbox.

Job Alerts

Get notified about new sustainability jobs

More at Gatik

Chase Vehicle Operator

Springdale$0k

Verification & Validation (V&V) Engineer – HIL, Simulation & Autonomy Validation

Mountain View$220k

Verification & Validation (V&V) Deployment Engineer

Mountain View$220k

Jobs in Mountain View, Canada

NPI Alignment Engineer

Aeva$179k

Senior Software Engineer, Autonomy Visualization

Nuro$291k

Software Engineer, Data Platform

Nuro$241k

More jobs at Gatik

Gatik logo

Chase Vehicle Operator

Gatik
NEW
SpringdaleSpringdale, Argentina
Full Time
6h
$0k-0k/hr
Gatik logo

Verification & Validation (V&V) Engineer – HIL, Simulation & Autonomy Validation

Gatik
NEW
Mountain ViewMountain View, Canada
Full Time
6h
$160k-220k
Gatik logo

Verification & Validation (V&V) Deployment Engineer

Gatik
NEW
Mountain ViewMountain View, Canada
Full Time
6h
$160k-220k

More jobs in Mountain View, Canada

Aeva logo

NPI Alignment Engineer

Aeva
NEW
Mountain ViewMountain View, Canada
Full Time
6h
$132k-179k
Nuro logo

Senior Software Engineer, Autonomy Visualization

Nuro
Mountain ViewMountain View, California
Full Time
May 6
$194k-291k
Nuro logo

Software Engineer, Data Platform

Nuro
Mountain ViewMountain View, California
Full Time
May 6
$160k-241k