Application opens on company website
Job Description Summary
The Senior AI Architect, MLOps / DevOps / Cloud Engineering is an enterprise technical authority responsible for defining how AI solutions are productionized, deployed, operated, monitored, secured, and continuously improved across Wind Engineering.
Building on the broader Senior AI Architect mandate, this role provides specialized leadership in MLOps, DevOps, cloud and hybrid infrastructure, distributed AI systems, CI/CD, observability, model lifecycle management, and production reliability. The role establishes the architectures, engineering practices, reusable components, and operational standards required to move AI models and workflows from experimentation into reliable, maintainable, and scalable engineering products.
This person partners closely with Digital/IT, ARC Foundry, enterprise platform teams, AI engineers, data engineers, and Embedded AI Architects to ensure solutions use approved infrastructure and integration patterns, are designed for sustainable operation, and can transition into GE Vernova ownership without dependence on fragile code, undocumented environments, or external support.
1. Define the MLOps, DevOps, and Cloud Architecture for Wind Engineering AI
• Define and maintain reference architectures for deploying and operating AI solutions across cloud, edge, on-premises, and hybrid environments.
• Define how AI workloads use enterprise environments, including approved cloud services, container platforms, model repositories, data platforms, APIs, engineering applications, and authentication services.
• Make architecture decisions across cloud, edge, and on-premises execution based on data sensitivity, latency, compute demand, cost, reliability, and engineering workflow requirements.
2. Establish Production-Grade AI Delivery Pipelines
• Design standardized CI/CD and continuous training patterns for AI-enabled engineering applications.
• Establish automated pipelines for code build, testing, model validation, security checks, packaging, deployment, and rollback.
• Define quality gates that prevent models or AI services from progressing into production unless they meet documented software, model-performance, data-quality, security, and engineering-validation criteria.
3. Own Model Lifecycle and Production Operations Standards
• Define the operating model for AI models from development and validation through deployment, monitoring, retraining, retirement, and replacement.
• Establish model-registration and versioning practices that preserve provenance, approval evidence, performance baselines, applicability limits, dependencies, and release history.
4. Build AI Observability, Reliability, and Drift-Management Practices
• Define observability standards for AI applications, model services, pipelines, APIs, workflows, and supporting infrastructure.
• Define performance baselines, service-level expectations, alert thresholds, and escalation paths for production AI solutions.
• Establish diagnostic practices that distinguish model issues from data, application, infrastructure, integration, or workflow failures.
• Define incident-response, rollback, recovery, and post-incident learning practices for production AI systems.
5. Embed Security, Governance, and Auditability by Design
• Integrate cybersecurity, identity, access control, secrets management, network, data-protection, and audit requirements into AI platform and deployment architectures.
• Partner with AI Governance, cybersecurity, Digital/IT, and platform teams to translate policies into enforceable technical controls.
6. Lead Technical Transfer and Sustainable GE Vernova Ownership
• Assess whether AI applications are technically ready to transition from external partners, research teams, or pilot environments into sustained GE Vernova operation.
• Require maintainable code, automated deployment, operating documentation, monitoring, test coverage, version history, and clearly assigned support ownership before transfer.
• Ensure reusable components and lessons learned are incorporated into enterprise standards and reference architectures.
7. Provide Portfolio Architecture Review and Technical Escalation
• Review subsystem AI designs for deployability, scalability, reliability, security, maintainability, observability, cost, and supportability.
• Identify production risks early, particularly where research prototypes, local infrastructure, manual processes, or undocumented dependencies could prevent scale.
8. Mentor Architects and Raise Production Engineering Capability
• Mentor AI Architects and AI engineering teams in cloud architecture, DevOps, MLOps, observability, testing, security, and production-readiness practices.
• Define competency expectations and practical development pathways for engineers responsible for building and maintaining production AI solutions.
Apply directly on GE Vernova's site
New jobs across data centers, grid, nuclear, storage, generation, and renewables. One email a week.
Get alerts for jobs like this one