Application opens on company website
GE is building operations teams focused on performance and availability of Compute and Network infrastructure consumed by all business segments.
In this role, you will: • Available to quickly respond and resolve critical service outages severely impacting consumers • Establish performance baseline, capacity thresholds, correlate events, and define monitoring/alerting criteria • Develop automated solutions to address potential problems before they result in a service interruption • Provide impact assessment and mitigation plan for changes going into the production environment • Investigate root cause of severe and systemic outages, identify corrective actions and apply across the enterprise • Develop availability measures that align with consumer experience to accurately assess the usability of crucial services • Build capacity models to baseline transactional load compared to resource performance and leverage data to predict overall system capacity while automating load placement to avoid outages • Identify thresholds for all critical links in the data path to quickly isolate where imbalances may result in potential outages • Analyze failure points in services to model risk level and resolution steps if failure occurs. Assist in driving architecture enhancements into system to mitigate potential failure points • Programmatically monitor for and remediate configuration drift of critical devices • Develop response plans to potential failure points and evaluate effectiveness during planned tests • Perform comprehensive operational health checks of the entire services to identify areas of concern and track activities to drive improvements at all levels of the architecture • Provide technical coaching and direction to more junior teammates
Apply directly on GE Vernova's site
New jobs across data centers, grid, nuclear, storage, generation, and renewables. One email a week.
Get alerts for jobs like this one