Do you think in playbooks and dashboards? Do you sleep better knowing an IAM platform is healthy, patched, and instrumented end to end? At Everpure, identity is the front door to everything we build, and we're looking for an IAM Site Reliability Engineer to keep that door running smoothly — and to make it better every day.
The Global Information Security Office (GISO) at Everpure is seeking an IAM Site Reliability Engineer to operate, automate, and continuously improve our enterprise Identity and Access Management services. This role is centered on keeping IAM platforms reliable and observable, automating operations, and making sure the rest of the organization always knows the state of the systems you run.
Enterprise IAM Platform Ownership : Own the end-to-end operation, high availability, and resilient performance of enterprise IAM platforms—including identity providers and lifecycle services—guaranteeing seamless, secure access for global users.
Infrastructure as Code & Workflow Automation : Build and maintain configuration management and automation frameworks using Terraform, Ansible, and Tines to streamline system patching, provisioning, and routine operations, eliminating manual overhead and deployment risk.
Observability, SLIs/SLOs & Incident Response : Establish, track, and tune observability tooling (metrics, logs, alerting) alongside SLIs, SLOs, and operational KPIs, driving proactive issue detection and rapid incident resolution during scheduled on-call rotations.
Change Governance & Cross-Functional Collaboration : Partner across technical and business teams to lead change management risk assessments, deploy new identity capabilities, and maintain actionable runbooks and architecture documentation to ensure transparent, repeatable operations.
Requirements
Production Automation & Infrastructure Skills : Hands-on proficiency operating production-grade enterprise infrastructure, using Infrastructure as Code (Terraform), configuration management (Ansible), and scripting (Python, PowerShell, or Bash) to automate workflows and identity lifecycle management.
Observability & Reliability Engineering : Demonstrated experience implementing telemetry using enterprise monitoring tools (such as Datadog, Prometheus, Grafana, or Splunk) and applying SLAs, SLOs, and operational KPIs to measure and elevate service health.
Incident Management & Root Cause Analysis : Expertise in incident response and structured problem management, with the ability to lead resolution efforts during service disruptions and implement effective preventive controls.
Technical Communication & Operational Documentation : Exceptional written and verbal communication skills to translate complex operations into clear updates for diverse audiences, paired with disciplined habits for creating maintainable runbooks, SOPs, and system documentation.
We are primarily an in-office environment and therefore, you will be expected to work from the Santa Clara office in compliance with Pure’s policies, unless you are on PTO, or work travel, or other approved leave.
Salary ranges are determined based on role, level and location. For positions open to candidates in multiple geographical locations, the base salary range is reflective of the labor market across the applicable locations.
This role may be eligible for incentive pay and/or equity.
There is no application deadline and we accept applications on an ongoing basis until the job is filled.
Ready to Apply?
Join Pure Storage and make an impact in renewable energy