Alarm.com is the leading cloud-based platform for smart security and the Internet of Things.
Key Responsibilities
Own reliability for core production systems, including deployments, on-call operations, and incident response.
Drive capacity planning across cloud and physical infrastructure.
Define and improve reliability goals and operational outcomes (for example availability and MTTR).
Lead observability improvements across metrics, logs, traces, alert quality, and runbook maturity.
Build automation and tooling that improve reliability and engineering velocity.
Partner with engineering, product, and business stakeholders to align reliability priorities with roadmap goals.
Shape database and data-processing reliability, including performance and operational safety.
Apply sound engineering judgment to balance delivery speed with long-term resilience.
Other duties as assigned
Requirements
Bachelor’s degree in Computer Science, Computer Engineering, a related field, or equivalent practical experience.
10+ years of professional software engineering experience, including production operations and on-call ownership.
Proven track record of improving measurable reliability outcomes in distributed systems.
Deep expertise in observability and production support practices.
Strong networking fundamentals for distributed systems, including TCP/IP, DNS, TLS, HTTP, and L4/L7 behavior.
Strong experience with distributed systems technologies such as Kubernetes, Kafka, and Redis.
Experience with cloud infrastructure and operations at scale; Azure experience is a plus.
Strong software engineering fundamentals in object-oriented design and software development lifecycle practices; C# and .NET experience is a plus.
Strong analytical, communication, and decision-making skills under production pressure.
Experience leading high-severity incident response and driving durable post-incident improvements is a plus.
Strong database and data-processing experience, including performance tuning, schema evolution, and rollback-safe deployment practices, is a plus.
Deep experience with advanced traffic management and troubleshooting in production (for example Envoy, ingress controllers, service mesh, global load balancing, and cross-region routing) is a plus.
Experience with video streaming systems, WebRTC, IoT device ecosystems, OpenVPN, and MQTT is a plus.
Benefits & Perks
Collaborate with outstanding people: We have a strong focus on teamwork, and we work to create a collaborative and welcoming environment that enables our teams to excel.
Make an immediate impact: You can expect to be given real responsibility for bringing new technologies to the marketplace. You will be empowered to perform as soon as you join the team!
Be Empowered: We don't want to micro-manage you. We want you to own stuff and bring your experience to make those products the best in class.
Long-term employment based on a permanent employment contract (CoE).
Attractive benefits package: including medical care, life insurance, sports package, annual budget for professional development ($2,000).
Ready to Apply?
Join Alarm.com and make an impact in renewable energy