Public summary
We are seeking a Principal Engineer specializing in Resilience to lead the development and management of incident response and operational excellence across a scalable commerce infrastructure. The role involves standardizing incident management processes, enhancing system visibility through metrics and dashboards, driving data-driven improvements, and managing organizational readiness for peak traffic events such as Black Friday. The position requires leadership in cross-functional initiatives and fostering organizational knowledge sharing with a hybrid work model based in Germany or Europe.
Location and work setup
- Location
- Berlin
- Remote status
- Hybrid
- German requirement signal
- No German Required Detected
- Detected job language
- English
Responsibilities
Lead end-to-end incident management processes including detection, response, communication, and post-mortem analysis; develop real-time monitoring metrics and dashboards to track system health; identify and implement process improvements using operational data; manage organizational preparedness for large-scale traffic events; coordinate cross-team resilience initiatives; collaborate closely with engineering leadership and principal engineers across various domains; promote communication, documentation, and training on resiliency and operational excellence.
Qualifications
7+ years experience in incident management and operational excellence; at least 5 years leading organization-wide resiliency and reliability initiatives; proven ability to manage high-stakes peak load events; strong data literacy for analyzing metrics and process improvements; excellent leadership, influence, and project management skills in Agile environments; fluent English communication skills; customer-focused with high self-awareness and mentoring passion; proven problem-solving ability and adaptability in complex technical environments.