• 5+ years of hands-on experience in Software Engineering, Site Reliability Engineering (SRE), Production Engineering, Incident Operations, or related technical roles.
• Strong software development background with recent, demonstrable experience building and maintaining production-grade applications and automation in Python.
• Experience working in on-call environments with SLA/SLO-driven operational responsibilities.
• Proven ability to operate effectively during high-severity, real-time production incidents.
• Solid understanding of distributed systems, cloud-native architectures, and large-scale production environments.
• Experience troubleshooting complex application, infrastructure, and service reliability issues.
• Advanced proficiency in Python development, including building automation, tooling, integrations, and operational services.
• Experience with software engineering best practices including testing, code reviews, CI/CD, and version control.
• Strong understanding of the Software Development Lifecycle (SDLC) and production reliability engineering principles.
• Familiarity with Kotlin is a plus.
• Experience with Slack automation and operational workflows is desirable. Incident Response & Reliability.
• Hands-on experience coordinating, managing, and resolving production incidents.
• Experience assessing customer impact, driving remediation efforts, and leading technical investigations.
• Ability to create and execute operational runbooks and automate repetitive operational tasks.
• Strong understanding of observability, monitoring, alerting, and incident response processes.
• Experience performing root cause analysis and driving continuous reliability improvements.
• Monitoring and observability platforms (e.g., Datadog, Chronosphere)
• Incident management platforms (e.g., PagerDuty, Rootly)
• APIs and service integrations
• Production debugging and root cause analysis
• Reliability engineering concepts including SLI/SLOs, error budgets, toil reduction, and automated remediation