Python Developer (SRE)

We are looking for Python Developer (SRE).

 

Rate: 160-180 PLN/h net + VAT (B2B)

100% remote

 

Project Description
Join a global technology organization focused on ensuring the reliability, stability, and operational excellence of large-scale production systems. The role sits at the intersection of Incident Operations, Site Reliability Engineering (SRE), and technical stakeholder communication, supporting real-time incident management, impact assessment, and operational improvements in a fast-paced, highly available environment. You will work closely with engineering and operational teams to maintain service reliability, improve incident processes, and drive automation initiatives

Responsibilities:

  • Monitor, triage, and coordinate responses to production incidents and operational alerts.
  • Act as a central coordination point between engineering teams and key stakeholders during incidents.
  • Assess incident impact, determine severity, and coordinate communications according to SLA commitments.
  • Manage incident lifecycles from detection through resolution and post-incident activities.
  • Maintain external-facing incident communications and status updates.
  • Support incident reporting, root cause analysis (RCA), and operational reviews.
  • Contribute to process improvements, automation initiatives, and operational tooling enhancements.
  • Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
  • Participate in reliability-focused development activities and support operational excellence initiatives.

We are looking for:

•  5+ years of hands-on experience in Software Engineering, Site Reliability Engineering (SRE), Production Engineering, Incident Operations, or related technical roles. 

•  Strong software development background with recent, demonstrable experience building and maintaining production-grade applications and automation in Python. 

•  Experience working in on-call environments with SLA/SLO-driven operational responsibilities.

•  Proven ability to operate effectively during high-severity, real-time production incidents. 

•  Solid understanding of distributed systems, cloud-native architectures, and large-scale production environments. 

•  Experience troubleshooting complex application, infrastructure, and service reliability issues.

•  Advanced proficiency in Python development, including building automation, tooling, integrations, and operational services. 

•  Experience with software engineering best practices including testing, code reviews, CI/CD, and version control. 

•  Strong understanding of the Software Development Lifecycle (SDLC) and production reliability engineering principles. 

•  Familiarity with Kotlin is a plus. 

•  Experience with Slack automation and operational workflows is desirable. Incident Response & Reliability.

•  Hands-on experience coordinating, managing, and resolving production incidents. 

•  Experience assessing customer impact, driving remediation efforts, and leading technical investigations. 

•  Ability to create and execute operational runbooks and automate repetitive operational tasks. 

•  Strong understanding of observability, monitoring, alerting, and incident response processes. 

•  Experience performing root cause analysis and driving continuous reliability improvements.

•      Monitoring and observability platforms (e.g., Datadog, Chronosphere)  

•        Incident management platforms (e.g., PagerDuty, Rootly) 

•         APIs and service integrations 

•         Production debugging and root cause analysis

•         Reliability engineering concepts including SLI/SLOs, error budgets, toil reduction, and automated remediation 

ID: 1874 job_post.published_on: 28/09/2026
announcement.apply