العودة لجميع الوظائف

SRE / Platform Engineer

DevOps Egypt — Fully Remote Part-time

الوصف الوظيفي

Reliability is a product feature. Users notice it only when it's gone.

ErthDev is looking for an SRE / Platform Engineer to improve the reliability, observability, and developer platforms behind the systems we run and deliver.

You'll help teams ship safely and operate confidently across environments.

What you'll do

  • Improve service reliability, SLOs, and incident response practices.
  • Build and maintain platform tooling for engineering teams.
  • Strengthen observability: metrics, logs, traces, and alerting.
  • Automate operational toil out of deployments and recovery.
  • Partner with DevOps, backend, and product teams on production readiness.
  • Run post-incident reviews focused on learning, not blame.
  • Improve environment consistency and release safety.
  • Support capacity, performance, and resilience planning.
  • Document runbooks and operational standards.
  • Help define what "production ready" means for ErthDev projects.

Reliability at ErthDev

We don't confuse activity with reliability.

We care about measurable service health, fast recovery, and platforms that make the right path the easy path.

What you'll get

  • Fully remote work from Egypt
  • Flexible working hours with collaboration overlap
  • Cairo office access when needed
  • Paid annual leave and sick leave
  • Performance bonuses and annual salary reviews
  • Learning and development support
  • Courses and certifications
  • Conference opportunities
  • Equipment and internet support
  • Medical insurance
  • Mentorship and career growth
  • Exposure to projects across different industries

Hiring process

Application → Initial Conversation → Practical Assessment → Technical Interview → Final Conversation → Offer

If you enjoy making systems boringly reliable, we'd like to meet you.

Apply through LinkedIn Easy Apply or the ErthDev careers page.

المتطلبات

What we're looking for

  • Strong production operations experience.
  • Solid Linux, networking, and cloud fundamentals.
  • Experience with observability stacks.
  • Familiarity with containers and deployment automation.
  • Incident management and troubleshooting skills.
  • Ability to automate repetitive operational work.
  • Clear communication under pressure.
  • Product empathy for developer experience.

Experience with Kubernetes, Terraform, Prometheus/Grafana, OpenTelemetry, AWS, and on-call practices is a plus.

قدم الآن