Site Reliability Engineer
Ontrac Solutions · Karachi
Job description
About the role
The Cloud Operations team is expanding its Site Reliability Engineering function. As an SRE you will combine pragmatic operations with software craftsmanship to keep user‑facing services and production systems running smoothly.
Key responsibilities
- Participate in an on‑call rotation, responding to production incidents and supporting service engineers.
- Automate infrastructure using Ansible, Puppet, Terraform and Kubernetes.
- Design, build and maintain core infrastructure that scales to hundreds of thousands of concurrent users.
- Develop and improve monitoring and alerting (e.g., Prometheus) to focus on symptoms rather than outages.
- Document actions and turn findings into repeatable automation.
- Debug production issues across services and stack layers.
- Plan infrastructure growth and migration from AWS VMs to cloud‑native, container‑based deployments on Kubernetes (EKS).
Required profile
- Cloud‑first mindset, regardless of public‑cloud provider.
- Security‑first attitude.
- Strong systems thinking: edge cases, failure modes, behaviours.
- Comfortable with Linux and Windows environments.
- Proactive, go‑for‑it attitude to fix broken things.
Required skills
- Linux, Windows
- Configuration management: Ansible, Puppet
- Programming: Python, Java, Go, Node.js
- Container & orchestration: Docker, Kubernetes, EKS
- Infrastructure as code: Terraform
- Load balancing: Nginx, HAProxy
- Monitoring: Prometheus
- Cloud platforms: AWS
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Pakistan.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 2 weeks from now
37 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Ontrac Solutions
Karachi