Tasks:
- * Analyze server performance and troubleshoot issues
- * Automate operations and bulk administration
- * Build and optimize monitoring and logging platforms
- * Build monitoring dashboards and alert rules
- * Conduct root-cause analysis and incident reviews
- * Coordinate cross-functional issue resolution
- * Define and improve SLA and SLO metrics
- * Deploy and maintain production environments
- * Design alerting and escalation mechanisms
- * Establish operational standards and procedures
- * Govern alert noise and incident follow-up
- * Operate Linux servers and cloud hosts
- * Plan high availability and disaster recovery
- * Respond to complex production incidents
- * Review architecture for reliability and capacity
Perks/Benefits:
Skills/Tech stack required:
[Alertmanager] [Ansible] [Capacity Planning] [CentOS] [Cloud Platforms] [DNS] [ELK Stack] [Enterprise Linux] [Grafana] [High Availability] [HTTP/HTTPS] [Incident Management] [Keepalived] [Kernel tuning] [KVM] [Linux] [Linux Kernel] [Linux Kernel Tuning] [Linux performance] [Linux performance troubleshooting] [LVS] [Nginx] [OpenSearch] [Performance Troubleshooting] [Prometheus] [Python] [Red Hat] [Red Hat Enterprise] [Red Hat Enterprise Linux] [Shell] [SLA SLO] [Systemd] [TCP/IP] [Terraform] [Ubuntu] [VMware] [Zabbix]
Educational requirements:
[Bachelor's Degree]
Role(s):
[Administrator] [Engineer] [Linux Systems Administrator] [Linux Systems Engineer] [Monitoring Engineer] [Reliability Engineer] [Senior Linux Systems Engineer] [Site Reliability Engineer] [Systems Administrator] [Systems Engineer]