Tasks:
- * Automate provisioning and operational processes
- * Build infrastructure as code
- * Design and manage AWS and OpenStack infrastructure
- * Design monitoring and observability systems
- * Develop AI assisted operational workflows
- * Diagnose data pipeline performance issues
- * Implement redundancy, failover, and disaster recovery
- * Improve platform reliability and availability
- * Lead incident response and root cause analysis
- * Maintain operational runbooks
- * Manage CI/CD pipelines
- * Operate and scale production Kubernetes clusters
- * Operate and tune data and streaming platforms
- * Perform capacity planning
- * Plan infrastructure upgrades and maintenance
Perks/Benefits:
Skills/Tech stack required:
[Agent Frameworks] [Ansible] [AWS] [Capacity Planning] [CI/CD] [CloudFormation] [Disaster Recovery] [Distributed Systems] [Elasticsearch] [ELK Stack] [Go] [Grafana] [Helm] [Incident Management] [Infrastructure as Code] [Jenkins] [Kafka] [Kubernetes] [Linux] [LLM Agent] [LLM agent frameworks] [Microservices] [Monitoring and observability] [MySQL] [Nginx] [Nifi] [OpenStack] [Prometheus] [Python] [Terraform] [Vertica] [Zookeeper]
Educational requirements:
[Bachelor's Degree] [Master's Degree] [PhD]
Role(s):
[Cloud Infrastructure Engineer] [Engineer] [Infrastructure Engineer] [Kubernetes Platform Engineer] [Platform Engineer] [Reliability Engineer] [Site Reliability Engineer]