Minimum of 5 years of experience in DevOps, SRE or Platform Engineering, including at least 3 years of hands-on experience with Infrastructure as Code (IaC) in AWS production environments.
Strong experience with Terraform (AWS Provider) and/or OpenTofu, including reusable modules, semantic versioning and provider upgrades.
Hands-on experience with Ansible.
Production experience with Amazon EKS, including cluster version upgrades, add-on upgrades and node group migrations.
Solid knowledge of AWS services, particularly EC2, S3 and IAM.
Strong Kubernetes skills, including Helm, ingress controllers, probes, Horizontal Pod Autoscaler (HPA) and PodDisruptionBudgets (PDBs).
Experience administering PostgreSQL databases, including backup and restore procedures, high availability and failover testing.
Experience upgrading self-hosted platforms, particularly HashiCorp Vault, Sonatype Nexus and GitLab running on virtual machines.
Experience with CI/CD using GitLab, including pipelines, merge requests, branch protection and approval workflows, as well as Jenkins.
Experience with monitoring and alerting using Prometheus/Grafana or equivalent tools.
Experience implementing Kubernetes backup and restore procedures using Velero.
Knowledge of security fundamentals, including IAM/RBAC, least privilege, secrets management with Vault and TLS.
Strong Linux administration and Bash scripting skills.
Familiarity with Jira, Confluence and Visual Studio Code.
Ability to work autonomously, take ownership of technical activities and manage changes in production environments.
Fluency in Portuguese and English.
Responsabilities:
Ensure the availability, security, stability and efficiency of an AWS-based enterprise data platform.
Plan and execute upgrades of self-hosted HashiCorp Vault, Sonatype Nexus and GitLab, ensuring backups, rollback plans and post-upgrade validation are in place.
Upgrade Amazon EKS clusters and add-ons, and migrate node groups from Amazon Linux 2 to Amazon Linux 2023 without service downtime.
Migrate the end-of-life ingress-nginx controller to a replacement solution while maintaining uninterrupted access to exposed services.
Upgrade Terraform/OpenTofu and the AWS Provider, validating that execution plans are non-destructive before applying changes.
Deliver critical version and dependency updates for components such as Kafka, Jenkins and Uptime Kuma.
Design and implement PostgreSQL high-availability solutions, including failover testing and recovery validation.
Strengthen the resilience of core platform components and regularly test Kubernetes backup and restore procedures using Velero, measuring recovery times
Define and implement availability improvements for critical services, providing evidence that changes do not degrade service performance.
Improve monitoring and alerting, and implement self-healing mechanisms, including probes, automatic restarts, autoscaling and automated remediation.
Optimise EKS costs and capacity across pre-production and production environments through rightsizing, scheduled start/stop and tools such as Kubecost.
Develop and maintain reusable Terraform and Ansible modules, ensuring semantic versioning, tagging, pipeline templates and automated testing.
Manage infrastructure changes through GitLab merge requests and approval workflows, enforcing branch protection and change control processes.
Secure secrets using Vault and enforce least privilege and segregation of duties.
Investigate and remediate security findings identified by Checkmarx, Fortify and other security scanning tools in IaC modules.
Produce technical designs and ITIL-aligned operational procedures covering upgrades, backup and restore, and observability.
Maintain technical documentation in Confluence and prepare monthly progress reports.
Take ownership of ongoing activities, conduct knowledge transfer sessions and collaborate with platform owners and third-party providers.
Coordinate access requirements, maintenance windows and production change approvals.
Work within a Scrum team, contributing to two-week sprints and using Jira and Confluence for task management, collaboration and documentation.
Carta de presentación
Inicia sesión para generar una carta de presentación para esta vacante.