Job Description:
Role: Cloud Engineer
Location: Hybrid - 4 days/week onsite in SE MI
Duration: Long-Term Contract
Description:
As a Software Engineer on the Cloud Engineering team, you will design, develop, and maintain the cloud platform, infrastructure, and automation tooling that powers our connected vehicle ecosystem. You will play a key role in improving reliability, scalability, security, and cost efficiency across a platform serving millions of vehicles worldwide.
Key Responsibilities
Platform Development & Automation
- Design and build cloud-native solutions that enable automated, repeatable shard provisioning.
- Develop tooling for:
- Kubernetes cluster bootstrapping
- Network configuration
- Service mesh deployment
- Health validation and readiness checks
- Build control plane capabilities that allow a new shard to be provisioned and production-ready within 24 hours with zero manual intervention.
Infrastructure as Code (IaC)
- Develop and maintain production-grade Terraform modules across AWS and GCP environments.
- Manage GitOps-based infrastructure workflows utilizing Atlantis or similar tools.
- Ensure infrastructure changes are properly reviewed, tested, and fully auditable before production deployment.
Kubernetes & Service Mesh Engineering
- Provision, manage, and automate Kubernetes clusters in AWS EKS and Google GKE environments.
- Create and maintain Helm charts supporting multi-shard deployments across regions and environments.
- Configure and operate Istio service mesh capabilities, including:
- Traffic management
- mTLS enforcement
- Observability and monitoring
Google Cloud Platform (GCP) Engineering
- Design and operate GCP-native platform services, including:
- Google Kubernetes Engine (GKE)
- Pub/Sub
- Cloud Storage
- Secret Manager
- Workload Identity Federation
- VPC Service Controls (VPC-SC)
- Support the phased migration of platform services from AWS to GCP.
API Gateway & Ingress Management
- Design, configure, and operate API gateway technologies such as Tyk, Apigee, or NGINX.
- Implement routing policies, rate limiting, mTLS enforcement, and SLI/SLO metric replication.
- Execute zero-downtime ingress migrations through automation and validation processes.
CI/CD & Deployment Automation
- Develop and maintain CI/CD pipelines using:
- ArgoCD
- Tekton
- Concourse
- Similar GitOps tools
- Automate deployment controls including:
- Cluster health validation
- Fleet distribution verification
- Rollback triggers
- Migration orchestration
Security & Compliance
- Implement cloud security best practices across AWS and GCP.
- Manage:
- IAM role isolation
- Least-privilege access controls
- Workload identity
- Automated access reviews
- Monitor and remediate:
- Certificate expiration risks
- Security policy drift
- Cloud quota thresholds
Observability & Monitoring
- Design and maintain monitoring, alerting, and dashboarding solutions using:
- Prometheus
- Grafana
- Datadog
- Google Cloud Operations
- Provide visibility into:
- Platform health
- Fleet distribution metrics
- Service reliability and SLO compliance
Automation & Tool Development
- Develop automation scripts and internal tooling using:
- Build tools supporting:
- Migration orchestration
- Fleet validation
- Runbook automation
- Capacity planning
Operational Excellence
- Participate in on-call support rotations.
- Troubleshoot production incidents and platform issues.
- Apply Site Reliability Engineering (SRE) principles to improve system resilience.
- Drive root cause analysis and permanent corrective actions.
Required Qualifications
Education
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
Experience
- 5+ years of professional experience in Cloud Engineering, Platform Engineering, DevOps, SRE, or Software Engineering.
- Experience building and operating production infrastructure at scale.
- 3+ years of experience managing cloud infrastructure in AWS and/or GCP.
Technical Skills
- Hands-on experience provisioning and managing Kubernetes clusters (EKS or GKE) in production environments.
- Strong experience with Terraform and Infrastructure as Code (IaC) best practices.
- Experience with GitOps workflows and infrastructure automation.
- Proficiency developing and maintaining Helm charts.
- Strong troubleshooting and problem-solving skills in cloud and distributed systems environments.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or similar solutions.