Skip to content
Jobsearch.ing

Senior Platform Engineer, Cloud Infrastructure

JobgetherRemote · CA

EngineeringSeniorFull time6+ yrs
Source-verified: read directly from this employer's own lever job board, not a repost.RemotePosted (3 days ago)Last verified (today)

At a glance

Location
Remote · CA
Workplace
Remote
Pay
Not published by the employer
Employment type
Full time
Experience
6+ years
Education
No degree requirement stated
Job family
Engineering
Seniority
Senior
Posted by employer
4 September 2026
Last verified open
7 September 2026
Work from
CA
Region and country
CA
Team
Security & IT
Listed via
Lever

What the employer wrote

Accountabilities • Design, build, and operate production Kubernetes clusters, including networking, workload isolation, resource management, scheduling, and multi-region architectures.

• Work deeply with Kubernetes internals, including CNI networking, NetworkPolicy enforcement, cluster behavior under load, resource quotas, and custom controllers or operators.

• Design and operate service mesh capabilities, including mTLS, workload identity, service-account authentication and authorization, and traffic management.

• Optimize containerized workloads for performance, cost efficiency, scalability, and effective resource utilization.

• Develop and maintain production-grade services, Kubernetes controllers, middleware, and platform components using Go, Python, or Java.

• Build HTTP, REST, and gRPC interfaces used by internal engineering teams, while maintaining strong unit and integration test coverage.

• Lead platform-level incident response, troubleshoot distributed systems using logs, metrics, traces, and profiling, and produce postmortems that drive lasting improvements.

• Define and implement SLOs, actionable alerting, dashboards, metrics, and distributed tracing to strengthen platform reliability.

• Own infrastructure as code at scale, creating reusable modules and evolving them as platform requirements change.

• Build and improve CI/CD and GitOps workflows that enable engineering teams to release software safely, frequently, and reliably.

• Plan and execute cloud migration initiatives, including moving production workloads between providers or environments while minimizing disruption and maintaining reliability.

• Partner with product, security, and infrastructure teams to gather requirements, evaluate trade-offs, conduct design reviews, and establish technical direction.

• Mentor engineers and contribute to higher standards of engineering, reliability, testing, and operational excellence.

Requirements

• 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including substantial experience operating production distributed systems.

• Demonstrated experience  building and operating production Kubernetes platforms , rather than simply deploying applications onto existing clusters.

• Strong production programming experience in  Go, Python, or Java , with the ability to work effectively in complex existing codebases.

• Experience taking ambiguous technical problems from initial design through production implementation and ongoing operation.

• Proven experience planning and executing cloud migrations involving production workloads across providers or environments.

• Degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.

• Strong understanding of Kubernetes networking, CNI, NetworkPolicy, resource management, scheduling, and cluster behavior.

• Hands-on experience with service mesh technologies such as Istio, Envoy, Linkerd, or equivalent, including mTLS and workload identity.

• Strong Linux fundamentals, including cgroups and resource management.

• Significant infrastructure-as-code experience using Terraform or an equivalent technology.

• Production experience with at least one major cloud platform such as AWS, GCP, or Azure; multi-cloud experience is highly valuable.

• Experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, and PromQL or comparable query languages.

• Production experience with relational databases, particularly PostgreSQL or managed PostgreSQL-compatible services, including an understanding of replication and failover.

• Strong Docker and container tooling experience across the software delivery lifecycle.

• Advanced debugging, troubleshooting, and performance-profiling capabilities.

• Experience with Kubernetes controllers, operators, API-server extensions, or Go testing frameworks such as Ginkgo and Gomega is an asset.

• Additional Python expertise and the ability to read Java are valuable.

• Experience with identity and access technologies such as SSO, Keycloak, OIDC, SAML, Vault, or cloud-based secrets management is preferred.

• Experience designing high-availability and disaster-recovery architectures across regions or cloud providers is an advantage.

• Familiarity with monorepos, Bazel, compliance frameworks such as SOC 2 or GDPR, and reliability considerations for LLM-backed systems is beneficial.

• Experience with Alibaba Cloud and large-scale data technologies such as Apache Spark or Apache Flink is highly preferred.

• Strong analytical and problem-solving skills, excellent written and verbal communication, and the ability to produce clear design documents and postmortems.

• Comfortable working independently in a distributed, fast-moving, highly technical environment while collaborating effectively across teams.

Benefits

• Fully remote position open to candidates in Canada.

• Senior-level opportunity with significant technical ownership and influence over cloud infrastructure.

• Opportunity to work with Kubernetes, cloud platforms, service mesh, observability, infrastructure as code, and modern software engineering practices.

• Exposure to complex, multi-quarter initiatives including cloud migrations and large-scale reliability improvements.

• Collaborative distributed environment with opportunities to mentor engineers and shape engineering standards.

• Flexible remote collaboration across an international technical team.

• Opportunity to contribute to innovative cloud-native infrastructure and emerging technologies, including AI-enabled systems.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether?    Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.     #LI-CL1

Where this record came from

Read from Jobgether's own Lever job board on , and last confirmed still open on . The employer published it on 4 September 2026. Jobsearch.ing did not write, edit or rank this posting, and does not vet the employer. View the original posting.

More roles like this one