Senior Kubernetes Engineer (SME)

Work type: Full time · Department: Engineering · Workplace: On-Site

Mumbai, Maharashtra, India

About the Role:

We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack - from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.

What you will be doing:

•      Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.

•      Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning

•      Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)

•      Configure and troubleshoot CNI plugins (Calico, Cilium) — pod networking, network policies, BGP, and eBPF dataplanes.

•      Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.

•      Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups

•      Deploy and operate service mesh (Istio, Cilium mesh) for traffic management

•      Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)

•      Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting

•      Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery

•      Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning

•      Own backup and disaster recovery (Velero, etcd snapshotting) and run DR drills

•      Provide on-call production support: monitor cluster health, troubleshoot incidents, and drive root-cause resolution

•      Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer

What we need to see:

•          Core Kubernetes & Infrastructure (5+ years)

•          Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet

•          Proven experience designing, deploying, and troubleshooting production clusters at scale

•          Hands-on with multi-cluster/multi-tenant tooling — Rancher, Kamaji, and/or vCluster (strongly preferred)

•          CNI expertise: Calico, Cilium — network policy design, BGP, VXLAN/IPIP encapsulation, eBPF

•          Container runtime internals: containerd, CRI-O, runc — configuration and troubleshooting

•          Strong Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management

•          Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS

•          Service mesh experience: Istio, Linkerd, or Cilium service mesh

•          CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments

•          Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)

Automation & Infrastructure-as-Code:

•          Helm chart authoring and lifecycle management

•          GitOps workflows: ArgoCD or FluxCD

•          IaC and configuration management: Terraform, Ansible

•          CI/CD pipeline integration for cluster and application delivery

•          Scripting proficiency: Bash and Python (Go a plus)

Ways to stand out from the rest:

•          Kubernetes certifications: CKA, CKAD, CKS

•          Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)

•          Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal

•          Experience with GPU-enabled clusters and AI/ML workload scheduling '

•          Familiarity with AI/LLM serving stacks on Kubernetes — vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning

•          General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints

•          Contributions to CNCF projects or an active open-source/GitHub presence

Minimum Qualifications:

•          Bachelor’s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)

•          5+ years of hands-on production Kubernetes experience

•          Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments

•          Solid grounding in containerd/OS-level troubleshooting

•          Production on-call and incident-response experience.

Soft Skills:

•          Strong problem-solving and debugging abilities under production pressure

•          Ownership mindset with accountability for platform reliability and stability

•          Cross-functional collaboration with application, security, operations and platform teams

•          Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem

•          Strong documentation and communication skills, with ability to defend design decisions in peer reviews

 

Made with

Senior Kubernetes Engineer (SME) – Mumbai | Neysa Careers