Software Development Engineer Infra - 1

Job type: Full Time · Department: Engineering · Work type: On-Site

Bellandur, Karnataka, India

About the Role

At Cashfree, infrastructure is not just about running servers — it is a software platform that powers our engineering organization.

As a Software Development Engineer Infra - 1, you will build the systems, platforms, automation, and developer tools that enable hundreds of engineers to build and operate products reliably at scale.

This is a hands-on engineering role at the intersection of software engineering, distributed systems, cloud infrastructure, Kubernetes, reliability, security, developer productivity, and AI.

You will design and build production software — APIs, backend services, controllers, CLI tools, automation frameworks, infrastructure abstractions, internal developer platforms, and AI-powered engineering tools.

What You Will Build

  • AI-powered tools for troubleshooting, automation, developer productivity, and infrastructure operations

  • Internal developer platforms and self-service infrastructure

  • Infrastructure APIs, control planes, and automation

  • Kubernetes operators and infrastructure abstractions

  • Deployment, provisioning, and environment orchestration

  • Production debugging, observability, and incident-response tools

  • Security, governance, and cost-optimization automation

The goal is to make infrastructure self-service, reliable, secure, efficient, and increasingly intelligent.

Core Responsibilities

1. Build Internal Infrastructure Platforms

  • Build APIs, services, CLIs, and tools that simplify infrastructure consumption for engineering teams

  • Develop self-service workflows for service onboarding, deployments, environments, and infrastructure

  • Create reusable abstractions over cloud and Kubernetes

  • Build AI-assisted developer and infrastructure workflows that reduce operational effort and improve engineering productivity

  • Own platform features end-to-end from design to production

2. Build Cloud & Kubernetes Infrastructure

  • Build reusable Terraform/Pulumi modules and infrastructure frameworks

  • Develop Kubernetes operators, controllers, and platform components

  • Design highly available multi-AZ and multi-region infrastructure

  • Automate provisioning, scaling, failover, and infrastructure lifecycle management

  • Explore AI-driven automation for infrastructure provisioning, optimization, and operational decision-making

3. Engineer Reliability & Production Systems

  • Design systems for availability, failure isolation, graceful degradation, and automated recovery

  • Define and improve SLOs, SLIs, error budgets, and reliability metrics

  • Build observability, diagnostics, and production troubleshooting capabilities

  • Build AI-assisted incident detection, triage, root-cause analysis, and remediation workflows

  • Drive improvements to MTTR, capacity, performance, and operational toil

4. Build Secure & Efficient Infrastructure

  • Build IAM, secrets management, policy, and compliance automation

  • Embed security and governance controls into infrastructure workflows

  • Build tooling to identify infrastructure waste and automate resource optimization

  • Explore AI-assisted detection of security, configuration, and cost anomalies

  • Balance cost, performance, reliability, and developer productivity

5. Solve Complex Infrastructure Problems

  • Investigate and solve challenging problems across distributed systems, networking, Kubernetes, and cloud infrastructure

  • Turn recurring operational problems into scalable engineering solutions

  • Identify opportunities to apply AI/LLMs to automate complex engineering workflows and reduce manual operations

  • Work closely with product engineering and security teams to design platform capabilities

  • Take ownership of systems through design, implementation, deployment, and production operations


Job Requirements

Software Engineering

  • 1+ years of experience in software engineering, backend engineering, platform engineering, SRE, or a similar role

  • Strong programming skills in Go, Python, Java, Rust, or a similar language

  • Strong understanding of software design, APIs, concurrency, data structures, and distributed systems

  • Experience building and operating production-grade services, tools, or automation

  • Strong debugging and problem-solving skills

Cloud & Infrastructure

  • Strong understanding of Linux, networking, cloud infrastructure, and distributed systems

  • Hands-on experience with AWS, GCP, or Azure

  • Experience with infrastructure automation and Infrastructure as Code

  • Strong understanding of Kubernetes and containerized workloads

  • Experience with Terraform, Pulumi, OpenTofu, or similar tools

Reliability & Security

  • Understanding of SLOs, SLIs, observability, scalability, and failure recovery

  • Experience with tools such as Prometheus, Grafana, Datadog, OpenTelemetry, or equivalent

  • Understanding of IAM, secrets management, authentication, authorization, and least privilege

  • Experience with Kubernetes operators/controllers, GitOps, Vault, policy-as-code, or cloud cost optimization is a plus

AI & Automation — Nice to Have

  • Experience building applications or internal tools using LLMs, AI agents, or AI APIs

  • Understanding of LLM-based automation, tool calling, RAG, or agentic workflows

  • Experience integrating AI into developer productivity, observability, incident management, or infrastructure workflows

  • Familiarity with MCP or similar approaches for connecting AI systems to engineering tools and infrastructure is a plus


What We Value

We look for engineers who build instead of manually operate, automate repetitive work, think in APIs and abstractions, design for failure and scale, and take ownership of systems end-to-end.

We also value engineers who are curious about using AI to fundamentally improve how infrastructure is built, operated, and debugged.

Build the platform. Automate the work. Make every engineer faster.

Made with