Software Development Engineer Infra - 1
Job type: Full Time · Department: Engineering · Work type: On-Site
Bellandur, Karnataka, India
At Cashfree, infrastructure is not just about running servers — it is a software platform that powers our engineering organization.
As a Software Development Engineer Infra - 1, you will build the systems, platforms, automation, and developer tools that enable hundreds of engineers to build and operate products reliably at scale.
This is a hands-on engineering role at the intersection of software engineering, distributed systems, cloud infrastructure, Kubernetes, reliability, security, developer productivity, and AI.
You will design and build production software — APIs, backend services, controllers, CLI tools, automation frameworks, infrastructure abstractions, internal developer platforms, and AI-powered engineering tools.
AI-powered tools for troubleshooting, automation, developer productivity, and infrastructure operations
Internal developer platforms and self-service infrastructure
Infrastructure APIs, control planes, and automation
Kubernetes operators and infrastructure abstractions
Deployment, provisioning, and environment orchestration
Production debugging, observability, and incident-response tools
Security, governance, and cost-optimization automation
The goal is to make infrastructure self-service, reliable, secure, efficient, and increasingly intelligent.
Build APIs, services, CLIs, and tools that simplify infrastructure consumption for engineering teams
Develop self-service workflows for service onboarding, deployments, environments, and infrastructure
Create reusable abstractions over cloud and Kubernetes
Build AI-assisted developer and infrastructure workflows that reduce operational effort and improve engineering productivity
Own platform features end-to-end from design to production
Build reusable Terraform/Pulumi modules and infrastructure frameworks
Develop Kubernetes operators, controllers, and platform components
Design highly available multi-AZ and multi-region infrastructure
Automate provisioning, scaling, failover, and infrastructure lifecycle management
Explore AI-driven automation for infrastructure provisioning, optimization, and operational decision-making
Design systems for availability, failure isolation, graceful degradation, and automated recovery
Define and improve SLOs, SLIs, error budgets, and reliability metrics
Build observability, diagnostics, and production troubleshooting capabilities
Build AI-assisted incident detection, triage, root-cause analysis, and remediation workflows
Drive improvements to MTTR, capacity, performance, and operational toil
Build IAM, secrets management, policy, and compliance automation
Embed security and governance controls into infrastructure workflows
Build tooling to identify infrastructure waste and automate resource optimization
Explore AI-assisted detection of security, configuration, and cost anomalies
Balance cost, performance, reliability, and developer productivity
Investigate and solve challenging problems across distributed systems, networking, Kubernetes, and cloud infrastructure
Turn recurring operational problems into scalable engineering solutions
Identify opportunities to apply AI/LLMs to automate complex engineering workflows and reduce manual operations
Work closely with product engineering and security teams to design platform capabilities
Take ownership of systems through design, implementation, deployment, and production operations
1+ years of experience in software engineering, backend engineering, platform engineering, SRE, or a similar role
Strong programming skills in Go, Python, Java, Rust, or a similar language
Strong understanding of software design, APIs, concurrency, data structures, and distributed systems
Experience building and operating production-grade services, tools, or automation
Strong debugging and problem-solving skills
Strong understanding of Linux, networking, cloud infrastructure, and distributed systems
Hands-on experience with AWS, GCP, or Azure
Experience with infrastructure automation and Infrastructure as Code
Strong understanding of Kubernetes and containerized workloads
Experience with Terraform, Pulumi, OpenTofu, or similar tools
Understanding of SLOs, SLIs, observability, scalability, and failure recovery
Experience with tools such as Prometheus, Grafana, Datadog, OpenTelemetry, or equivalent
Understanding of IAM, secrets management, authentication, authorization, and least privilege
Experience with Kubernetes operators/controllers, GitOps, Vault, policy-as-code, or cloud cost optimization is a plus
Experience building applications or internal tools using LLMs, AI agents, or AI APIs
Understanding of LLM-based automation, tool calling, RAG, or agentic workflows
Experience integrating AI into developer productivity, observability, incident management, or infrastructure workflows
Familiarity with MCP or similar approaches for connecting AI systems to engineering tools and infrastructure is a plus
We look for engineers who build instead of manually operate, automate repetitive work, think in APIs and abstractions, design for failure and scale, and take ownership of systems end-to-end.
We also value engineers who are curious about using AI to fundamentally improve how infrastructure is built, operated, and debugged.
Build the platform. Automate the work. Make every engineer faster.
Autofill from resume
Save time by uploading your resume. (Only PDF or DOCX format supported)