Reliability is a feature.
I engineer it.
Nine years, three domains, one obsession: production that stays up.
I started in 2016 wearing every hat at once — infrastructure, networks, firewalls, releases — the kind of environment where if something broke, you fixed it, because there was nobody else. That built an instinct I still rely on: understand the whole system, not just your layer.
From there I went deep into cloud-native: designing and running Kubernetes platforms on AWS EKS and GCP GKE, building GitOps pipelines with ArgoCD, Helm, Jenkins and GitHub Actions, wiring service meshes, and hardening supply chains for PCI, HIPAA and SOC compliance — while mentoring engineers along the way.
Today I lead observability for live Mutual Fund & Equity trading platforms — where seconds of latency are measured in money. I define the SLOs and SLIs, build the Splunk and Dynatrace dashboards leadership actually uses, and run the RCAs when things get interesting.
Platform builder
Multi-cloud Kubernetes, GitOps delivery, service mesh — designed and run in production, not in a lab.
Observability-first SRE
SLOs, error budgets, synthetic monitoring and predictive dashboards that catch issues before users do.
Security & compliance aware
Image scanning, hardened pipelines, PCI / HIPAA / SOC support — shipped fast without cutting corners.
Calm under fire
Incident response and RCA in high-pressure financial environments. Trading hours don't wait.
Tech stack, battle-tested.
Every tool below has survived real production incidents with me.
Cloud Platforms
Containers & Orchestration
CI/CD & GitOps
Observability & SRE
Automation & IaC
Logging, Data & Network
Nine years of keeping the lights on.
From bare-metal networks to multi-cloud Kubernetes to real-time trading observability.
- Own reliability for live Mutual Fund & Equity trading platforms in a high-pressure financial environment.
- Designed and implemented SLOs / SLIs for latency, availability and service health across critical trading services.
- Built advanced Splunk dashboards with predictive insights — real-time monitoring for engineering and business visibility.
- Created Dynatrace health/performance dashboards and proactive alerting, cutting response times and downtime.
- Built synthetic monitors automating user-journey validation; drive RCA and continuous observability improvements.
- Designed DevOps solutions for multiple production-grade environments; managed Kubernetes on AWS EKS & GCP GKE.
- Built CI/CD with Jenkins & GitHub Actions; automated deployments via ArgoCD + Helm (GitOps).
- Implemented Prometheus/Grafana monitoring and centralized logging with ELK / EFK / Graylog.
- Ran Istio and Nginx Ingress traffic management; image vulnerability scanning with Clair & Anchore.
- Led cloud cost optimization and mentored engineers on DevOps practices and architecture.
- Built and deployed Java applications with Maven / Gradle on Jenkins; managed Tomcat and production releases.
- Designed pfSense firewall with dual-WAN failover and OpenVPN for secure remote access.
- Managed DNS and load balancing with F5 BIG-IP; network monitoring via ntopng.
- Owned patching, upgrades and security hardening across the estate with minimal downtime.
Problems solved, at production scale.
Selected work — architecture, problem, impact.
Trading Observability Command Center
Mutual Fund & Equity trading services had fragmented visibility — incidents were detected reactively, often by users, in an environment where latency = money.
Defined SLOs/SLIs for latency, availability and health. Built predictive Splunk dashboards, Dynatrace performance boards, proactive alerting and synthetic user-journey monitors.
▲ Issues caught before users notice · faster incident response · single pane of glass for engineering + business.
Multi-Cloud Kubernetes GitOps Platform
Multiple production environments across AWS and GCP with manual, drift-prone deployments and inconsistent monitoring.
Standardized on EKS + GKE with ArgoCD + Helm GitOps delivery, Jenkins/GitHub Actions CI, Istio traffic management, Prometheus/Grafana monitoring and centralized ELK/EFK logging.
▲ Deployment frequency · zero config drift (Git as source of truth) · consistent operations across clouds.
Secure CI/CD Supply Chain (DevSecOps)
Container images shipped to regulated environments (PCI, HIPAA, SOC) without automated vulnerability gates — audit risk and manual review bottlenecks.
Embedded Clair and Anchore image scanning into Jenkins/GitHub Actions pipelines, enforced policy gates pre-deploy, and hardened the release path to support compliance initiatives.
▲ Vulnerabilities blocked before production · audit-ready pipelines · compliance shipped without slowing delivery.
High-Availability Network & Release Infrastructure
Single points of failure across internet uplinks, remote access and load balancing threatened business continuity.
Designed pfSense firewall with dual-WAN failover, OpenVPN remote access, F5 BIG-IP load balancing and DNS, ntopng network monitoring, plus hardened patching and release processes.
▲ Survived uplink failures with zero business interruption · secure remote workforce · minimal-downtime releases.
Certified across cloud & network.
Validated by Google, Amazon and Cisco.
Professional Cloud Architect
Solutions Architect — Associate
CCNA
AWS Associate Training
Numbers that move the needle.
What nine years of reliability engineering looks like.
Talk to my terminal.
Recruiter-friendly CLI. Try help, skills or sudo hire-me.
What teams say about working with me.
From engineering peers, leads and stakeholders.
When trading hours are live and something looks off, Rajni's dashboards are the first place everyone looks — and usually the reason we caught it early.
Platform Manager
Trading / Capital Markets
He didn't just set up our Kubernetes clusters — he made deployments boring. GitOps, monitoring, alerts: everything just works.
Engineering Lead
Cloud Platform Team
Calm in incidents, thorough in RCAs, generous as a mentor. The engineers he coached still follow his runbooks.
Delivery Manager
Fintech Programs
Let's build something reliable.
Open to DevOps, SRE, Platform and Cloud engineering roles. Response SLO: < 24 hours.