SRE · DEVOPS · PLATFORM ENGINEERING

Reliable systems.
Clear evidence.

I’m Anil Kumar Tangirala, a Senior DevOps and Service Reliability Engineer building cloud platforms, automation, and dependable production delivery.

Hyderabad, India · SRE / DevOps at Oracle

Full-length professional portrait of Anil Kumar Tangirala wearing a navy blazer, ivory shirt and classic watch
Anil Kumar TangiralaExecutive portrait

01 / ABOUT

I build the calm
behind production.

Eight years across cloud, SRE, and DevOps—turning delivery paths, operational evidence, and incident response into systems teams can trust.

8Years of platform
engineering
4Cloud platforms
in practice
99%Incidents resolved
within SLA at Oracle

02 / EXPERIENCE

Three chapters.
One operational mindset.

2018 — 2023

CLOUD & DEVOPS ENGINEERING

Artisto
Technologies

DevOps Engineer · Vizag

MISSION

Modernize delivery and infrastructure for e-commerce systems across AWS, Azure, and GCP.

RESPONSIBILITIES
  • Built CI/CD and GitHub Actions workflows
  • Provisioned EC2, S3, RDS and multi-cloud infrastructure
  • Automated configuration with Terraform, Ansible and Python
  • Built Docker and Kubernetes deployment workflows
+60% delivery efficiency−35% deployment errors−50% deployment time
01
2023 — 2025

RELEASE & SERVICE RELIABILITY

Mobile Programming
India

DevOps Engineer · Bangalore

MISSION

Make AWS and OCI healthcare environments repeatable to deploy, support, and recover.

RESPONSIBILITIES
  • Built GitLab CI/CD pipelines for AWS and OCI delivery
  • Led EC2-to-EKS migration and standardized OKE rollout checks
  • Standardized incident diagnostics with SOPs and RQ logs
  • Automated server updates, builds, and deployments
+40% deployment efficiency−30% incident response time−50% manual work
02
2025 — NOW

SRE / PLATFORM RELIABILITY

Oracle
India

Senior DevOps Engineer (SRE) · Hyderabad

MISSION

Own production reliability, releases, diagnostics, and recovery across multi-region Kubernetes/OKE SaaS platforms.

RESPONSIBILITIES
  • Lead incidents, changes, releases, and on-call operations
  • Deliver RQ, upgrades, and controlled deployments across ECP, ECP-RT, OSDMC and ICON
  • Build log surveillance, dashboards, SLO signals, and alerting for production services
  • Develop and review infrastructure and deployment automation
−50% time to detect−30% time to resolve≥75% reduction in triage time
03

03 / CAREER PROJECTS

Named projects.
Concrete engineering work.

Earlier engagements at Artisto, drawn from my original project-based resume. Current capabilities and outcomes follow my latest SRE resume.

OCT 2022 — APR 2023 / ARTISTO

FDA Platform Development (HCLS)

AWS platform infrastructure and delivery for healthcare workloads.

  • Designed UAT and production environments with EC2, S3, RDS and Elastic Beanstalk; provisioned microservices infrastructure with Terraform.
  • Implemented AWS Config, Security Hub, WAF and IAM controls; automated administration with Python, Bash and Ansible.
  • Built Jenkins, Docker and Kubernetes delivery pipelines, reducing deployment time by 50%; implemented ELK, Prometheus and CloudWatch monitoring.

AWS · Terraform · Jenkins · Docker · Kubernetes · Ansible

AUG 2021 — SEP 2022 / ARTISTO

Ebates Online Shopping Portal

Weekly UAT and production releases for an e-commerce platform.

  • Architected AWS infrastructure using VPC, EC2, ECS, SQS, SNS and S3; provisioned MySQL and Aurora on RDS with Terraform.
  • Built Jenkins CI infrastructure and Docker build agents; maintained registries, image builds and Tomcat/Apache deployments.
  • Coordinated release branches, build requirements, smoke checks, patches and urgent hotfixes with developers and QA.

AWS · VMware · Terraform · Jenkins · Docker · Tomcat

SEP 2020 — JUL 2021 / ARTISTO

Sentinel EMS

Build infrastructure, CI/CD and production deployment support.

  • Implemented Jenkins and Git integration; managed repositories, branches and collaboration practices to reduce merge conflicts.
  • Maintained Maven, SonarQube and Ansible delivery tooling across development, test and production.
  • Automated repetitive operations with Shell and Python; maintained log backups and responded to high-severity escalations.

Git · Jenkins · Maven · Ansible · AWS · Python

JUL 2019 — AUG 2020 / ARTISTO

JFK Health System

Healthcare deployment automation and resilient AWS application hosting.

  • Configured IAM policies, EC2 security groups, ELB and Route 53 failover and latency-based routing.
  • Built Docker/AWS delivery pipelines, Kubernetes test environments and Tomcat hosting; integrated Git, Jenkins and Nexus.
  • Automated remote configuration with Ansible and supported emergency releases and on-call maintenance.

AWS · Kubernetes · Docker · Jenkins · Ansible · Tomcat

OCT 2018 — JUN 2019 / ARTISTO

Key Performance Indicators

Build and Release Engineer responsibilities for repeatable application delivery.

  • Automated Maven builds and Jenkins nightly, release and production-branch jobs.
  • Configured distributed Jenkins on Linux/Unix build machines with plugins and Apache Tomcat.
  • Automated promotion across development, QA, staging and production; coordinated build releases and bug reviews.

Maven · Jenkins · Linux · Shell · Ansible · Tomcat

ADDITIONAL ARTISTO EXPERIENCE

Artrya ML Platform Infrastructure

Infrastructure engineering for a healthcare machine-learning platform.

  • Engineered and deployed supporting infrastructure and integrated Jupyter into deployment workflows.
  • Owned the infrastructure and deployment layer supporting the ML platform.

Jupyter · Platform engineering · Infrastructure automation

CURRENT / ARCHITECTURE TO EXECUTION

Production ownership,
from code to recovery.

INFRASTRUCTURE & REVIEW

Platform architecture

Terraform, Python, Bash and OCI Resource Manager automation for multi-region, multi-tenant platforms. Personal ownership of infrastructure and deployment code, pipeline reviews, IAM boundaries and configuration-drift checks.

DELIVERY & QUALIFICATION

Controlled releases

Reusable GitLab templates, security checks and promotion gates. OCI DevOps and Shepherd delivery, Playwright release qualification, schema-change coordination, telemetry validation and rollback signals.

OBSERVABILITY & AI

Evidence-led response

Grafana SLO/SLI views, Streamlit uptime and cost dashboards, OpenSearch monitors and OpenTelemetry. MCP correlation and custom Codex workflows support investigation with human-approved production decisions.

RESILIENCE & CONNECTIVITY

Recovery strategy

Multi-region disaster-recovery runbooks, replication validation and failover tests. LiveKit deployments, WebRTC troubleshooting, and automated jitter and packet-loss analysis for real-time services.

04 / ORIGINAL PROJECTS

Tools for the
critical moment.

Original, runnable reference projects: deterministic operational checks, bounded retrieval, citation validation, and human-reviewed decisions.

01 / INCIDENT RESPONSE

SignalDesk

Interactive incident workspace for SLO analysis, evidence preflight, dependency impact, logs, ownership and runbook retrieval.

14 workflowsPython ASTRAG
OPEN PROJECT ↗
02 / KUBERNETES

Platform
Readiness RAG

Workload-specific reviews connected to inspectable platform standards and cited operational guidance.

Policy checksOfflineJSON
OPEN PROJECT ↗
03 / TERRAFORM

Change
Risk RAG

Curated findings for destructive changes, public ingress, protection removal, and uncertainty—without raw plan values.

Risk rulesRunbooksSafe review
OPEN PROJECT ↗

05 / HOW I WORK

Make production
safer to change.

01

Investigate

Correlate service symptoms, application logs, platform health, database signals, and deployment context to turn an alert into a testable incident hypothesis.

02

Review readiness

Review rollout plans, MOPs, SOPs, pipeline inputs, environment assumptions, failure paths, recovery steps, and operational documentation before production change.

03

Execute & recover

Coordinate release qualification, controlled upgrades, validation, triage, and corrective action across engineering, QA, infrastructure, security, and support teams.

04

Harden the path

Convert recurring gaps into guardrails: monitoring, alert tuning, credential-expiry awareness, safer secret handling, dry-run expectations, and reusable runbooks.

06 / TOOLKIT

CLOUD

AWS · OCI · AZURE · GCP

SHIP

GITLAB CI/CD · GITHUB ACTIONS · JENKINS · ARGO CD

RUN

KUBERNETES · TERRAFORM · ANSIBLE · HELM · ISTIO

SEE

PROMETHEUS · GRAFANA · OPENSEARCH · OPENTELEMETRY

LEARN

PYTHON · BASH · MCP · CUSTOM CODEX SKILLS

07 / AI-ASSISTED OPERATIONS

AI assistance
with operational guardrails.

I build evidence-based diagnostic workflows and evaluate LLM-assisted investigation, intelligent runbooks, and predictive reliability. Production decisions remain human validated.

AWS SOLUTIONS ARCHITECT — ASSOCIATE · ORACLE DATABASE@AWS ARCHITECT PROFESSIONAL · OCI OBSERVABILITY PROFESSIONAL · OCI MULTICLOUD ARCHITECT PROFESSIONAL · ORACLE AI AUTONOMOUS DATABASE CLOUD PROFESSIONAL · OCI GENERATIVE AI PROFESSIONAL · ORACLE AI VECTOR SEARCH PROFESSIONAL · GLOBAL PRODUCT SECURITY & AI SECURITY ADVANCED (ORACLE) · GENAI FOR DEVOPS PRACTITIONERS (COURSERA)

08 / LET’S CONNECT

Let’s make
reliability visible.

Hyderabad, IndiaLinkedIn ↗GitHub ↗