Architecture that operates.
Access, blast radius, recovery and ownership belong in the design—not in the postmortem.
SRE / DEVOPS / PLATFORM ENGINEERING
ANIL KUMAR TANGIRALA
I build the infrastructure, delivery systems and operational clarity that help teams make change with confidence.

Senior DevOps Engineer · Oracle
SRE / DEVOPS / PLATFORM ENGINEERINGFrom the first Terraform module to the last recovery check, I connect the decisions that make a platform dependable.
I work across cloud architecture, infrastructure automation, release engineering and production reliability. My journey spans e-commerce, healthcare and multi-region SaaS.
Hands-on infrastructure and deployment code ownership. Thoughtful reviews. Human-validated production decisions.
Read my complete résumé ↗Access, blast radius, recovery and ownership belong in the design—not in the postmortem.
Reusable pipelines, explicit qualification and deliberate promotion gates.
Separate observed facts, hypotheses and unknowns. Validate recovery through service behavior.
Production responsibilities and original reference projects.
Architecture, implementation and operational impact.
My current Oracle work spans ECP, ECP-RT, OSDMC and ICON—connecting infrastructure code, controlled delivery and day-to-day reliability.
Develop and review Terraform, Python, Bash and GitLab CI/CD code. Automate provisioning and OKE lifecycle operations with OCI Resource Manager across multi-region, multi-tenant platforms.
Cloud platform & IoT connectivity
Infrastructure automation and operational support for a platform that connects cloud services with real-world devices.
I develop Terraform, Python and Bash automation for provisioning and OKE lifecycle management. My work includes reviewing infrastructure and delivery code for multi-region, multi-tenant environments; application-code changes remain with the relevant developers.
I troubleshoot ECP IoT connectivity across carrier and satellite networks and automate jitter and packet-loss analysis. The focus is on distinguishing network behavior from cloud or workload symptoms, rather than treating every failure as a deployment issue.
I deploy LiveKit and troubleshoot WebRTC across ECP and ECP-RT. This complements the infrastructure work with service-level investigation.
Execute release qualification across cloud services, infrastructure components and edge artifacts, combining pipeline dry runs with tenant-lifecycle validation to prepare changes for controlled promotion.
Carry out procedure-led database migration and credential-rotation operations alongside platform image upgrades. Validate release prerequisites through health analysis, credential-expiry checks and monitoring-rule maintenance, coordinating application-level changes with the relevant owners.
Real-time communications & release operations
Production delivery and qualification alongside hands-on LiveKit and WebRTC deployment and troubleshooting.
I participate in platform deployments and upgrades across environments, with responsibility for release execution and qualification. Release completion includes qualification, not simply a successful pipeline status.
Post-upgrade activation checks and qualification tasks expose integration issues that require coordination with product engineering. I investigate failures and work with developers on defect resolution in partnership with the application owners.
I deploy LiveKit and troubleshoot WebRTC behavior. Jitter and packet-loss analysis connects the user-facing communication symptom to the underlying network and service evidence.
Execute patch upgrades and multi-region production deployments, verify scheduled OCI compute maintenance, and carry releases through post-upgrade activation qualification. Connect infrastructure lifecycle work with service readiness rather than treating pipeline completion as the end of a release.
Platform automation & operational monitoring
Cloud infrastructure engineering backed by recurring monitoring checks and production-service investigation.
Develop and review infrastructure and deployment automation using Terraform, Python, Bash and OCI Resource Manager, supporting provisioning and OKE lifecycle operations within the shared platform-engineering remit.
Verify detection rules as part of recurring service operations, complementing dashboards with checks on the monitoring controls that support incident detection.
Investigate identity-service connectivity alerts, correlate the available operational evidence and coordinate recovery with the relevant service owners.
Perform tenant storage-capacity changes with pre-change validation, tenant-status checks and post-change verification. Bring the same controlled execution discipline to capacity operations as to service releases.
Investigate partner-token generation failures and identity-service connectivity alerts, using operational evidence to narrow the affected integration and coordinate the next recovery action.
Infrastructure code & Kubernetes lifecycle
One of the four platforms covered by my current Oracle infrastructure and deployment engineering responsibilities.
Contribute Terraform, Python and Bash automation to repeatable platform provisioning through OCI Resource Manager, with infrastructure changes represented and reviewed as code.
Support OKE lifecycle management within a multi-platform engineering remit, applying infrastructure and deployment expertise to the Kubernetes environment while collaborating with application owners.
Build and review infrastructure and GitLab CI/CD pipeline code for multi-region, multi-tenant platforms, connecting provisioning workflows with repeatable delivery practices.
ENGINEERING DEPTH / SHARED ORACLE RESPONSIBILITIES
Eight areas of hands-on work behind the four platforms. Shared responsibilities across my Oracle role, from infrastructure implementation and design review to release execution and production recovery.
I review Terraform changes, pipeline approvals, IAM boundaries and schema migrations as part of OCI DevOps and Shepherd delivery. Custom promotion gates validate architecture integrations and detect configuration drift between staging, pre-production and production.
Focus: infrastructure design review, deployment engineering and integration readiness.
I develop Terraform, Python and Bash automation with OCI Resource Manager for platform provisioning and OKE lifecycle management across ECP, ECP-RT, OSDMC and ICON.
Deliverables: provisioning automation, infrastructure changes and Kubernetes lifecycle workflows.
I build reusable GitLab CI/CD templates and shared libraries with security scans, linting and compliance checks. I own OCI DevOps and Shepherd production deployment and upgrade responsibilities within my role.
Deliverables: reusable GitLab templates, shared libraries and controlled promotion workflows.
I build Playwright-based Release Qualification automation and coordinate defect resolution. The work links release execution to integration behavior and schema-change validation.
Deliverables: RQ automation, qualification findings and coordinated release completion.
I build Streamlit uptime and cloud-cost dashboards, Grafana SLO/SLI monitoring, and OpenSearch/Kibana threshold monitors with Slack webhooks.
Reported result: threshold monitoring and webhook alerting reduced detection time by 45%.
I engineered an OpenSearch and Prometheus MCP correlation server with the SRE team and built custom Codex diagnostic workflows. I also built an AI review agent for infrastructure and deployment changes.
Reported results: at least 75% less triage time and 50% less initial response time.
I standardized multi-region disaster-recovery runbooks and failover tests, validating replication and recovery across OCI block storage and Oracle AI Autonomous Databases.
Deliverables: DR runbooks, failover tests and operational procedures.
I lead on-call incident, change and problem management. Telemetry correlation and cross-stack root-cause analysis connect the immediate response with longer-term operational improvement.
Reported results: MTTD down 50%, MTTR down 30%, and incidents resolved within SLA improving from approximately 82% to 99%.
ACROSS MY ORACLE ROLE
OCI DevOps and Shepherd promotion gates, IAM and schema-change review, reusable GitLab templates, Playwright RQ automation, SLO/SLI dashboards, MCP-assisted diagnostics, disaster-recovery runbooks and on-call leadership. Production actions retain human validation.
Gather context. Identify dependencies, ownership and the operational constraints that shape the design.
Implement infrastructure and deployment automation. Review assumptions, permissions and failure paths.
Validate configuration, integrations and service health. Keep promotion gates and recovery plans explicit.
Observe production, investigate with evidence, verify recovery and feed lessons into the platform.
From build and release foundations to multi-region platform ownership.
Complete career profile ↗Technology is the means.
Dependable systems are the outcome.
AWS · OCI · Azure · GCP
Terraform · Ansible · Python · Bash
Kubernetes · OKE · EKS · Helm · Docker
GitLab CI/CD · GitHub Actions · Jenkins · Argo CD
Prometheus · Grafana · OpenTelemetry · OpenSearch
SLOs · Incident response · DR · Release qualification
THE NEXT CHAPTER